Skip to content

Note · Economics

Measure cost per accepted outcome

Compare the cost of an AI workflow against the outcomes your team accepts, including review effort, evaluation conditions, and the scope of the measurement.

What you will have

At the end of these steps you will have a working cost measure for one AI workflow. It shows what the workflow costs for each outcome your team accepts, with review effort and rejected work shown beside it.

The workflow owner uses it to compare options and decide what to change. Finance uses it to read AI spending in terms of finished work.

Before you start

Have these in hand before you count anything.

  • One workflow your team already performs and can judge.
  • Agreed acceptance criteria for the outcome, written before the comparison begins.
  • A named reviewer who checks each outcome against those criteria.
  • Access to the cost records for the workflow: service fees, infrastructure, and time records.
  • An evaluation period that every scenario will share.

If the criteria are not yet written, start with Agree the test before the build. The measure depends on them.

Step 1: Define an accepted outcome

Begin with a unit of work your team can recognize and evaluate. A reconciled invoice, a reviewed contract summary, or a resolved support request are typical units.

Agree the criteria before comparing approaches. Record who checks the outcome and which conditions apply.

A person approves any consequential action or production release associated with the workflow. That approval is part of what makes an outcome accepted.

Signs the unit is well chosen

  • Your team already produces it and can tell an acceptable one from one that needs rework.
  • It has a clear start and end, so it can be counted without debate.
  • One reviewer can judge it against the criteria in a single sitting.
  • It matters to the business, so a change in its cost is worth a decision.

Step 2: Draw the measurement boundary

Write down which costs the measure includes. Use the same boundary for every scenario you compare.

Review time, infrastructure, and service fees may have different owners. List each cost, its owner, and where its record lives.

Keep the comparison boundary explicit. A model’s usage charge alone does not describe every cost involved in delivering the accepted work.

Costs to consider inside the boundary

  • Engineering and delivery fees for the period
  • Model usage charges and hosting
  • Infrastructure and storage
  • Time your staff spend reviewing, correcting, and approving outputs
  • Time spent handling exceptions the system passes to a person

Costs outside the boundary are listed too, with a note on why. That keeps later readers from comparing figures built on different terms.

Step 3: Count accepted, revised, and rejected work

Our engagement measure is total engagement cost divided by the outcomes accepted under the agreed criteria. For one workflow, the same idea applies over an evaluation period.

Record three counts for the period, using the same reviewer and criteria throughout.

  • A, outcomes accepted as delivered
  • V, outcomes accepted after revision by a person
  • R, outcomes rejected

Track rejected or revised work beside the accepted count. Rejected work stays in the cost, so the measure reflects what acceptance took.

Step 4: Record the review effort

Record how much review the outcome requires. Time logged per review, or a sampled estimate agreed in advance, both work if applied consistently.

Write the total review time for the period as T, and its cost as T × [reviewer rate]. Include it inside the boundary from Step 2.

Review effort is easy to leave out because it sits in staff time, not on an invoice. A setup with lower usage charges may need more review, and the measure makes that visible.

Step 5: Calculate the measure

Add every cost inside the boundary for the period. Call the total C.

Divide by the outcomes your team accepted. Whether V counts toward the accepted total is a choice you make once and record.

  • Cost per accepted outcome = C ÷ (A + V), if revised work counts as accepted
  • Cost per accepted outcome = C ÷ A, if only work accepted as delivered counts

Report R, V, and T beside the result. Two setups with the same cost per accepted outcome can carry very different review loads.

Step 6: Change one decision at a time

This creates a question the team can investigate: which change improved the accepted result within the agreed cost boundary?

Change one thing per period: the model, the prompt, the business rule, or the review procedure. Rerun the same criteria over a comparable period, then compare.

Your workload owner approves each change before it reaches production. The result may support a new model, adaptation, or retaining the current setup, and the decision follows the agreed measure.

Keep a change log

Record each change beside the figures it produced. A short entry is enough.

  • The date and the one thing that changed
  • The person who approved the change
  • The resulting C, A, V, R, and T for the period
  • Any difference in conditions, such as a new type of case arriving

Check your work

Before you share the figure, confirm four things.

  • Were the acceptance criteria written before the comparison started?
  • Does every scenario use the same boundary, reviewer, and evaluation period?
  • Is review time inside the cost, with rejected and revised work reported beside it?
  • Could someone outside the team reproduce the figure from your records?

If the answer to any of these is no, record the difference next to the figure. A measure with stated limits is more useful than one without them.

Put it to work

Start with an AI Workload Evaluation to assess one workload’s applicability and next steps. It is free for one workload. Any deeper candidate comparison has a separately agreed scope.

How We Engage explains how this measure is reported across an engagement. It covers fixed-price and hourly work, the Acceptance Charter, and how changes are decided.

Full engagement terms are finalized in a Master Services Agreement.

Super Intelligence Newsletter

The frontier of AI, once a month.

The models, research, and releases that change what a business can build, with the sources and the question worth testing.

Read recent issues

One email a month. Unsubscribe anytime.

Related research

Acceptance

Agree the test before the build

Define the outcome, evaluation examples, acceptance measure, and human approver before building an AI workflow with Sophrono and your operating team.

Human approval

Design the human approval point

Give governed AI workflows a clear approval owner, explicit action boundaries, and review evidence before they connect to your business systems and records.

Build the system that compounds.

Book a time