Skip to content

Note · Acceptance

Agree the test before the build

Define the outcome, evaluation examples, acceptance measure, and human approver before building an AI workflow with Sophrono and your operating team.

The situation

An AI project needs a clear answer to one question: what must happen for your team to accept the work?

That question often arrives late. A system is built, a demonstration goes well, and then the team asks whether it is ready. Without a test agreed in advance, the answer depends on who is in the room and what they happen to try.

The moment that matters is earlier. It comes after the workflow is chosen and before any build begins. That is when the outcome, its test, and the person who decides can all be agreed in writing.

What we do

We write the test alongside the outcome, before the build is scoped or priced. Your team and ours agree four things for each deliverable.

The outcome, in the language of the work

Start with the business outcome. Identify the records that represent completed work and the person who can judge them.

A useful scope describes both the intended output and the conditions under which it will be used. “Draft a reply to each new supplier query” is a start. The scope also says which queries, from which inbox, and what the reply must contain.

The evaluation method and its examples

For each deliverable, record the evaluation method, the representative examples, and the acceptance threshold you will agree with the delivery team.

Keep those examples distinct from material used to adapt the system. The review should ask whether the system handles work it was not configured around.

A typical document workflow might check extracted fields against records your team has already reviewed. The client and delivery team agree the field definitions and the threshold together.

The acceptance threshold

The threshold states what result counts as accepted. It is written as a measure your team can check, such as [share of fields matching the reviewed record] at or above [agreed threshold].

The figure itself is a business decision. It reflects what the workflow is worth and what an error costs, so your owner sets it with our engineers.

The human approver

Agents can prepare work within assigned permissions. A named person approves consequential actions and the production release.

Document how an exception reaches that person. Specify which evidence accompanies the request and what happens when approval is withheld.

Why it works

Agreeing the test first changes the conversation from opinion to evidence. Each party knows what it is building toward and how the result will be read.

Scope becomes something you can price

A test with a threshold defines the work. That makes fixed-price scoping possible, because both sides know when a deliverable is complete.

When the scope is still open, the same discipline applies to hourly work. The backlog is reviewed against the outcomes the team is trying to reach.

Rework surfaces early

Writing the test exposes gaps in the brief. Missing field definitions, unclear exceptions, and unowned decisions appear before code exists.

Fixing a definition on paper costs a conversation. Fixing it after the build costs a revision cycle and a second review.

Review effort goes where it counts

Reviewers know what to check and in what order. They spend their time on the agreed examples and the exceptions, rather than on open-ended exploration.

Decisions have an owner

The person who signs the test is the person who accepts the result. That keeps the decision with someone who knows the work and answers for it.

It also gives the delivery team one clear point of contact for questions about scope. Disagreements are settled against the written record.

Later changes have a baseline

Record the version, evaluation conditions, results, and known limits. This creates a basis for reviewing later changes against the same accepted behavior.

When a model, prompt, or integration changes, the team reruns the same test. A person compares the new result with the accepted one and approves the change, or declines it.

How to apply it

You can use this practice on your next AI project, with or without us.

  1. Pick one workflow. Narrow the first test to a single unit of work your team already performs and can judge.
  2. Name the judge before the build. Choose the person whose acceptance ends the work. Confirm they have time to review the evidence.
  3. Set aside the examples. Collect representative cases, including difficult ones, and keep them out of any material used to configure the system.
  4. Write the threshold as a measure. Express it so someone outside the delivery team could run the check and reach the same result.
  5. Write down the withheld path. Agree what happens when the approver declines: revision, escalation to another person, or a change in scope.
  6. Sign it before work starts. A short written record, signed by your owner, is enough. Update it when scope changes, before the work changes.

Questions to test your draft

  • Could a new team member run the test from the written description alone?
  • Does every consequential action in the workflow have a named approver?
  • Are the evaluation examples separate from anything used to build the system?
  • Does the record say what happens when a result falls short?

If any answer is no, the test is not yet ready to sign.

Put it to work

Our System Blueprint connects this evaluation plan to the architecture. It names the systems of record, the human approvers, and the acceptance criteria, so the build can be priced against them.

The Sample Acceptance Charter shows how an outcome can be paired with a test, a threshold, and a signature. Use it as a starting structure for your own.

To see how the test connects to cost, read Measure cost per accepted outcome. For the approval side of the same workflow, read Design the human approval point.

Full engagement terms are finalized in a Master Services Agreement.

Super Intelligence Newsletter

The frontier of AI, once a month.

The models, research, and releases that change what a business can build, with the sources and the question worth testing.

Read recent issues

One email a month. Unsubscribe anytime.

Related research

Human approval

Design the human approval point

Give governed AI workflows a clear approval owner, explicit action boundaries, and review evidence before they connect to your business systems and records.

Economics

Measure cost per accepted outcome

Compare the cost of an AI workflow against the outcomes your team accepts, including review effort, evaluation conditions, and the scope of the measurement.

Build the system that compounds.

Book a time