Skip to content

VI · Metron ariston · Measure is best

Every workflow teaches the next.

Every autonomous workflow should produce evaluation data.

Each run records the decision, the correction, the outcome, and the cost.

In practice

What it means for your system.

Every run records what was decided, what a person corrected, and what it cost. Reviewed by your people, that record becomes the test data for the next improvement.

  • Each run records its input, decision, correction, outcome, and cost.
  • Reviewed corrections become test cases for later changes.
  • An improvement is accepted when it measures better on those cases, and a person approves it.

The word for it

Empeiria

ἐμπειρία

Empeiria is Greek for “experience, knowledge gained by doing.” Each accepted run adds to what the next one knows.

Why it matters

Every correction is a lesson, if someone writes it down.

Your team already corrects AI output every day. Each correction costs staff time, and without a record, that time buys one fixed answer and nothing more.

A recorded run turns that review effort into an asset. It shows where the system is right, where it needs a person, and what each run costs to deliver.

That record also changes how improvement is funded. Instead of paying for changes on opinion, you pay for changes measured against your own past work.

Read it accurately

Scope of the rule.

Improvements arrive in reviewed batches.
Runs are recorded continuously, and improvements arrive in reviewed batches. A named person approves each one before release.
The record holds what the work needs.
The record covers what a run decided, changed, and cost. Sensitive fields follow your access rules and privacy obligations.
A correction is the most useful entry.
A correction is the most useful entry in the record. It shows where the system needs a person, and where the next improvement should start.

How it is measured

Every rule is something you can check.

Input, context, decision, action, correction, outcome, cost, time, and exceptions recorded for every run.

In an engagement

Where the rule is applied.

The rule is checked at each stage of the work, from the first design to the system in operation.

  1. In the design

    The run record is specified with the workflow, field by field. Reviewers get a place to note why they corrected an output.

  2. In the build

    Recording is built into each run, and the review screen captures correction reasons. Each record links to the version of the system that produced it.

  3. At acceptance

    Acceptance confirms that a complete record exists for every run. Reviewed records from the first runs form the starting test set.

  4. In operation

    Iteration Sprints turn recorded runs into measured improvements. Model Training & Fine-Tuning trains on reviewed corrections when the evaluation shows it pays, and your owner approves each release.

Check your own system

Five questions to ask this week.

Each one has a yes or no answer. A no marks where to start.

  • Can you retrieve the full record of any single past run?
  • Do reviewers record why they corrected an output, not only what they changed?
  • Does someone read the corrections as a set, looking for patterns?
  • Can you say what one run costs, including the review time it needed?
  • Is each improvement linked to the recorded runs that justified it?

Questions

Does recording every run slow the workflow down?

The record is written as the run happens, in the background. Reviewers add a correction reason in the screen they already use.

Who decides which corrections become improvements?

An engineer proposes changes from the patterns in the record. Your owner approves each one before it is released.

Where is the run record kept, and who owns it?

It is kept in systems you control, under your access rules. The record and the test cases built from it belong to you.

How do we check this rule on our own system?

Pick any recent run and ask for its record. It should show the input, context, decision, action, correction, outcome, cost, time, and any exceptions.

How many runs are needed before improvements begin?

There is no fixed count. Improvement starts once the reviewed record shows a clear pattern, with enough cases to measure a change against.

Does the record include the cost of review time?

Yes. Each record includes the time a run took, review included. The cost of one run then reflects the people involved as well as the model.

Build the system that compounds.

Book a time