Pillar 4 of 5
AI that improves through measured, approved changes.
Your people's reviewed work becomes the material the next version learns from. Every candidate is tested, and a person approves every release.
How does an owned model get better at your work?
Through a loop your people control. Each turn starts with real work and ends with a release decision made by a person.
- 01WorkCapture outcomes
- 02Human checkpointReviewPeople approve examples
- 03LearnCurate the dataset
- 04RetrainVersion the candidate
- 05TestEvaluate held-out work
- 06Human checkpointReleasePeople approve deployment
The loop turns daily work into evidence. Over time, your evaluation set grows to cover more of the cases your business sees, and each release is measured against it.
- WorkThe system handles real tasks, and each run records the decision, any correction, the outcome, and the cost.
- ReviewPeople accept, correct, or reject the output. Only reviewed examples move forward.
- LearnAccepted examples are curated into a versioned dataset with documented rights. Reviewed corrections also become tests.
- RetrainA candidate model is trained or configured on the new dataset version.
- TestThe candidate is measured on held-out work it has never seen, and on every earlier correction.
- ReleaseA named approver reviews the results and decides whether the candidate goes live.
Every run leaves a record
The loop learns only from what it can see. Each run writes down four things for review and testing.
The record keeps the decision with its context. A reviewer can see what the system did, why, and what it cost, without rebuilding the case from memory.
What each run records
- The decision: what the system produced or recommended, with its inputs.
- The correction: what a reviewer changed, and the reason given.
- The outcome: whether the work was accepted, corrected, or rejected.
- The cost: model and infrastructure use for that run.
Reviewed corrections become tests
When a reviewer corrects the output, that correction is evidence of what right looks like. After review, it is added to the evaluation set as a test case.
Every future candidate runs against those tests. A mistake your team corrected once is checked in every release that follows.
Tests are added and retired only with approval, so the evaluation set stays a deliberate record of your standards.
What Sophrono builds
The loop is engineering, not a promise. These are the parts that make it run.
Capture
Run records that keep input, context, decision, correction, outcome, cost, and time together.
A review queue
A place where reviewers accept or correct examples, with rights checked before anything is used.
Versioned datasets
Each dataset version records what was added, what was removed, and who approved it.
An evaluation harness
Held-out tests that measure every candidate on the same agreed criteria.
A release checklist
The evidence an approver needs, in one place, before any version goes live.
A rollback route
The previous accepted version, ready to restore if the new one underperforms.
This follows Canon rule VI: Every workflow teaches the next.
Use cases by industry
Where a learning loop fits
Typical applications, not past client work or results. Each one shows where reviewed work feeds the next version, and who approves.
Insurance agency
Submission intake
Account managers correct fields drawn from applications. Reviewed corrections become tests, and the agency’s operations lead approves each new version.
Medical billing
Denial categories
Billing specialists correct suggested denial reasons. Sensitive records stay inside the agreed boundary, and the billing manager approves each release.
Distributor
Order line matching
Representatives correct item matches on customer orders. New item codes enter the tests, and the sales operations manager approves each version.
Accounting firm
Transaction categories
Staff accountants correct suggested categories. A manager accepts corrections into the dataset and approves every release.
Property management
Maintenance triage
Property managers correct urgency calls against the written policy. Each reviewed correction becomes a test, and a named manager approves releases.
Law firm
Intake summaries
Attorneys edit draft matter summaries. Accepted edits become examples, and a supervising attorney approves each version before use.
Where people approve
Self-improving here means improving through changes people approve. The model learns only from reviewed examples, and every candidate is tested on held-out work.
Three decisions belong to named people: accepting examples and tests, releasing a new version, and restoring a previous one. None is delegated to the system.
What release acceptance means
- The dataset has documented rights and versions.
- The candidate clears the agreed held-out tests.
- Every earlier correction still passes as a test.
- The approver can inspect the results and limitations.
- A rollback route is ready before deployment.
- The Model Passport is updated with the new version.
How improvement is measured
Each measure is tracked on the same evaluation set, so a change in the number reflects a change in the system. Thresholds are agreed with you; none are published here.
Task quality
Acceptance rate against the same criteria, including the edge cases your reviewers flag.
Correction rate
How often reviewers change the output, and which kinds of change they make.
Earlier corrections
Whether every correction already turned into a test still passes.
Review effort
The time people spend checking and correcting the work.
Operating cost
Cost per accepted outcome under the agreed conditions.
Release history
Each version, its results, and its approver, recorded in the Model Passport.
Governance: history you can inspect
Every release leaves a record. The Model Passport lists each version with its data, results, known limits, and approver.
Research moves into the loop the same way. New techniques are adopted only after they clear technical, security, licensing, and economic review, with rollback preserved.
Ownership of datasets, models, and records is set in the Master Services Agreement.
See how new work is screened in Research.
Who runs the loop
Your team or an agreed stewardship team maintains the evaluation and release process. Your named approver authorizes each deployment.
AI Stewardship supports ongoing operation. Iteration Sprints focus on one agreed measure at a time.
Stewardship comes in four arrangements: Foundation, Continuity, Operations, and Custom. The scope of each is agreed with you before work begins.
When a loop is not worth running
A loop needs steady work and people to review it. For low-volume tasks, or when no one can approve releases, a periodic evaluation may serve you better.
If a rented model already clears your criteria, measure it on your evaluation set and keep it. The free AI Workload Evaluation helps decide.
How to start
Begin with one workload that already runs, or soon will. Each step is priced before it begins.
- Ask an engineerA free 15-minute conversation about the workload.
- Evaluate the workloadA free AI Workload Evaluation of fit and next steps.
- Capture and reviewRun records and a review queue go in place.
- Approve the first releaseA named approver reads the evidence and decides.
When working code would help first, the Agentic Engineering Diagnostic is $750. It includes one hour of agent coding, the code delivered, an Agentic AI Blueprint, and one hour of consultation.
Questions
Does every example enter training?
No. People review relevance, rights, and quality before accepting examples. Rejected examples can still inform the evaluation set.
How often does the model retrain?
On a schedule agreed for the workload, based on how much reviewed work accumulates. Each release still needs tests and approval.
How do you keep the model from learning a mistake?
Only reviewed examples enter the dataset, and each candidate is tested on held-out work. A person inspects the results before release.
Who decides whether an improvement is adopted?
A named person on your side. They decide only after the candidate measures better on the agreed tests, including every earlier correction.
What if the candidate performs worse?
Keep the accepted version. Review the results before deciding whether to change the dataset or the method.
Can we roll back a release?
Yes. The previous accepted version is ready before any deployment, and the Model Passport records each version.
Can a rented model be part of the loop?
Yes. The evaluation set measures any candidate, rented or owned. Prompts and configurations improve through the same tests and approval.
What data leaves our environment?
Only what the documented data boundary allows. The loop can run inside your environment when your policies require it.
Pillar 4 of 5