Skip to content

ITERUMIT-er-um

Iteration Sprints

Iterum is Latin for “again,” the root of “iterate.” Each sprint improves your live system against criteria agreed in advance.

A live system that gets better, one agreed sprint at a time.

Is this for you?

  • AI features already in production, built by anyone
  • Systems that work but should work better
  • Teams that want predictable, fixed-price improvement

The situation

The system works. It should work better.

Your AI system is live, and your team relies on it. Some outputs need correction, some cases reach people more often than you would like, and the cost per accepted outcome could be lower.

Changes made without a baseline are hard to judge. A change can help one kind of case and quietly affect another, and the effect shows up later.

You want improvement you can see and approve, one change at a time, at a price agreed before the work starts.

Our approach

One metric, one baseline, one decision.

Each sprint starts from a documented baseline of the system as it runs today. We agree one metric, the test set, and the threshold a change must meet before release.

Each sprint builds one candidate change, such as clearer instructions, better context, a new check, or retraining where it pays. The candidate is measured against the baseline under the same conditions.

Your approver reads the evidence and decides. The change is accepted and released, or it is rolled back and the baseline stays in place.

Use cases by industry

Where this service fits.

Typical applications across industries. They show where the service applies, not past client work or results.

  • Software company

    Ticket routing accuracy in a support system

    The live system routes incoming tickets to product teams, and each misroute adds a handoff before the customer hears back. The sprint takes routing accuracy as its metric, measured on recent tickets with the correct team recorded. The candidate adds clearer category definitions, and the support lead approves the release only if the agreed threshold is met.

  • Marketing agency

    Campaign brief drafts accepted by account leads

    An agent drafts campaign briefs from client intake notes, and account leads rewrite many of them before work starts. The sprint measures the share of briefs accepted with light edits, using recent briefs and the leads’ corrections. The candidate supplies each client’s approved brand guidelines as context, and the operations director accepts or rolls back the change.

  • Home health

    Visit note summaries checked for required details

    The live system summarizes clinician visit notes for care coordinators, and the sprint metric is completeness of the required items. A clinical supervisor checks baseline and candidate on a sample of recent visits, and sensitive records stay inside the existing boundary. The candidate adds an automated completeness check, and the supervisor approves the release only if the agreed threshold is met.

  • Restaurant group

    Invoice line extraction for the food cost report

    An extraction system reads supplier invoices for the weekly food cost report, and line items with unusual units need correction. The sprint metric is field accuracy on line items, measured against invoices the finance team has already corrected. The candidate adds a check against each supplier’s price list and pack sizes, and the controller decides whether it is released.

  • Agriculture and food producer

    Cost per accepted order in email intake

    A live system reads customer orders from email and prepares them for the order desk to confirm. The sprint metric is cost per accepted order, with field accuracy reported alongside so quality stays visible. The candidate sends routine orders to a smaller model, and the order desk manager approves only if accuracy holds at the agreed threshold.

  • Veterinary group

    Appointment requests routed to the right clinic

    Requests from the website and phone transcripts are classified and routed across several clinics and service teams. The sprint measures routing accuracy on a test set of recent requests, with urgent cases checked separately. The candidate rewrites the classification instructions, and the practice manager approves the release after reviewing every urgent case in the comparison.

  • Title company

    Exceptions flagged in title commitment review

    An AI system reviews title commitments and flags items for an examiner, who confirms or clears each one. The sprint metric is whether the flagged exceptions match what examiners mark, measured on files already closed. The candidate adds a check for recurring exception types, and the senior examiner accepts or rolls back the change from the side-by-side report.

  • Equipment rental

    Quote drafts from a model trained on past quotes

    The live system drafts rental quotes, and sales staff correct the same kinds of terms repeatedly, such as damage waivers and delivery fees. Once instruction and context changes have been tried, a sprint can test a model trained on approved past quotes. The sales manager approves the release only if the share of quotes accepted without correction meets the agreed threshold.

What you receive

Agreed criteria up front

Each sprint starts with the improvements it must deliver and how they will be measured.

Measured improvement

Fixes driven by evaluation data, retraining where it pays, and new capabilities.

A sprint report

What changed, what the measurements show, and what to do next.

How it works

  1. Choose the measureSelect one operating metric and document the baseline.
  2. Scope the changeAgree the task, test set, and limits of the sprint.
  3. Build and measureCompare the candidate with the baseline on the agreed evaluation.
  4. Accept or roll backYour approver reviews the result and authorizes the release or rollback.

How success is measured

The measures your approver signs.

Each measure goes into the acceptance criteria with its test data, threshold, and the person who checks it.

The sprint metric
The one measure the sprint is set to move, such as accuracy, acceptance, routing, or cost. It is reported for baseline and candidate on the same test set.
Regressions in other measures
Whether the candidate made any other agreed measure worse. Each regression is reported with the cases behind it, so the approver sees the trade before deciding.
Cases that moved
Which individual cases changed, in either direction, between baseline and candidate. Reviewers read a sample from each group to confirm the change is real.
Cost per accepted outcome
What each accepted output costs before and after the change. It shows whether a gain in quality carries a running cost the business is willing to accept.
Rollback readiness
Whether the baseline can be restored through the agreed process. The rollback route is tested before release and kept ready until the change is accepted.

Where care is needed

What we watch, and how it is handled.

A fair comparison
Baseline and candidate are useful to compare only under the same conditions. We fix the test set, the inputs, and the settings before the candidate is built, and keep them fixed.
A test set that reflects today’s work
The work a system sees changes over time, with new products, policies, and customers. We refresh the test set from recent run records with your approver before each sprint.
Side effects of a change
A change aimed at one kind of case can affect another. We report every agreed measure alongside the sprint metric and show the individual cases that moved.
Systems built by another team
Many live systems were built by someone else. We document the baseline as it runs, learn the operating process from whoever runs it, and change nothing without approval.
Retraining where it pays
Retraining costs more than an instruction change and takes more care to roll back. We try simpler changes first and propose training when the evidence supports the cost.

Who does what

Your team decides. We engineer.

Your team

  • Name the approver who accepts or rolls back each change
  • Agree the metric and the threshold for each sprint
  • Provide access to the live system, its logs, and recent run records
  • Share corrections and feedback from the people who use the system
  • Decide what the next sprint should address

Sophrono

  • Document the baseline before any change is made
  • Find where the system misses, using run records and corrections
  • Build and test one candidate change within the agreed scope
  • Report baseline and candidate results under the same conditions
  • Keep the rollback route ready until the change is accepted

At the end

The decisions you make next.

The service ends with evidence and a choice. Each option is yours, and none is assumed.

  1. Accept and release

    Your approver accepts the change, it is released through the agreed process, and the new result becomes the baseline for the next sprint.

  2. Roll back and keep the baseline

    The change is withdrawn, the baseline stays live, and the sprint report records why the candidate fell short and what might work instead.

  3. Run another sprint

    Choose the next metric from the sprint report, or try a different approach to the same one with a fresh candidate.

  4. Move to ongoing care

    When steady upkeep matters more than separate sprints, AI Stewardship keeps the system monitored and current. Sprints can continue alongside it for larger changes.

What “accepted” means

Measured against criteria you agree to in advance.

  • The baseline and candidate use the agreed evaluation conditions.
  • The report states the change in the selected metric.
  • The agreed threshold is met before a release is accepted.
  • A human approver records acceptance or rollback.

Full engagement terms are finalized in a Master Services Agreement.

See a sample Acceptance Charter

Request a sprint conversation

Name the metric to improve.

One metric per sprint. We reply with what we need to see first: the live system and its records.

Describe the work in plain words. Please leave confidential records and passwords out.

  • A senior engineer reads every request
  • A reply by email with the next step
  • No obligation until scope and price are agreed

Not ready to scope this? Ask an engineer first: a free 15-minute call that names the agentic systems that could fit.

The Canon rule behind this service

Iteration Sprints answers to Canon VI.

Questions

What if the measure does not improve?

Then the change is not released. We review the evidence with you and agree whether another approach deserves a sprint.

Can you improve a system someone else built?

Yes. The first step is agreeing a baseline on the system as it runs today, so every sprint is measured against it.

How are fees agreed?

Scope, fees, and payment terms are documented for the engagement. Full terms are finalized in a Master Services Agreement.

Why one metric per sprint?

A single metric makes the decision clear. Any regression we observe in other measures is reported alongside the result.

What kinds of change can a sprint make?

Clearer instructions, better context or retrieval, new tools, added checks, a different model, or retraining. We try the simpler changes first, because they are easier to measure and roll back.

Can sprints run one after another?

Yes. Each sprint starts from a documented baseline, so every accepted change is measured against the system as it then runs.

Who changes the production system?

No change reaches production until your approver accepts it. Release and rollback follow the process agreed with whoever operates the system.

What do we need before the first sprint?

A live system, access to its run records, and an owner who can choose the metric. If no baseline exists yet, the first sprint can establish one.

How do sprints relate to AI Stewardship?

Sprints make agreed, measured improvements one at a time. Stewardship keeps the system measured and current between them, and the two can run together.

What makes a good sprint metric?

A measure your team already cares about that the run records can support. Accuracy, acceptance without rewriting, routing, and cost per accepted outcome are typical choices.

Who supplies the test set?

We build it with you from recent run records and the corrections your team made. Your approver agrees the test set before the candidate is built.

How long is a sprint?

Each sprint is a fixed-length cycle, and its length is agreed alongside the metric and scope. The fixed price covers that agreed scope.

Can a sprint add a new capability?

Yes, when the capability can be measured. The sprint agrees how it will be tested, and the release follows the same approval as any other change.

What does the sprint report contain?

What changed, the baseline and candidate results side by side, the cases that moved, and a recommendation for the next sprint. Your approver’s decision is recorded with it.

Do you need access to our production system?

We need access to the live system, its logs, and recent run records, under permissions agreed with whoever operates it. Candidates are tested outside production until your approver accepts a release.

Can a sprint lower running cost rather than raise quality?

Yes. Cost per accepted outcome can be the sprint metric, with quality measures reported alongside. The change is released only if quality holds at the agreed threshold.

What if two metrics matter equally?

Choose one for this sprint and report the other alongside it as a guard. The next sprint can take the second metric as its own.

What happens to the test set after a sprint?

It is documented with the sprint report and the baseline. Later sprints, and any later model change, can be measured against it.

Who needs to be involved on our side?

The approver, whoever operates the system, and the people whose corrections feed the test set. The approver’s time goes mostly into agreeing the metric and reading the report.

How do you find where the system misses?

We read the run records and the corrections people made, and group the misses by cause. The most common cause that one change can address becomes the sprint’s focus.

A live system that gets better, one agreed sprint at a time.

Book a time