V · Metron ariston · Measure is best
Reliability is engineered.
Context, tools, permissions, and validation determine production reliability.
Built from context, permissions, and automated validation.
In practice
What it means for your system.
Reliability comes from what surrounds the model. That means the context an agent receives, the tools it may use, the permissions it holds, and the checks on its output.
- Agents receive only the context and tools the task needs.
- Every change passes regression checks before release.
- Cost, quality, and outcome are traced for every run.
The word for it
Asphaleia
ἀσφάλεια
Asphaleia is Greek for “steadiness, literally not stumbling.” Reliability comes from checks designed into the system from the start.
Why it matters
Reliability is built around the model, not found inside it.
A demonstration shows what a model can do on a good day. Production asks what it does every day, with messy inputs, changing prompts, and new model versions.
Unreliable output has a cost that is easy to miss. Staff recheck the work, errors reach customers, and people start routing around the system they paid for.
So reliability is specified, built, and measured like any other requirement. A release goes out when it clears the tests the owner agreed, and the run records show whether it keeps clearing them.
Read it accurately
Scope of the rule.
- Reliability is a measured rate.
- Reliability is a measured rate against agreed tests. Exceptions are expected, recorded, and routed to a person.
- People review the exceptions.
- Automated validation checks each output first. People review the exceptions and approve consequential actions, so their time goes where judgment is needed.
- The test set grows with the work.
- The test set grows as new cases appear in production. Your owner approves each addition, so the bar reflects the work as it is.
How it is measured
Every rule is something you can check.
Regression gates on every change; cost, quality, and outcome traced for every run.
In an engagement
Where the rule is applied.
The rule is checked at each stage of the work, from the first design to the system in operation.
In the design
Reliability targets are written as acceptance criteria, with the test data that measures them. Each agent’s context, tools, and permissions are specified.
In the build
Validation gates and permissions are built alongside the features. Custom Agentic Systems traces cost, quality, and outcome for every run.
At acceptance
A Production Pilot tests one workflow against signed criteria in production conditions. Failure cases are documented with the exceptions they route to a person.
In operation
Regression gates run on every change, including model version updates. Run records show whether each workflow keeps clearing the agreed tests.
Check your own system
Five questions to ask this week.
Each one has a yes or no answer. A no marks where to start.
- Does each agent receive only the records its task requires?
- Is there a set of test cases every change must pass before release?
- Would a model version update be tested before it reached production?
- Can you see the cost of each run, as well as the total bill?
- Do you know what share of outputs your team accepts without rework?
Delivered through
The services that apply this rule.
Apply the rule
Check it against your own system.
Questions
Is reliability mainly about choosing a better model?
The model matters. Context, tools, permissions, and validation shape production results too, and those parts are yours to engineer.
Can you add this to agents we already run?
Yes. Agent Assurance adds evaluation, tracing, and regression gates to agents already in production, whoever built them. Owners approve the gates and the actions that need a person.
What happens when a gate fails?
The change does not ship. Engineers review the failing cases, and the owner decides whether to fix the change or update the test.
How do we check this rule on our own system?
Ask whether a regression gate runs on every change, including prompt and model updates. Then pick a recent run and look for its cost, quality, and outcome.
What goes into a regression gate?
Test cases drawn from your own work, each with its expected result. A change ships only when it clears them, and your owner approves any change to the cases.