Skip to content

Pillar 1 of 5

Your data is the part nobody else can copy.

Anyone can download an open model. Only you have your documents, your decisions, and your corrections. We turn them into training and evaluation data that stays inside the boundary you set.

What makes a model yours?

Open-source and open-weight models are available to every business. What sets one system apart is the material it learns from and the tests it must pass.

That material already exists in most businesses. It sits in resolved requests, reviewed documents, and the decisions your experts make every day.

This pillar turns that work into an asset: examples with documented rights, and an evaluation set that measures any model on your terms.

The model is available to everyone. The examples of your work done well are available only to you.

It is Canon rule IV: Your data is worth more than you think. See what your data could be worth.

What counts as training data

Any record of the work done well, with the right answer and its basis, can support a specific task.

Documents and their outcomes

Contracts, forms, and reports paired with what your team decided about them.

Requests and resolutions

Tickets, emails, and cases alongside how they were resolved and why.

Expert decisions

The judgments your senior people make, with the reasons they give.

Corrections to AI output

Every edit a reviewer makes shows what the task requires. Corrections are often the most useful examples.

Rights, consent, and quality decide which examples are suitable. Examples without a clear basis for use stay out of the dataset.

Typical applications

Where useful examples already sit

Typical applications by industry, not past client work or results. They show the kind of record that becomes training and evaluation data.

Insurance agency

Coverage answers paired with policy records

Replies to coverage questions, each linked to the policy wording it relied on. An account manager confirms every expected answer.

Medical billing

Denials paired with the appeals that followed

Denial reasons matched with the appeals specialists wrote. Sensitive records stay inside the boundary your data owner approved.

Distributor

Purchase orders paired with corrected sales orders

Customer orders in many layouts, matched to the sales orders representatives entered. Their corrections often teach the most.

Engineering consultancy

Submittal reviews with specification references

Engineer comments on submittals, each tied to the specification section it cites. A senior engineer confirms the house standard.

Property management

Maintenance requests with assigned urgency

Resident requests paired with the urgency staff assigned under the written policy. A property manager settles disputed cases.

Accounting firm

Client documents with confirmed categories

Receipts, statements, and invoices paired with the categories staff accountants confirmed. A reviewing accountant approves the held-out test set.

How the material is prepared

Five steps turn working records into a dataset. The last step belongs to your data owner.

  1. FindIdentify representative work and the authoritative records behind it.
  2. CleanRemove duplicates, errors, and information outside the agreed scope.
  3. LabelRecord the expected answer and its basis, reusing existing review work.
  4. SplitHold back a test set the model never trains on.
  5. ApproveYour data owner approves the dataset and its permitted uses.

Labeling rarely starts from zero. Most businesses already review this work, so the expected answer often exists in a system of record.

The split step matters most for trust. A test set the model has never seen is what makes each result a fair measure.

Governance: rights, versions, and the boundary

Sensitive data needs an agreed processing boundary before anything moves. The preparation can run inside your own environment when your policies require it.

Data Protection & Readiness maps the rights and handling requirements. Ownership terms for the resulting datasets are set in the Master Services Agreement.

The evaluation set is an asset too

Of everything this pillar produces, the test set may last longest.

An evaluation set built from your work measures every future model, open or rented, against the same criteria. When a new model is released, you can test it on your work before anyone changes the system.

It also keeps model choice grounded. The decision to own or rent a model for a task rests on your own measurements, not on general rankings.

Treat it like any other business asset. Give it an owner, version it, and add new cases as the work changes.

  • Document the expected answers.
  • Record dataset versions and permitted uses.
  • Hold the test set apart from training.
  • Review each model change against the same agreed criteria.

How data readiness is measured

A dataset is ready when it measures well against criteria your data owner agreed. Thresholds are set per task before preparation begins, and each result is reported with its sample.

  1. Coverage of the workWhether the examples represent each case type the task meets, including the unusual ones.
  2. Agreement between reviewersWhether two experts give the same expected answer on a shared sample.
  3. Rights coverageWhether every example carries a documented basis for its intended use.
  4. Test set separationWhether any held-out example also appears in the training material.
  5. CurrencyWhether the examples reflect current policy, pricing, and practice.

Where people approve

Preparation is engineering work. The decisions about your records belong to named people on your team.

Before anything is copied

Your data owner approves which records are in scope and where processing happens.

Where records leave gaps

Your reviewers confirm the expected answer when the system of record does not settle it.

Before a dataset is used

Your data owner approves each dataset version and the uses it permits.

When the test set changes

A named approver accepts every change to the held-out set, so results stay comparable.

Corrections captured from live work follow the same rule. A person reviews each one before it becomes a candidate example.

What stays with your business

The work reads from your systems of record and leaves them in place. What it produces is documented so your team can use it without us.

Full engagement terms are finalized in a Master Services Agreement. That agreement sets ownership, portability, and permitted uses for each deliverable.

  • The prepared training dataset, with its version history.
  • The held-out evaluation set and its expected answers.
  • The record of sources, rights, and approved uses.
  • The preparation steps, written down so they can be repeated.

When training on your data is not the answer

Some tasks improve most through clearer instructions, better context, or retrieval from your records. Those come first, and they need no training run.

If a rented model already clears your acceptance criteria within your data rules, keep it. Your evaluation set still has value: it tells you when that changes.

When examples are limited

Start by measuring what you have. A focused evaluation shows which examples would help most, so collection effort goes where it counts.

How to start

Begin with one workload. Each step ends with a decision your team makes from the evidence.

When working code would help, the Agentic Engineering Diagnostic is $750. It includes one hour of agent coding, the code, an Agentic AI Blueprint, and one hour of consultation.

Questions

How much data do we need?

It depends on the task and the method. The evaluation measures what you have and shows which examples would help most.

Does our data leave our environment?

Only within the boundary your data owner approves. Preparation can run inside your environment when your policies require it.

Can we use personal or client data?

Only where rights and consent cover the intended use. Data Protection & Readiness maps those requirements before any data moves.

Who owns the datasets and evaluation set?

Ownership terms are set in the Master Services Agreement for each engagement. Each dataset records its sources and permitted uses.

Do our reviewers need to label everything again?

Usually not. Existing review decisions often hold the expected answer. We reuse that work and ask reviewers to confirm the gaps.

Can the evaluation set test rented models too?

Yes. It measures any candidate, open or closed, on the same held-out work and the same acceptance criteria.

What if our records disagree with each other?

That is common, and worth finding early. Labeling surfaces each disagreement, and your experts decide the expected answer.

Can we withdraw examples later?

Yes. Each example records its source, so your data owner can remove it and approve a new version. Retraining a model on that version is a separate decision.

Do you need access to our live systems?

Read access to the agreed records is usually enough. Access is agreed in writing, with the narrowest permissions the work needs.

Your data is your edge. Own the AI built on it.

Book a time