PAIDEIApie-DAY-uh
Model Training & Fine-Tuning
Paideia was the Greek education that formed a person’s judgment. Our training forms a model’s judgment on your own data.
A model trained on your work, owned by you.
Is this for you?
- Repeated tasks with specialized language
- Businesses with high-quality examples of the work done well
- Workloads where consistency, privacy, or cost justify an owned model
The situation
A general model knows a great deal. Your work asks for something specific.
Your team handles a task that repeats, with its own vocabulary, formats, and judgment calls. General-purpose models cover much of it; the details particular to your business are where they need the most guidance.
You may also have reasons to run a model you control. Privacy, consistency, and the cost of calling a hosted model at volume all point in that direction.
The question is whether training a model on your own examples pays for itself, and how you would know if it did. You want that answered with evidence before the training budget is spent.
Our approach
Measure first. Train when the evaluation shows it pays.
We start with your data: its rights, its quality, and a held-out evaluation set that stays with you. That set measures your current approach first, then every candidate model on your own work.
Better prompting and retrieval are tried before any training begins. When they clear the agreed bar, the work can end there; when they do not, training is scoped on firm evidence.
When training goes ahead, an open-source or open-weight model is adapted to your data and measured on the held-out set. A Model Passport records the result, and a named person approves release.
Use cases by industry
Where this service fits.
Typical applications across industries. They show where the service applies, not past client work or results.
- Medical billing
Claim denial reasons classified for rework
Denied claims arrive with reason codes and free-text notes that billing specialists read, classify, and route for rework. Prompting and retrieval of the denial rules are measured first, on a held-out set of past denials the specialists labeled themselves. If a gap remains, an open-weight model is fine-tuned on approved examples, and a specialist still decides every rework.
- Wealth management
Meeting notes drafted in the firm’s house style
Advisers record client meetings in notes that follow the firm’s structure, vocabulary, and record-keeping conventions. Instructions and examples are measured first against notes the operations lead has accepted. If training is warranted, a model fine-tuned on approved notes runs where the firm controls it, and advisers review every note before filing.
- Architecture firm
Specification sections drafted from office standards
Specification writers adapt the firm’s standard sections to each project’s materials, assemblies, and performance requirements. A held-out set of completed sections measures retrieval of the office standards first, then any trained candidate on the same sections. Where a fine-tuned model closes the gap, a specification writer still reviews every draft before it enters the project manual.
- Software company
Support tickets summarized for the engineering queue
Support engineers turn customer tickets into clear, reproducible issue summaries for the product team, many times a day. The evaluation set pairs past tickets with summaries the engineers accepted. If prompting falls short, a smaller model is fine-tuned or distilled to lower running cost, and a support lead approves each version.
- Food producer
Quality inspection notes classified by defect type
Inspectors write free-text notes on incoming ingredients and finished lots, using terms particular to the operation and its products. Labeled past notes form the held-out set, and retrieval of the defect guide is measured before any training. A fine-tuned model can classify notes consistently across shifts and sites, and the quality manager approves every hold or rejection.
- Veterinary group
Discharge instructions drafted from visit records
Veterinarians write discharge instructions that follow the group’s conventions for medications, home care, and follow-up visits. Past instructions the veterinarians approved form the evaluation set, and the current drafting approach is measured on them first. If a fine-tuned model clears the agreed bar, a veterinarian still reviews and signs every set of instructions before the client leaves.
- Market research firm
Open-ended survey responses coded to client frameworks
Analysts code open-ended survey responses against each client’s coding framework, a repeated task with its own vocabulary. A held-out set of responses coded by senior analysts measures prompting first, then each fine-tuned candidate. The trained model proposes codes at volume, and an analyst reviews flagged and sampled responses before any results are reported.
- Mortgage broker
Loan documents classified and key fields extracted
Brokers receive pay stubs, bank statements, and disclosures in many layouts, and privacy obligations are reviewed before any example is used. The evaluation set pairs past documents with fields a processor confirmed. Where the model runs on hardware the broker owns, quantization is measured too, and a processor approves every extracted value before it is used.
What you receive
A trained model you own
An open-source or open-weight model adapted to your data, rules, and corrections.
An evaluation set
The test suite that measures this model, and every future one, on your work.
The method that fits the work
Sharper specifications, context, and retrieval come first. Training happens when the evaluation shows it pays.
How it works
- Prepare the dataReview rights, quality, and the held-out evaluation split.
- Choose the methodCompare prompting, retrieval, and training against the task.
- Train a candidateRecord data versions, settings, and the base model.
- EvaluateMeasure quality on the agreed held-out examples.
- Approve the releaseDeliver the Model Passport and obtain human release approval.
How success is measured
The measures your approver signs.
Each measure goes into the acceptance criteria with its test data, threshold, and the person who checks it.
- Quality on the held-out set
- Each candidate is scored on the evaluation set you keep, using the measures you agreed. The current approach is scored the same way, so every comparison is like for like.
- Agreement with your expert
- A named person compares a sample of outputs with what a correct result looks like. Their judgment confirms the scores reflect the work.
- Behavior outside agreed scope
- Cases outside the model’s agreed scope are tested to confirm they route to a person. The results are recorded in the Model Passport.
- Cost per accepted output
- Running cost is measured against the outputs a reviewer accepts. It is compared with the current approach at the volume you expect.
- Reproducibility
- The base model, data versions, and configuration are recorded for each candidate. The record is checked so the result can be rebuilt and reviewed later.
Where care is needed
What we watch, and how it is handled.
- Data rights
- Training needs the right to use each example for that purpose. We review the rights and origin of the data with you before any of it enters a training set.
- A fair evaluation
- A model is measured fairly only on examples it did not see in training. The held-out set stays separate, and it stays with you.
- Base model licenses
- The base model’s license shapes what you may do with the trained version. We choose bases whose terms suit your intended use and record them in the Passport.
- Sensitive records
- Some examples carry personal or confidential details. We agree how those details are removed or substituted, and where training runs, before training begins.
- Change over time
- Base models improve and your work changes. Each new version is measured on the same held-out set, and a named person approves each release.
Who does what
Your team decides. We engineer.
Your team
- Provide examples of the work done well, with the rights to use them
- Name a person who confirms what a correct output looks like
- Agree the quality measures used on the held-out set
- Keep the evaluation set and review each candidate’s results
- Approve any version before it is deployed
Sophrono
- Review data rights, quality, and the evaluation split with you
- Measure prompting and retrieval before proposing training
- Train candidates on open-source or open-weight base models
- Record each candidate in a Model Passport, including its limits
- Report every agreed measure, whether it favors training or not
At the end
The decisions you make next.
The service ends with evidence and a choice. Each option is yours, and none is assumed.
Release the model
Your approver reads the Model Passport, including its limits, and approves the version for deployment within its agreed scope.
Keep the improved prompting
When sharper instructions and retrieval clear the agreed bar, the work ends there. The training budget is kept for a later decision, and the evaluation set stays ready for it.
Improve against one measure
Iteration Sprints compare each new candidate with the accepted baseline, one agreed measure at a time. Each sprint’s criteria are agreed before it starts.
Keep it measured in production
AI Stewardship watches the agreed signals after release and prepares retraining when the evidence supports it. Your approver decides each release.
Before we start
What to have ready.
- A description of the task and what a correct result looks like
- Examples of the work, with notes on where they came from
- What you know about your rights to use that data for training
- Where the model will need to run, if you already know
- The person who will approve a model for release
Related services
What “accepted” means
Measured against criteria you agree to in advance.
- The dataset has a documented rights and version record.
- The held-out evaluation reports the agreed quality measures.
- The Model Passport identifies the model, configuration, and limits.
- A named person approves the version before deployment.
Full engagement terms are finalized in a Master Services Agreement.
Scope a training run
Describe the task and the data.
The data decides whether training is worth it. We reply with what we would assess first.
- A senior engineer reads every request
- A reply by email with the next step
- No obligation until scope and price are agreed
Not ready to scope this? Ask an engineer first: a free 15-minute call that names the agentic systems that could fit.
The Canon rule behind this service
Model Training & Fine-Tuning answers to Canon VI.
Where this fits
- Learn Research
- Explore free AI Workload Evaluation
- Try one workflow Diagnostic
- Plan and validate Model Training & Fine-Tuning
- Build Custom Agentic Systems
- Improve and operate AI Stewardship
Not ready yet? AI Workload Evaluation. After this: AI Stewardship or Iteration Sprints.
Questions
How much data is needed?
The task and variation in the examples determine the requirement. We assess the available material before setting a scope.
Which base models do you train?
Open-source and open-weight models, chosen for the task and for a license that suits your intended use. The base license shapes the rights you receive.
Why keep the evaluation set separate from the training data?
A model is measured fairly only on examples it did not see in training. Keeping the set with you means every future model is judged on the same work.
What does the Model Passport record?
The base model, the data versions used, the training configuration, and the known limits. It gives anyone reviewing the model later a plain record of what was released.
Does the model keep learning after release?
Only through a new version. New training data goes through the same evaluation, and a named person approves each release before it is deployed.
How is the price set?
The price is fixed, and it is scoped once the data is assessed. You see the scope and the price before training begins.
What if we do not have enough examples?
The data assessment shows whether the examples support training. If they do not, we recommend a better route, such as retrieval or collecting confirmed examples first.
Do we always end up training a model?
No. Sharper specifications, context, retrieval, a different model, or the approach you use today are measured first. Training happens when the evaluation shows it pays.
What are distillation and quantization?
Distillation trains a smaller model to reproduce a larger one’s results on your task. Quantization stores a model’s weights more compactly so it runs on less hardware. Both are measured on the held-out set like any other candidate.
Where does the trained model run?
Where your privacy, cost, and volume requirements point: on hardware you own or in an environment you control. Hardware is itemized separately, and you buy and own it.
Who owns the trained model?
You do, within the terms of the base model’s license. The Model Passport records the base and the terms that apply. Full engagement terms are finalized in a Master Services Agreement.
How do we choose the quality measures?
Together, from what a correct output means in your work. Your named expert confirms them before the baseline is measured, and every candidate is judged on the same measures.
Can a trained model use our documents as well?
Yes. Training and retrieval work together: training shapes how the model handles your task, and retrieval supplies current facts from your documents. Each combination is measured on the held-out set.
Who confirms what a correct output looks like?
A person you name, who knows the work well. They confirm the examples used for training and the answers in the evaluation set, so the measures reflect your standards.
Can we retrain the model later without Sophrono?
The Model Passport records the base model, data versions, and configuration, and the evaluation set stays with you. That record gives any capable team a clear starting point for the next version.
How is the model connected to our systems?
Integration is scoped as part of the work or as a separate build, depending on where the model will run. Either way, the boundary of what the model may do is agreed first.
What happens if a later base model is better?
A new base can be trained and measured on the same held-out set. It is released only if it clears the agreed measures and your approver signs off.