Skip to content

PAIDEIApie-DAY-uh

Model Training & Fine-Tuning

Paideia was the Greek education that formed a person’s judgment. Our training forms a model’s judgment on your own data.

A model trained on your work, owned by you.

Is this for you?

  • Repeated tasks with specialized language
  • Businesses with high-quality examples of the work done well
  • Workloads where consistency, privacy, or cost justify an owned model

The situation

A general model knows a great deal. Your work asks for something specific.

Your team handles a task that repeats, with its own vocabulary, formats, and judgment calls. General-purpose models cover much of it; the details particular to your business are where they need the most guidance.

You may also have reasons to run a model you control. Privacy, consistency, and the cost of calling a hosted model at volume all point in that direction.

The question is whether training a model on your own examples pays for itself, and how you would know if it did. You want that answered with evidence before the training budget is spent.

Our approach

Measure first. Train when the evaluation shows it pays.

We start with your data: its rights, its quality, and a held-out evaluation set that stays with you. That set measures your current approach first, then every candidate model on your own work.

Better prompting and retrieval are tried before any training begins. When they clear the agreed bar, the work can end there; when they do not, training is scoped on firm evidence.

When training goes ahead, an open-source or open-weight model is adapted to your data and measured on the held-out set. A Model Passport records the result, and a named person approves release.

Use cases by industry

Where this service fits.

Typical applications across industries. They show where the service applies, not past client work or results.

  • Medical billing

    Claim denial reasons classified for rework

    Denied claims arrive with reason codes and free-text notes that billing specialists read, classify, and route for rework. Prompting and retrieval of the denial rules are measured first, on a held-out set of past denials the specialists labeled themselves. If a gap remains, an open-weight model is fine-tuned on approved examples, and a specialist still decides every rework.

  • Wealth management

    Meeting notes drafted in the firm’s house style

    Advisers record client meetings in notes that follow the firm’s structure, vocabulary, and record-keeping conventions. Instructions and examples are measured first against notes the operations lead has accepted. If training is warranted, a model fine-tuned on approved notes runs where the firm controls it, and advisers review every note before filing.

  • Architecture firm

    Specification sections drafted from office standards

    Specification writers adapt the firm’s standard sections to each project’s materials, assemblies, and performance requirements. A held-out set of completed sections measures retrieval of the office standards first, then any trained candidate on the same sections. Where a fine-tuned model closes the gap, a specification writer still reviews every draft before it enters the project manual.

  • Software company

    Support tickets summarized for the engineering queue

    Support engineers turn customer tickets into clear, reproducible issue summaries for the product team, many times a day. The evaluation set pairs past tickets with summaries the engineers accepted. If prompting falls short, a smaller model is fine-tuned or distilled to lower running cost, and a support lead approves each version.

  • Food producer

    Quality inspection notes classified by defect type

    Inspectors write free-text notes on incoming ingredients and finished lots, using terms particular to the operation and its products. Labeled past notes form the held-out set, and retrieval of the defect guide is measured before any training. A fine-tuned model can classify notes consistently across shifts and sites, and the quality manager approves every hold or rejection.

  • Veterinary group

    Discharge instructions drafted from visit records

    Veterinarians write discharge instructions that follow the group’s conventions for medications, home care, and follow-up visits. Past instructions the veterinarians approved form the evaluation set, and the current drafting approach is measured on them first. If a fine-tuned model clears the agreed bar, a veterinarian still reviews and signs every set of instructions before the client leaves.

  • Market research firm

    Open-ended survey responses coded to client frameworks

    Analysts code open-ended survey responses against each client’s coding framework, a repeated task with its own vocabulary. A held-out set of responses coded by senior analysts measures prompting first, then each fine-tuned candidate. The trained model proposes codes at volume, and an analyst reviews flagged and sampled responses before any results are reported.

  • Mortgage broker

    Loan documents classified and key fields extracted

    Brokers receive pay stubs, bank statements, and disclosures in many layouts, and privacy obligations are reviewed before any example is used. The evaluation set pairs past documents with fields a processor confirmed. Where the model runs on hardware the broker owns, quantization is measured too, and a processor approves every extracted value before it is used.

What you receive

A trained model you own

An open-source or open-weight model adapted to your data, rules, and corrections.

An evaluation set

The test suite that measures this model, and every future one, on your work.

The method that fits the work

Sharper specifications, context, and retrieval come first. Training happens when the evaluation shows it pays.

How it works

  1. Prepare the dataReview rights, quality, and the held-out evaluation split.
  2. Choose the methodCompare prompting, retrieval, and training against the task.
  3. Train a candidateRecord data versions, settings, and the base model.
  4. EvaluateMeasure quality on the agreed held-out examples.
  5. Approve the releaseDeliver the Model Passport and obtain human release approval.

How success is measured

The measures your approver signs.

Each measure goes into the acceptance criteria with its test data, threshold, and the person who checks it.

Quality on the held-out set
Each candidate is scored on the evaluation set you keep, using the measures you agreed. The current approach is scored the same way, so every comparison is like for like.
Agreement with your expert
A named person compares a sample of outputs with what a correct result looks like. Their judgment confirms the scores reflect the work.
Behavior outside agreed scope
Cases outside the model’s agreed scope are tested to confirm they route to a person. The results are recorded in the Model Passport.
Cost per accepted output
Running cost is measured against the outputs a reviewer accepts. It is compared with the current approach at the volume you expect.
Reproducibility
The base model, data versions, and configuration are recorded for each candidate. The record is checked so the result can be rebuilt and reviewed later.

Where care is needed

What we watch, and how it is handled.

Data rights
Training needs the right to use each example for that purpose. We review the rights and origin of the data with you before any of it enters a training set.
A fair evaluation
A model is measured fairly only on examples it did not see in training. The held-out set stays separate, and it stays with you.
Base model licenses
The base model’s license shapes what you may do with the trained version. We choose bases whose terms suit your intended use and record them in the Passport.
Sensitive records
Some examples carry personal or confidential details. We agree how those details are removed or substituted, and where training runs, before training begins.
Change over time
Base models improve and your work changes. Each new version is measured on the same held-out set, and a named person approves each release.

Who does what

Your team decides. We engineer.

Your team

  • Provide examples of the work done well, with the rights to use them
  • Name a person who confirms what a correct output looks like
  • Agree the quality measures used on the held-out set
  • Keep the evaluation set and review each candidate’s results
  • Approve any version before it is deployed

Sophrono

  • Review data rights, quality, and the evaluation split with you
  • Measure prompting and retrieval before proposing training
  • Train candidates on open-source or open-weight base models
  • Record each candidate in a Model Passport, including its limits
  • Report every agreed measure, whether it favors training or not

At the end

The decisions you make next.

The service ends with evidence and a choice. Each option is yours, and none is assumed.

  1. Release the model

    Your approver reads the Model Passport, including its limits, and approves the version for deployment within its agreed scope.

  2. Keep the improved prompting

    When sharper instructions and retrieval clear the agreed bar, the work ends there. The training budget is kept for a later decision, and the evaluation set stays ready for it.

  3. Improve against one measure

    Iteration Sprints compare each new candidate with the accepted baseline, one agreed measure at a time. Each sprint’s criteria are agreed before it starts.

  4. Keep it measured in production

    AI Stewardship watches the agreed signals after release and prepares retraining when the evidence supports it. Your approver decides each release.

Before we start

What to have ready.

  • A description of the task and what a correct result looks like
  • Examples of the work, with notes on where they came from
  • What you know about your rights to use that data for training
  • Where the model will need to run, if you already know
  • The person who will approve a model for release

What “accepted” means

Measured against criteria you agree to in advance.

  • The dataset has a documented rights and version record.
  • The held-out evaluation reports the agreed quality measures.
  • The Model Passport identifies the model, configuration, and limits.
  • A named person approves the version before deployment.

Full engagement terms are finalized in a Master Services Agreement.

Scope a training run

Describe the task and the data.

The data decides whether training is worth it. We reply with what we would assess first.

Describe the work in plain words. Please leave confidential records and passwords out.

  • A senior engineer reads every request
  • A reply by email with the next step
  • No obligation until scope and price are agreed

Not ready to scope this? Ask an engineer first: a free 15-minute call that names the agentic systems that could fit.

The Canon rule behind this service

Model Training & Fine-Tuning answers to Canon VI.

Questions

How much data is needed?

The task and variation in the examples determine the requirement. We assess the available material before setting a scope.

Which base models do you train?

Open-source and open-weight models, chosen for the task and for a license that suits your intended use. The base license shapes the rights you receive.

Why keep the evaluation set separate from the training data?

A model is measured fairly only on examples it did not see in training. Keeping the set with you means every future model is judged on the same work.

What does the Model Passport record?

The base model, the data versions used, the training configuration, and the known limits. It gives anyone reviewing the model later a plain record of what was released.

Does the model keep learning after release?

Only through a new version. New training data goes through the same evaluation, and a named person approves each release before it is deployed.

How is the price set?

The price is fixed, and it is scoped once the data is assessed. You see the scope and the price before training begins.

What if we do not have enough examples?

The data assessment shows whether the examples support training. If they do not, we recommend a better route, such as retrieval or collecting confirmed examples first.

Do we always end up training a model?

No. Sharper specifications, context, retrieval, a different model, or the approach you use today are measured first. Training happens when the evaluation shows it pays.

What are distillation and quantization?

Distillation trains a smaller model to reproduce a larger one’s results on your task. Quantization stores a model’s weights more compactly so it runs on less hardware. Both are measured on the held-out set like any other candidate.

Where does the trained model run?

Where your privacy, cost, and volume requirements point: on hardware you own or in an environment you control. Hardware is itemized separately, and you buy and own it.

Who owns the trained model?

You do, within the terms of the base model’s license. The Model Passport records the base and the terms that apply. Full engagement terms are finalized in a Master Services Agreement.

How do we choose the quality measures?

Together, from what a correct output means in your work. Your named expert confirms them before the baseline is measured, and every candidate is judged on the same measures.

Can a trained model use our documents as well?

Yes. Training and retrieval work together: training shapes how the model handles your task, and retrieval supplies current facts from your documents. Each combination is measured on the held-out set.

Who confirms what a correct output looks like?

A person you name, who knows the work well. They confirm the examples used for training and the answers in the evaluation set, so the measures reflect your standards.

Can we retrain the model later without Sophrono?

The Model Passport records the base model, data versions, and configuration, and the evaluation set stays with you. That record gives any capable team a clear starting point for the next version.

How is the model connected to our systems?

Integration is scoped as part of the work or as a separate build, depending on where the model will run. Either way, the boundary of what the model may do is agreed first.

What happens if a later base model is better?

A new base can be trained and measured on the same held-out set. It is released only if it clears the agreed measures and your approver signs off.

A model trained on your work, owned by you.

Book a time