Skip to content

Pillar 2 of 5

Open models, trained on your work, with weights you keep.

Choose the model through measurement, adapt it to a bounded task, and release it only after your approver accepts the results.

Which model should your work run on?

There are more capable models available than any team can test by hand. The useful question is narrower: which model clears your acceptance criteria for this task, at a cost and under terms that suit you?

Open-source and open-weight models add an option that rented models do not: weights you can keep, adapt, and run where you choose. This pillar is how one of them becomes yours.

Ownership also changes the operating picture. The model changes only when your approver releases a version, and it runs where your data rules allow.

What Sophrono builds

An owned model is more than a set of weights. It arrives with the evidence and records to operate it.

An adapted model

An open-source or open-weight model adapted to your task, your formats, and your rules.

An evaluation set

The held-out tests that measure this model, and every future candidate, on your work.

A Model Passport

The base model, license, data versions, evaluation results, known limits, and release approvals.

A deployment package

The model prepared for the environment you approved: your cloud account, your hardware, or a hybrid.

Typical applications

Bounded tasks an adapted model can own

Typical applications by industry, not past client work or results. Each is one task, measured on held-out work before release.

Insurance agency

Client email sorted by request type

A compact model adapted to the agency’s request types. An account manager reviews each routed request before anyone replies.

Medical billing

Remittance details extracted into billing fields

A model adapted to remittance layouts, running inside the agreed boundary for sensitive records. A billing specialist reviews flagged items.

Accounting firm

Transactions categorized to each client’s accounts

A model adapted to the firm’s categorization habits. A staff accountant approves each batch before it is posted.

Law firm

Clauses extracted for contract review

A model adapted to the clause types the firm tracks. An attorney checks every extraction.

Engineering consultancy

Review comments drafted in the house style

A model adapted to the firm’s comment conventions. The reviewing engineer approves or edits each comment before issue.

Manufacturer

Inspection notes classified by defect type

A model adapted to the plant’s defect vocabulary. A quality lead confirms each classification before it enters the quality record.

How a model becomes yours

The method follows the evidence. Training happens when the evaluation shows it pays.

  1. EvaluateCompare candidates against representative work.
  2. Choose the methodTry instructions, context, and retrieval before training.
  3. AdaptTrain on rights-cleared examples for the agreed task.
  4. TestMeasure the candidate on a held-out test set.
  5. ApproveYour approver reviews results before a version is released.

Adaptation can mean fine-tuning a model on your examples or distilling a larger model's behavior into a smaller one. It can also mean quantizing a model to run on less memory. Each is chosen against the same acceptance criteria.

A smaller model that clears the bar can be the better business choice. It may cost less to run and fit more environments.

Model Training & Fine-Tuning delivers this work at a fixed price, scoped once the data is assessed.

How a candidate is measured

Every candidate, open or closed, runs the same held-out tests on your work. Your approver agrees the measures and thresholds before testing, and results are reported measure by measure.

  1. Task accuracyWhether outputs match the expected answers your experts agreed on held-out work.
  2. Difficult casesHow the candidate handles incomplete, unusual, and out-of-scope inputs, including when it defers to a person.
  3. Comparison with todayThe same tests run on the rented model or manual process you use now.
  4. Speed where it will runTime per item, measured in the environment you approved.
  5. Cost per accepted outcomeWhat each accepted output costs to produce on that environment.

Where people approve

Engineering produces the evidence. Named people on your team make each decision a model depends on.

  1. The task and its acceptance criteriaThe workflow owner agrees them before any candidate is tested.
  2. The base modelIts license is reviewed against your intended use, and your approver accepts the choice.
  3. The training data versionYour data owner approves the examples and their permitted uses.
  4. Each releaseA named approver reviews the held-out results and the Model Passport before deployment.
  5. Rollback and retirementYour approver decides when a version is replaced. The previous accepted version stays available.

Governance: the license, the passport, the agreement

Available weights and commercial-use rights are different questions. Each base model's license is reviewed against your intended use before any work begins.

Ownership, license, and portability terms for the adapted model are set in the Master Services Agreement. The base model's own license continues to apply.

What your business keeps

Model Training & Fine-Tuning delivers a trained model you own. The base model’s own license continues to apply.

Full engagement terms are finalized in a Master Services Agreement. It sets portability and permitted uses for every deliverable.

  • The adapted model, prepared for the environment you approved.
  • The held-out evaluation set that measured it.
  • The Model Passport for each released version.
  • The training configuration and the data versions it used.

Choose around the task

Model families differ by the kind of work they handle. Each is evaluated on your own material.

  • Text and reasoning

    Classification, extraction, and drafting measured against task-specific criteria.

  • Vision and documents

    Images and document structure assessed on representative material.

  • Speech and retrieval

    Audio and search tasks evaluated against the operating requirements.

When a rented model is the better fit

Sophrono also works with closed frontier models across task types. For broad, changing, or low-volume work, a rented model may clear your criteria with less effort.

When that is the measured result, we recommend it. The evaluation set you built still measures the next candidate when the market moves.

Start with the free AI Workload Evaluation for one workload. Deeper model comparison is scoped separately, with senior engineering at $375 an hour.

Models for the work you need to do.

The directory is being prepared. We help with frontier models across task types, including open and closed models. Start with one workload evaluation.

Questions

Who owns the weights?

Model Training & Fine-Tuning delivers a trained model you own. The base model license still applies, and full engagement terms are finalized in a Master Services Agreement.

Are open-source and open-weight models good enough for our work?

It depends on the task. We measure candidates against your acceptance criteria and recommend whichever option clears them, open or closed.

Do you always fine-tune?

No. Clearer instructions, context, and retrieval come first. Training happens when the evaluation shows it pays.

What happens when a newer model is released?

Test it on your evaluation set. If it measures better, adapt it through the same process, and your approver decides whether to release it.

Can you adapt a model we already use?

Often, if its license permits adaptation for your intended use. We review the license first, then measure the adapted version like any other candidate.

Where does the model run?

In your cloud account, on hardware you buy, or in a hybrid arrangement. The choice follows your measured requirements.

What does the Model Passport record?

The base model, license, data versions, evaluation results, known limits, and release approvals. Each released version has one.

Can an owned model run alongside a rented one?

Yes. Different tasks can use different models, each measured on its own work and kept within the documented data boundary.

What if a released version measures worse later?

The evaluation set shows the change. Your approver can return to the previous accepted version while the cause is investigated.

Your data is your edge. Own the AI built on it.

Book a time