From a retiring model to one you own
Most migrations take three to five weeks. Nothing touches production until a parity report says PASS and you sign off on it.
What happens, week by week
Cost check
A 30-minute call about what the model does, how many calls it handles and how fast it must answer. We price every option, including cheap API models, with your real volumes.
Parity check
You send a few hundred labelled examples. We fine-tune an open model, test it on a held-out slice and send a PASS / FAIL / INCONCLUSIVE report with confidence intervals and per-class results.
Migration
- Train on your full labelled set, with a separate dev set for tuning decisions
- Agree the test set and margin in writing before the final run
- Replay a sample of your logged traffic offline and compare answers side by side
- Quantize and load-test on the GPU it will run on, at your peak rate
Cutover
Traffic moves in stages: 1%, 5%, 25%, then 100%. A small control group stays on your old path so drift shows up early. A kill switch returns all traffic in one config change.
You own it
You receive the weights, the eval set, the scripts that produced every number, and a runbook. Hosting with us is optional.
Inputs
- Labelled examples: inputs plus the answers you'd accept. The data you fine-tuned on is ideal.
- A sample of production inputs, with personal data removed if needed
- Your current model's outputs on the same inputs, as the comparison
- Volume, latency target and peak requests per second
- One engineer for about two hours a week
Deliverables
- Parity report with the verdict, intervals, per-class results and the worst errors
- Model weights (LoRA adapter and open base model) that you keep
- An OpenAI-compatible endpoint, hosted by us or deployed in your cloud
- Eval scripts and test set, so you can rerun every number
- Runbook: rollout steps, kill switch, monitoring and retraining
Good fit, and not
Good fit
- Classification, routing, tagging and extraction tasks
- A fine-tuned OpenAI model on a retirement schedule
- Labelled data you already own, or a team that can label a few hundred items
- A need to keep data in a region or off third-party APIs
Not a fit
- Open-ended chat or agents where frontier reasoning is the product
- Tasks with no way to say what a correct answer is
- Teams with an in-house ML platform team (you can do this yourselves)
Start with the free check
You'll know within a week whether your model can move.