ILLATE
How it works

From a retiring model to one you own

Most migrations take three to five weeks. Nothing touches production until a parity report says PASS and you sign off on it.

Timeline

What happens, week by week

Day 1 · Free

Cost check

A 30-minute call about what the model does, how many calls it handles and how fast it must answer. We price every option, including cheap API models, with your real volumes.

Week 1 · Free

Parity check

You send a few hundred labelled examples. We fine-tune an open model, test it on a held-out slice and send a PASS / FAIL / INCONCLUSIVE report with confidence intervals and per-class results.

Weeks 2–4 · Fixed fee

Migration

  • Train on your full labelled set, with a separate dev set for tuning decisions
  • Agree the test set and margin in writing before the final run
  • Replay a sample of your logged traffic offline and compare answers side by side
  • Quantize and load-test on the GPU it will run on, at your peak rate
Week 4–5

Cutover

Traffic moves in stages: 1%, 5%, 25%, then 100%. A small control group stays on your old path so drift shows up early. A kill switch returns all traffic in one config change.

Handover

You own it

You receive the weights, the eval set, the scripts that produced every number, and a runbook. Hosting with us is optional.

What we need from you

Inputs

  • Labelled examples: inputs plus the answers you'd accept. The data you fine-tuned on is ideal.
  • A sample of production inputs, with personal data removed if needed
  • Your current model's outputs on the same inputs, as the comparison
  • Volume, latency target and peak requests per second
  • One engineer for about two hours a week
What you get

Deliverables

  • Parity report with the verdict, intervals, per-class results and the worst errors
  • Model weights (LoRA adapter and open base model) that you keep
  • An OpenAI-compatible endpoint, hosted by us or deployed in your cloud
  • Eval scripts and test set, so you can rerun every number
  • Runbook: rollout steps, kill switch, monitoring and retraining
Fit

Good fit, and not

Good fit

  • Classification, routing, tagging and extraction tasks
  • A fine-tuned OpenAI model on a retirement schedule
  • Labelled data you already own, or a team that can label a few hundred items
  • A need to keep data in a region or off third-party APIs

Not a fit

  • Open-ended chat or agents where frontier reasoning is the product
  • Tasks with no way to say what a correct answer is
  • Teams with an in-house ML platform team (you can do this yourselves)

Start with the free check

You'll know within a week whether your model can move.

Book a parity check