Small models, measured honestly, owned by the people who use them
ILLATE helps AI teams run their repetitive model work on small open models they own. We measure on real workloads before anything ships.
Renting your model is fragile
Many teams built a feature on a fine-tuned API model. The task was narrow, the model was small, and it worked. Then the provider retired the base model, changed the price or ended fine-tuning, and the feature broke on someone else's schedule.
For narrow tasks, a small open model trained on your own labels usually matches the API model, and you can keep it as long as you like. The hard part is proving it before you switch. That proof is what we build: an evaluation that you agree with, run on your data, with statistics that hold up.
- Company
- ILLATE
- Founded
- 2026
- Based in
- New Delhi, India
- Calls
- US and EU business hours
- Invoicing
- USD, via Wise or Payoneer
- Contact
- poojith@illate.dev
Poojith Devan
Founder and inference engineer · New Delhi
On every engagement you talk directly to the person who trains and serves your model. No account managers, no hand-offs.
- Built a ternary-quantized LLM runtime. A packed Triton kernel took decoding from 1.5 to 15.5 tokens per second and cut VRAM from 15.3 to 7.4 GiB.
- Fixed-point (Q1.15) inference for low-precision hardware
- Participant in AWS's Trainium kernel contest
- M.Sc. in Artificial Intelligence; MCA in Generative AI in progress
How we work
Evidence over claims
Every number on this site comes from a script you can rerun. Client reports include confidence intervals and the cases where our model is worse.
Say no early
If a cheaper API or a simple classifier already does the job, we say so in the free check, even though it costs us the work.
You can leave
You keep the weights, the eval set and the scripts. Hosting with us is a convenience, never a lock-in.
Talk to the engineer, not a sales team
Tell us what your fine-tuned model does. We reply within one business day.