ILLATE
Engineering notes

What we measured, and how

Short write-ups from our own runs. Each one gives the numbers, how they were produced and where they stop being true.

5 Oct 2026

A 4B open model at 94% on 77-way support triage

Banking77 end to end: data split, LoRA training, paired-bootstrap parity against a free baseline, serving latency on one L4, and the limits of the result.

→
5 Oct 2026

Zero invalid labels: constrained decoding for classification

Our first smoke test invented 68 labels in 200 answers. A token trie over the label set took that to zero, and set the concurrency limit we serve at.

→
5 Oct 2026

When a dedicated GPU beats the API bill, and when it doesn't

Break-even for short classification calls is about 1.8M calls a month on one L4. Below that, the reason to own a model is control, not price.

→

Want these numbers for your model?

The free parity check produces the same report on your data: accuracy with intervals, per-class results, latency and cost.

Get a free parity check