A 4B open model at 94% on 77-way support triage
This is the same test we run for clients, on a public dataset so you can check our work. It covers what we trained, how we measured it, what it costs to serve, and where it falls short.
Results on 3,080 held-out messages
| Model | Accuracy | 95% CI | Macro-F1 | Invalid | Prompt tokens |
|---|---|---|---|---|---|
| Qwen3-4B + LoRA, served by vLLM | 93.9% | 93.1–94.7 | 93.9% | 0 | 41 |
| Qwen3-4B + LoRA, offline (Transformers) | 94.0% | 93.1–94.8 | 94.0% | 0 | 41 |
| TF-IDF + logistic regression | 89.3% | 88.1–90.3 | 89.3% | 0 | – |
| Qwen3-4B, no fine-tuning, labels in prompt | 63.1% | 61.4–64.7 | 62.0% | 44 | 437 |
The fine-tuned model scores 4.7 points above the free baseline (paired 95% CI +3.7 to +5.8). It was worse in 0.0% of bootstrap resamples. The pre-agreed non-inferiority margin was 3 points.
Data and training
Dataset. Banking77 (PolyAI, CC BY 4.0): 13,083 customer messages to a bank, each labelled with one of 77 intents. We kept the official 3,080-message test set untouched and split the official training set into 9,000 training and 1,003 development messages.
Model. Qwen3-4B-Instruct-2507 (Apache 2.0) with a rank-32 LoRA adapter on all attention and MLP projections. Two epochs, learning rate 2e-4, batch 16. Training took 22 minutes on one NVIDIA L4. Loss is computed on answer tokens only.
Prompt. One short instruction and the customer's message, 41 tokens on average. The fine-tuned model has learned the label set, so it doesn't need the 77 category names in every call. That's 10 times fewer input tokens than the untuned prompt.
Decoding. Outputs are constrained to the 77 valid labels, so the model can't invent a category. Without that constraint, our early smoke test produced 68 invalid labels in 200 answers.
How we decide PASS or FAIL
A single accuracy number hides noise. Every report gives a 95% confidence interval from 10,000 bootstrap resamples of the test set. Comparisons use a paired bootstrap: both models are scored on the same resampled items, which removes the noise from item difficulty.
Before any training we agree a margin with the client, usually 2 or 3 points. The verdict is:
- PASS when the lower end of the paired interval is above minus the margin.
- FAIL when the upper end is below minus the margin.
- INCONCLUSIVE otherwise, which means more test data is needed. We never round it up to a PASS.
Test data is held out from training and from every tuning decision. We cap ourselves at about ten experiment runs per client, so the final score isn't the luckiest of a hundred tries.
Where the fine-tuned model still misses
Most remaining errors sit between labels that real people also confuse. These are its most common mistakes on the test set.
| True label | Predicted | Count |
|---|---|---|
| fiat_currency_support | exchange_via_app | 5 |
| card_arrival | card_delivery_estimate | 4 |
| balance_not_updated_after_bank_transfer | transfer_timing | 4 |
| topping_up_by_card | top_up_reverted | 3 |
| get_disposable_virtual_card | disposable_card_limits | 3 |
It loses most F1 against the free baseline on receiving_money (91.4% vs 97.4%, 40 items). A client report lists every class where the new model is worse, not only the total.
Latency and throughput on one L4
vLLM with the LoRA adapter behind an OpenAI-compatible endpoint, constrained to the 77 labels. Each row runs test messages at a fixed number of requests in flight.
| In flight | Throughput | p50 | p95 | GPU cost per 1M calls |
|---|---|---|---|---|
| 1 | 4.8 req/s | 194 ms | 269 ms | $46.49 |
| 8 | 29.7 req/s | 254 ms | 345 ms | $7.49 |
| 16 | 50.1 req/s | 310 ms | 485 ms | $4.44 |
| 32 | 86.6 req/s | 351 ms | 546 ms | $2.57 |
| 64 | 56.2 req/s | 780 ms | 3,045 ms | $3.96 |
GPU cost assumes $0.80 per L4-hour and a fully used GPU. Above 32 requests in flight, constrained decoding over 77 labels becomes the bottleneck and latency climbs, so we cap concurrency per replica and add replicas instead.
When owning is cheaper, and when it isn't
One always-on L4 costs about $584 a month and can carry roughly 130 million of these calls. A team making 1.5 million short classification calls a month pays OpenAI about $480 for a fine-tuned gpt-4.1-mini, which is less than a dedicated GPU.
For small classification workloads, the reason to move is control, not price. Your model stops being deprecated and you keep the weights. We share one GPU across several clients' adapters, which brings hosting under what you pay today. High-volume or long-output workloads save outright.
Every engagement starts with a cost check on your real traffic: calls per month, tokens per call, peak rate. It also compares the cheap API options. If one of those passes your eval for less, we tell you to use it.
What this benchmark doesn't show
- Banking77 is clean, single-label and English. Real tickets are messier, so we measure on your data before promising anything.
- We didn't fine-tune a gpt-4.1-mini on the same data here. Your current model is the comparison that matters, and the parity check runs against its real outputs and your labels.
- Latency was measured inside the same region as the GPU. Add your network round trip.
- The first request after the GPU scales to zero takes about a minute. Production endpoints keep a warm replica.
Reproduce it
modal run scripts/modal_run.py --stage full # train, eval, parity modal run --detach scripts/modal_bench.py # vLLM serving + load test
The scripts are in the ILLATE repository. Clients receive them with their report, so every number can be rerun.