ILLATE
Benchmark · Banking77 · Engineering note

A 4B open model at 94% on 77-way support triage

This is the same test we run for clients, on a public dataset so you can check our work. It covers what we trained, how we measured it, what it costs to serve, and where it falls short.

Measured 5 Oct 2026NVIDIA L4 on ModalvLLM 0.31Qwen3-4B-Instruct-2507Parity: PASS
01 Summary

Results on 3,080 held-out messages

ModelAccuracy95% CIMacro-F1InvalidPrompt tokens
Qwen3-4B + LoRA, served by vLLM93.9%93.1–94.793.9%041
Qwen3-4B + LoRA, offline (Transformers)94.0%93.1–94.894.0%041
TF-IDF + logistic regression89.3%88.1–90.389.3%0–
Qwen3-4B, no fine-tuning, labels in prompt63.1%61.4–64.762.0%44437
Parity: PASS

The fine-tuned model scores 4.7 points above the free baseline (paired 95% CI +3.7 to +5.8). It was worse in 0.0% of bootstrap resamples. The pre-agreed non-inferiority margin was 3 points.

02 Setup

Data and training

Dataset. Banking77 (PolyAI, CC BY 4.0): 13,083 customer messages to a bank, each labelled with one of 77 intents. We kept the official 3,080-message test set untouched and split the official training set into 9,000 training and 1,003 development messages.

Model. Qwen3-4B-Instruct-2507 (Apache 2.0) with a rank-32 LoRA adapter on all attention and MLP projections. Two epochs, learning rate 2e-4, batch 16. Training took 22 minutes on one NVIDIA L4. Loss is computed on answer tokens only.

Prompt. One short instruction and the customer's message, 41 tokens on average. The fine-tuned model has learned the label set, so it doesn't need the 77 category names in every call. That's 10 times fewer input tokens than the untuned prompt.

Decoding. Outputs are constrained to the 77 valid labels, so the model can't invent a category. Without that constraint, our early smoke test produced 68 invalid labels in 200 answers.

03 Statistics

How we decide PASS or FAIL

A single accuracy number hides noise. Every report gives a 95% confidence interval from 10,000 bootstrap resamples of the test set. Comparisons use a paired bootstrap: both models are scored on the same resampled items, which removes the noise from item difficulty.

Before any training we agree a margin with the client, usually 2 or 3 points. The verdict is:

  • PASS when the lower end of the paired interval is above minus the margin.
  • FAIL when the upper end is below minus the margin.
  • INCONCLUSIVE otherwise, which means more test data is needed. We never round it up to a PASS.

Test data is held out from training and from every tuning decision. We cap ourselves at about ten experiment runs per client, so the final score isn't the luckiest of a hundred tries.

04 Errors

Where the fine-tuned model still misses

Most remaining errors sit between labels that real people also confuse. These are its most common mistakes on the test set.

True labelPredictedCount
fiat_currency_supportexchange_via_app5
card_arrivalcard_delivery_estimate4
balance_not_updated_after_bank_transfertransfer_timing4
topping_up_by_cardtop_up_reverted3
get_disposable_virtual_carddisposable_card_limits3

It loses most F1 against the free baseline on receiving_money (91.4% vs 97.4%, 40 items). A client report lists every class where the new model is worse, not only the total.

05 Serving

Latency and throughput on one L4

vLLM with the LoRA adapter behind an OpenAI-compatible endpoint, constrained to the 77 labels. Each row runs test messages at a fixed number of requests in flight.

Throughput · req/speak 86.6 at 32
0 25 50 75 100 1 8 16 32 64 requests in flight 4.8 29.7 50.1 86.6 56.2
p95 latency3 s past the knee
0 0.8s 1.6s 2.4s 3.2s 1 8 16 32 64 requests in flight 546 ms 3,045 ms
In flightThroughputp50p95GPU cost per 1M calls
14.8 req/s194 ms269 ms$46.49
829.7 req/s254 ms345 ms$7.49
1650.1 req/s310 ms485 ms$4.44
3286.6 req/s351 ms546 ms$2.57
6456.2 req/s780 ms3,045 ms$3.96

GPU cost assumes $0.80 per L4-hour and a fully used GPU. Above 32 requests in flight, constrained decoding over 77 labels becomes the bottleneck and latency climbs, so we cap concurrency per replica and add replicas instead.

06 Cost

When owning is cheaper, and when it isn't

One always-on L4 costs about $584 a month and can carry roughly 130 million of these calls. A team making 1.5 million short classification calls a month pays OpenAI about $480 for a fine-tuned gpt-4.1-mini, which is less than a dedicated GPU.

For small classification workloads, the reason to move is control, not price. Your model stops being deprecated and you keep the weights. We share one GPU across several clients' adapters, which brings hosting under what you pay today. High-volume or long-output workloads save outright.

Every engagement starts with a cost check on your real traffic: calls per month, tokens per call, peak rate. It also compares the cheap API options. If one of those passes your eval for less, we tell you to use it.

07 Limits

What this benchmark doesn't show

  • Banking77 is clean, single-label and English. Real tickets are messier, so we measure on your data before promising anything.
  • We didn't fine-tune a gpt-4.1-mini on the same data here. Your current model is the comparison that matters, and the parity check runs against its real outputs and your labels.
  • Latency was measured inside the same region as the GPU. Add your network round trip.
  • The first request after the GPU scales to zero takes about a minute. Production endpoints keep a warm replica.

Reproduce it

modal run scripts/modal_run.py --stage full      # train, eval, parity
modal run --detach scripts/modal_bench.py        # vLLM serving + load test

The scripts are in the ILLATE repository. Clients receive them with their report, so every number can be rerun.