ILLATE
API reference · v1

The endpoint API

Every migrated model runs on a dedicated endpoint that speaks the OpenAI chat completions format. If your code calls OpenAI today, it calls ILLATE after changing the base URL and key.

OpenAI-compatibleHTTPS · JSONServed by vLLMPer-client endpoints

Overview

Each client gets one hostname and one or more models behind it. Endpoints are issued at the end of a migration, once the parity report says PASS and you sign off. There is no public endpoint and no self-serve sign-up.

Base URL
https://<client>.api.illate.dev/v1
Format
OpenAI chat completions, request and response
Endpoints
POST /chat/completions, GET /models
Auth
Bearer key, one per environment
Self-hosting
The same stack deploys into your AWS, GCP or Azure account with your own base URL

Authentication

Send your key in the Authorization header. Keys are issued per environment (for example staging and production), can be rotated on request, and only work on your own hostname.

Authorization: Bearer $ILLATE_KEY

We never ask for your OpenAI, Anthropic or Google keys. Calls that need a frontier model stay in your code, with your key.

Chat completions

Send the same messages you send your current fine-tuned model. The prompt your model was trained with is fixed at migration, so most clients send only the user message.

POST /v1/chat/completions

{
  "model": "acme-triage",
  "messages": [
    {"role": "user", "content": "I ordered a card but it has not arrived."}
  ]
}

The response has the standard shape. message.content is always one label from your schema, or defer.

{
  "id": "chatcmpl-9f2c…",
  "object": "chat.completion",
  "model": "acme-triage",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "card_arrival"},
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 41, "completion_tokens": 4, "total_tokens": 45}
}

Supported request fields are model, messages, max_tokens, temperature, seed, logprobs, top_logprobs and user. Other OpenAI fields behave as they do in vLLM's OpenAI-compatible server.

Labels and defer

Decoding is constrained to your label set, so the model can't return a label that doesn't exist. In our Banking77 benchmark that meant 0 invalid answers in 3,080, against 44 for the same model without fine-tuning or constraints.

When the model's probability for its best label falls below a threshold, the endpoint returns defer instead of guessing. We set the threshold on your development set during migration, to a defer rate you choose (often 5 to 10%). Your code then takes its existing path:

r = client.chat.completions.create(model="acme-triage", messages=msgs)
label = r.choices[0].message.content
if label == "defer":
    label = classify_with_current_provider(ticket)   # your code, your key

Pass logprobs: true if you want the token probabilities behind each answer for your own monitoring.

Models and versions

GET /v1/models lists the models on your endpoint. Each retrain creates a new dated version, such as acme-triage@2026-11-02. The plain name points at the version you approved, and it only moves after a new PASS report and your written sign-off. Pin a dated version if you want to control the switch yourself.

No forced deprecations. A version you rely on keeps running until you ask us to retire it. The weights are yours either way.

Errors

Errors use OpenAI's JSON shape, so existing retry and logging code handles them unchanged.

{"error": {"type": "rate_limit_error", "code": "rate_limited",
           "message": "Above the agreed concurrency for this endpoint."}}
StatusMeaningWhat to do
400Malformed request, or input longer than the model's limitFix the request; don't retry
401Missing, wrong or rotated keyCheck the key for this environment
404Unknown model or versionCall GET /v1/models
429Above the agreed concurrencyRetry with backoff, or ask us to add a replica
5xxServer error or a replica restartingRetry with backoff, or fall back as for defer

Limits

Limits are set per endpoint from your measured peak, and written into the hosting agreement.

Concurrency
Up to 32 requests in flight per replica, where our L4 benchmark peaked at 86.6 req/s. Higher peaks get more replicas, not a deeper queue.
Latency
p95 of 546 ms at 32 in flight on one L4 for a 41-token prompt, measured in-region. Add your network round trip.
Input length
2,048 tokens per request by default; raised on request
Warm capacity
Production endpoints keep at least one warm replica, so there is no cold start
Timeout
Requests are cut off after 30 seconds

Numbers in this reference come from the Banking77 benchmark. Your endpoint's numbers are measured on your traffic before cutover and included in your report.