How it works
decision-machine-1 is not one large model. It is two small encoder models behind one API, plus a scheduler that keeps them busy.
Each call runs one pass over your text. It uses no sampling loop, no retry chain and no prompt parser. That is where the speed comes from.
The two models
The classifier scores labels. The extractor finds spans. Every capability maps to one of those two jobs.
Labels are the interface. The classifier reads your label text and scores it against the input. The words yes and no carry no meaning for it. See Writing good statements and labels.
What the classifier does
The classifier scores every label independently in the same pass, so adding labels adds their length to the input and nothing else.
The API then normalises those scores. Decisions and probabilities gives the exact formulas.
What the extractor does
The extractor finds the best span for each label you supply. extract sends your JSON Schema as a structure. answer sends each question as a label, plus one decoy label named other. The decoy sharpens the spans, and the API discards it. entities sends each type as a label. verify checks one field against one value.
The extractor is multilingual. The classifier performs best in English. See Languages.
One pass per call
Every capability call does the same three things:
The API builds the labels
The API turns your request into labels or structures. It validates the field limits first, so bad input fails in under 0.2 s.
Inference runs on a pool of GPUs, with CPU boxes as fallback. The scheduler tries a GPU slot first. Each GPU serves several calls in flight and batches classifier calls into one forward pass.
Chunking at 2,000 characters
Both models slow down past roughly 500 tokens. The pipeline splits any text over 2,000 characters on whitespace, so every chunk stays in the linear regime.
A single over-long token becomes its own chunk. extract also collapses runs of two or more spaces to " | " first, so table cells read as distinct spans.
Cost is linear in characters. Long text and chunking covers the practical guidance on pre-splitting.
Measured latency
Two sets of numbers. The first is the model time per call on the GPU pool, from the x-inference-ms header, short inputs.
Label count does not move the classifier time. Extra statements, questions or texts ride in the same batch.
The second set is end-to-end wall clock, measured from a laptop against production on 2026-09-18, median of 7 runs. Inputs were 51 to 199 characters. These numbers include TLS, internet transit and the API hop. The last column is the x-inference-ms header: the model time alone.
Every single-pass call lands near 0.4 s. The model takes 36 to 54 ms of that; the rest is the network and the API hop.
Batch the cheap axis. Extra statements or questions ride in one runner call and add only their own length to the billed input. Extra texts cost one call each. See Batching.
What you can measure yourself
Every capability response carries the input size and the model processing time in headers. Use them to predict cost, to size your batches and to watch latency.
x-input-tokens is the input tokens billed for this call. x-input-chars is the number of input characters, and it sums every item for texts. x-inference-ms is the model processing time for this request, summed over the model calls it made. The inference service measures it around the model itself, so it excludes queueing, network and the API layer. Compare it with your own wall clock to see how much of the total the network takes. The OpenAI-compatible routes report the request’s input tokens inside usage instead. Input tokens cost $0.04 per million and output tokens cost $0. See Pricing.
Backpressure
The scheduler answers 529 overloaded when no slot frees in time. A failed call marks its slot down for 30 seconds and retries once on another slot. A second failure returns 502 runner_error.
Both are transient. Retry with backoff. See Errors.
Next
- Long text and chunking — how chunking changes your results, and when to pre-split.
- Decisions and probabilities — the formulas behind every number.
- Limits and rate limits — field limits and the per-organization rate limits.