How it works

Two small encoder models run on GPU, one pass over your text per call, with chunking at 2,000 characters.

decision-machine-1 is not one large model. It is two small encoder models behind one API, plus a scheduler that keeps them busy.

Each call runs one pass over your text. It uses no sampling loop, no retry chain and no prompt parser. That is where the speed comes from.

The two models

ModelJobCapabilities
The classifierScores every label you send in one pass over the textYes / no, Classify, Rate, and the tool pick in function calling
The extractorFinds the best span in the text for each label you sendAnswer, Extract, Entities, Verify

The classifier scores labels. The extractor finds spans. Every capability maps to one of those two jobs.

Labels are the interface. The classifier reads your label text and scores it against the input. The words yes and no carry no meaning for it. See Writing good statements and labels.

What the classifier does

The classifier scores every label independently in the same pass, so adding labels adds their length to the input and nothing else.

The API then normalises those scores. Decisions and probabilities gives the exact formulas.

What the extractor does

The extractor finds the best span for each label you supply. extract sends your JSON Schema as a structure. answer sends each question as a label, plus one decoy label named other. The decoy sharpens the spans, and the API discards it. entities sends each type as a label. verify checks one field against one value.

The extractor is multilingual. The classifier performs best in English. See Languages.

One pass per call

Every capability call does the same three things:

1

The API builds the labels

The API turns your request into labels or structures. It validates the field limits first, so bad input fails in under 0.2 s.

2

The scheduler leases a slot

A single scheduler holds every inference slot. It leases one call per slot and queues the rest. It opens the circuit when a box fails.

3

The model runs once

The model reads the text once and returns scores or spans. The API normalises, rounds to three decimals, and returns typed JSON.

Inference runs on a pool of GPUs, with CPU boxes as fallback. The scheduler tries a GPU slot first. Each GPU serves several calls in flight and batches classifier calls into one forward pass.

Chunking at 2,000 characters

Both models slow down past roughly 500 tokens. The pipeline splits any text over 2,000 characters on whitespace, so every chunk stays in the linear regime.

CapabilityBehaviour over 2,000 characters
yes-no, classify, rateA label’s score is its max over chunks. A claim that holds only in the last paragraph still scores high.
extractPer-chunk records merge. A single-record structure folds into one record, first non-empty value per field. A multi-record structure concatenates.
entities, verifyThe extractor runs overlapping windows and remaps the offsets back to your text.

A single over-long token becomes its own chunk. extract also collapses runs of two or more spaces to " | " first, so table cells read as distinct spans.

Cost is linear in characters. Long text and chunking covers the practical guidance on pre-splitting.

Measured latency

Two sets of numbers. The first is the model time per call on the GPU pool, from the x-inference-ms header, short inputs.

Call groupModel time
classify / yes-no / rate (classifier, any label count)~54 ms
entities / answer / extract / verify (extractor)~36 ms
classify-tree, 2 levels~108 ms

Label count does not move the classifier time. Extra statements, questions or texts ride in the same batch.

The second set is end-to-end wall clock, measured from a laptop against production on 2026-09-18, median of 7 runs. Inputs were 51 to 199 characters. These numbers include TLS, internet transit and the API hop. The last column is the x-inference-ms header: the model time alone.

CallObservedModel time
yes-no, one statement0.42 s54 ms
yes-no, 2 statements0.49 s107 ms
yes-no, 2 texts x 2 statements0.49 s214 ms
classify, 3 labels0.43 s54 ms
classify-tree, 2 levels0.77 s108 ms
rate, 4 levels0.42 s53 ms
answer, 1 question0.40 s36 ms
answer, 3 questions0.40 s36 ms
entities, 3 types0.42 s36 ms
verify0.40 s36 ms
extract, 9 properties0.41 s44 ms
extract, 2 texts0.44 s154 ms
chat json_schema0.39 s37 ms
chat tools, 2 tools0.72 s88 ms
chat forced tool_choice0.40 s36 ms
Every 400 / 404under 0.1 s

Every single-pass call lands near 0.4 s. The model takes 36 to 54 ms of that; the rest is the network and the API hop.

Batch the cheap axis. Extra statements or questions ride in one runner call and add only their own length to the billed input. Extra texts cost one call each. See Batching.

What you can measure yourself

Every capability response carries the input size and the model processing time in headers. Use them to predict cost, to size your batches and to watch latency.

curl -sS -D - -o /dev/null -X POST https://api.milliseconds.ai/v1/decision-machine-1/classify \
-H "Content-Type: application/json" \
-H "authorization: Bearer test_sk-..." \
-d '{
"text": "My card was charged twice this month.",
"labels": ["billing", "bug", "feature request"]
}' | grep -iE "x-input|x-inference"
x-input-chars: 37
x-input-tokens: 10
x-inference-ms: 412

x-input-tokens is the input tokens billed for this call. x-input-chars is the number of input characters, and it sums every item for texts. x-inference-ms is the model processing time for this request, summed over the model calls it made. The inference service measures it around the model itself, so it excludes queueing, network and the API layer. Compare it with your own wall clock to see how much of the total the network takes. The OpenAI-compatible routes report the request’s input tokens inside usage instead. Input tokens cost $0.04 per million and output tokens cost $0. See Pricing.

Backpressure

The scheduler answers 529 overloaded when no slot frees in time. A failed call marks its slot down for 30 seconds and retries once on another slot. A second failure returns 502 runner_error.

Both are transient. Retry with backoff. See Errors.

Next