Transcribe audio

Send one audio file as the raw request body. Any format ffmpeg decodes works (mp3, m4a, wav, ogg, opus, flac, webm). The call is synchronous and returns the whole transcript, with word timestamps and, when `diarize` is true, speaker labels. One call takes up to 60 minutes of audio. Price: $0.12 per audio hour with diarization, $0.10 without, per second rounded up. The response carries `x-input-tokens` (input tokens billed) and `x-inference-ms`.

Authentication

AuthorizationBearer

A milliseconds API key, test_sk-{namespace}-{entropy} or prod_sk-{namespace}-{entropy}. Every /v1 route requires it.

Query parameters

diarizebooleanOptionalDefaults to false

Label speakers (SPEAKER_00 to SPEAKER_03). Default false. Send true for speaker labels, at the higher price.

Request

The audio file bytes. Content type audio/* or application/octet-stream.

Response

The transcript
languageany

Always null: the model does not report a language code.

segmentslist of objects
durationdouble
Audio duration in seconds. Billing uses it.
textstring
The segment texts joined with spaces.

Errors

400
Bad Request Error
415
Unsupported Media Type Error
502
Bad Gateway Error
503
Service Unavailable Error