Transcribe

POST the audio file as the raw body to /v1/rakeaudio/transcribe. Read segments, words, speakers and the full text back.

The request

POST https://api.milliseconds.ai/v1/rakeaudio/transcribe?diarize=true
PartValue
BodyThe audio file bytes, unchanged. Not JSON, not base64, not a URL
Content-Typeaudio/* (for example audio/mpeg, audio/wav, audio/mp4, audio/ogg) or application/octet-stream
AuthorizationBearer test_sk-... or Bearer prod_sk-.... See Authentication
diarize (query)false (default) returns the transcript without speakers. true labels speakers, at the higher price

The API decodes every common audio format: mp3, m4a, wav, ogg, opus, flac and webm. You do not need to convert the file.

curl -X POST "https://api.milliseconds.ai/v1/rakeaudio/transcribe?diarize=true" \
-H "authorization: Bearer test_sk-..." \
-H "Content-Type: audio/wav" \
--data-binary @call.wav

The SDKs cover decision-machine-1 only. Call rakeaudio with any HTTP client, as above.

The response

This is a real response to a 19-second support call with two speakers. The word lists are cut short.

{
"language": null,
"segments": [
{
"start": 0,
"end": 2.4,
"text": "Hello, thanks for calling support.",
"words": [
{ "word": "Hello,", "start": 0, "end": 0.72, "speaker": "SPEAKER_00" },
{ "word": "thanks", "start": 0.72, "end": 1.04, "speaker": "SPEAKER_00" },
{ "word": "for", "start": 1.04, "end": 1.2, "speaker": "SPEAKER_00" }
],
"speaker": "SPEAKER_00"
},
{
"start": 4,
"end": 10.24,
"text": "Hi, my invoice for September shows two charges of $49, and I only have one account.",
"words": [
{ "word": "Hi,", "start": 4, "end": 4.56, "speaker": "SPEAKER_01" },
{ "word": "my", "start": 4.56, "end": 4.8, "speaker": "SPEAKER_01" }
],
"speaker": "SPEAKER_01"
}
],
"duration": 19.0075,
"text": "Hello, thanks for calling support. How can I help you today? Hi, my invoice for September shows two charges of $49, and I only have one account. I see the duplicate charge. I will refund one of them now. You should see it in three to five business days. Great. Thank you very much for your help."
}

The full response held eight segments. The agent spoke as SPEAKER_00 and the customer as SPEAKER_01.

FieldMeaning
textThe whole transcript: the segment texts joined with spaces
segments[]Sentence-like parts of the transcript, in time order
segments[].start, endSeconds from the start of the file
segments[].speakerThe speaker who says most of the segment. Only with diarize=true
segments[].words[]Every word of the segment, with word, start and end
words[].speakerThe speaker of the word. Only with diarize=true
durationThe audio duration in seconds. Billing uses it
languageAlways null

With diarize=false, the response has the same shape without any speaker field.

Speaker labels are per call. SPEAKER_00 is the first voice in this file. It does not identify a person across files.

Response headers

HeaderMeaning
x-input-tokensThe input tokens billed for this call. The example above billed 15,840
x-inference-msThe model processing time in milliseconds. The example above took 282

Next