Transcribe
POST the audio file as the raw body to /v1/rakeaudio/transcribe. Read segments, words, speakers and the full text back.
The request
The API decodes every common audio format: mp3, m4a, wav, ogg, opus, flac and webm. You do not need to convert the file.
The SDKs cover decision-machine-1 only. Call rakeaudio with any HTTP client, as above.
The response
This is a real response to a 19-second support call with two speakers. The word lists are cut short.
The full response held eight segments. The agent spoke as SPEAKER_00 and the customer as SPEAKER_01.
With diarize=false, the response has the same shape without any speaker field.
Speaker labels are per call. SPEAKER_00 is the first voice in this file. It does not identify a person across files.