rakeaudio
Audio in, transcript out: word timestamps, up to four speaker labels, 25 European languages, in one synchronous call.
rakeaudio turns an audio file into a transcript. You send the file. You get the text back, split into segments, with a start and end time on every word. With diarization on, every word and every segment also carries a speaker label.
What you get
Languages
Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish and Ukrainian.
One file can mix languages, for example English and French in one call.
How fast
A 19-second call with two speakers returned in 1.0 second. A one-hour file with diarization returned in 32 seconds, on 2026-09-27. Allow up to five minutes per call in your client timeout.
What it costs
Billing uses the exact audio duration, rounded up to a whole token. See Pricing.
When not to use it
- Your file is longer than 60 minutes. Split it into parts. Speaker labels are per call, so
SPEAKER_00in one part is not the same person asSPEAKER_00in the next. - You need a live stream. The call takes a finished file. It returns when the whole file is done.
- You need more than four speakers. Diarization labels at most four.
- Your language is not in the list above. The model covers European languages only.