rakeaudio

Audio in, transcript out: word timestamps, up to four speaker labels, 25 European languages, in one synchronous call.

rakeaudio turns an audio file into a transcript. You send the file. You get the text back, split into segments, with a start and end time on every word. With diarization on, every word and every segment also carries a speaker label.

curl -X POST "https://api.milliseconds.ai/v1/rakeaudio/transcribe?diarize=true" \
-H "authorization: Bearer test_sk-..." \
-H "Content-Type: audio/mpeg" \
--data-binary @meeting.mp3

What you get

FeatureDetail
TranscriptThe full text, plus segments with start, end and text
Word timestampsEvery word in words, with start and end in seconds from the start of the file
SpeakersWith diarize=true (opt-in), up to four speakers, labeled SPEAKER_00 to SPEAKER_03
Languages25 European languages. The model detects the language. It does not report it: language is always null
InputOne audio file of up to 60 minutes, as the raw request body
CallSynchronous. The response holds the whole transcript

Languages

Bulgarian, Croatian, Czech, Danish, Dutch, English, Estonian, Finnish, French, German, Greek, Hungarian, Italian, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Russian, Slovak, Slovenian, Spanish, Swedish and Ukrainian.

One file can mix languages, for example English and French in one call.

How fast

A 19-second call with two speakers returned in 1.0 second. A one-hour file with diarization returned in 32 seconds, on 2026-09-27. Allow up to five minutes per call in your client timeout.

What it costs

ModePrice per audio hour
With speakers (diarize=true)$0.12
Without speakers (diarize=false, the default)$0.10

Billing uses the exact audio duration, rounded up to a whole token. See Pricing.

When not to use it

  • Your file is longer than 60 minutes. Split it into parts. Speaker labels are per call, so SPEAKER_00 in one part is not the same person as SPEAKER_00 in the next.
  • You need a live stream. The call takes a finished file. It returns when the whole file is done.
  • You need more than four speakers. Diarization labels at most four.
  • Your language is not in the list above. The model covers European languages only.