rakeaudio: audio transcription

A second model joins decision-machine-1 on the same API key. POST /v1/rakeaudio/transcribe takes an audio file as the raw request body. It returns the transcript with word timestamps and, by default, up to four speaker labels.

  • 25 European languages, detected from the audio.
  • Up to 60 minutes of audio per call, in one synchronous call.
  • $0.12 per audio hour with speakers, $0.10 without (diarize=false), billed as input tokens.

See rakeaudio and Transcribe.

A docs layout for several models

The docs now describe milliseconds.ai as a platform. Start here and Platform cover every model. decision-machine-1 and rakeaudio each have their own section. Existing page URLs do not change. The one exception is How it works, which moves to /concepts/how-it-works with a redirect.