AssemblyAI
Speech transcription and audio analysis APIs, with different billing for files and streaming sessions.
- Recorded, streaming, and synchronous transcription
- Audio analysis and speaker-related options
- Streaming session time includes idle connections
AssemblyAI provides speech-to-text services and additional analysis for spoken material. Applications can submit a recording, transcribe an incoming stream, or use supported synchronous endpoints.
Selecting a transcription workflow
Recorded transcription fits completed audio files. Streaming fits live audio, and analysis options can add speaker information or extract structured information from speech. Available languages, features, and concurrency limits depend on the selected model and account.
How pricing works
Recorded transcription is billed by submitted audio duration. Streaming is billed for the time the session connection remains open, including idle time; close it when the conversation ends. Extra analysis features and multichannel audio can add charges. For local processing without a hosted transcription bill, see Whisper .