From the archive

ElevenLabs introduces Scribe speech-to-text

Scribe adds file transcription in 99 languages, with speaker labels, word timestamps, and audio-event tags.

By Chat Overview Published Updated

ElevenLabs introduced Scribe, its first standalone speech-to-text model. The launch added transcription alongside the company's speech-generation products, with support for 99 languages.

More structure in a transcript

Scribe returns information that helps turn an audio file into something searchable or editable:

  • Word-level timestamps for locating a passage in the recording.
  • Speaker labels for distinguishing participants.
  • Audio-event tags for sounds such as laughter.

The application programming interface (API) returns structured transcripts. The ElevenLabs dashboard also accepts uploaded audio and video files, giving editors a way to try the service without writing an integration.

File transcription at launch

The announcement concerned uploaded or submitted audio. ElevenLabs said a low-latency version for real-time applications would come later. File transcription and live conversational speech therefore had different availability at this point.

The ElevenLabs profile covers current speech products and billing. The Whisper profile provides a useful local-processing alternative for transcription workloads.

Source