From the archive
OpenAI releases Whisper for multilingual transcription
OpenAI releases speech-recognition models and code for transcription, language detection, and translation into English.
OpenAI released Whisper, a speech-recognition system trained on 680,000 hours of multilingual audio. The release includes model weights and inference code, allowing transcription to run on a local machine or a self-managed server.
More than English transcription
Whisper handles several related tasks in one model:
- Transcribe speech in its original language.
- Identify the spoken language.
- Translate speech from supported languages into English.
The model processes audio in short chunks and produces text with task and timing information. That makes it useful for recorded interviews, subtitles, and searchable audio collections.
Running it locally
The release gives applications a way to process audio without sending it to a hosted transcription endpoint. Running the model still requires suitable hardware and a transcription workflow; downloadable weights do not remove those operating costs. Names, unfamiliar vocabulary, and difficult recordings still need a transcript check.
The Whisper profile covers the project and its deployment options. The OpenAI profile covers the company's hosted products and developer access.