From the archive
Whisper adds a faster turbo transcription model
Whisper turbo reduces the decoder size for faster transcription, while language performance and translation support differ.
Whisper maintainer Jong Wook Kim announced the large-v3-turbo model on October 1, 2024. The optimized model reduced the decoder from 32 layers to four to improve transcription speed.
A smaller decoder for speech recognition
The model retained multilingual transcription training while reducing the work required to generate the transcript. The maintainer reported performance similar to large-v2 across languages, with larger differences for some languages, including Thai and Cantonese.
The package version 20240930 or later included the model. The command-line tool also changed its default model to turbo.
Transcription and translation differ
Turbo's additional training excluded translation data. The announcement therefore described it as a transcription model and cautioned against expecting the same results for speech translation.
The release added a faster local transcription option, with model choice still depending on the language, recording quality, and task.