From the archive

Whisper adds a faster turbo transcription model

Whisper turbo reduces the decoder size for faster transcription, while language performance and translation support differ.

By Chat Overview Published Updated

Whisper maintainer Jong Wook Kim announced the large-v3-turbo model on October 1, 2024. The optimized model reduced the decoder from 32 layers to four to improve transcription speed.

A smaller decoder for speech recognition

The model retained multilingual transcription training while reducing the work required to generate the transcript. The maintainer reported performance similar to large-v2 across languages, with larger differences for some languages, including Thai and Cantonese.

The package version 20240930 or later included the model. The command-line tool also changed its default model to turbo.

Transcription and translation differ

Turbo's additional training excluded translation data. The announcement therefore described it as a transcription model and cautioned against expecting the same results for speech translation.

The release added a faster local transcription option, with model choice still depending on the language, recording quality, and task.

Source