From the archive

AssemblyAI releases Universal-2 for clearer transcripts

Universal-2 focuses on names, repeated numbers, formatting, and timestamps in speech-to-text output.

By Chat Overview Published Updated

AssemblyAI released Universal-2, a speech-to-text model focused on the details that make a transcript useful in an application. The changes address names, numerical strings, formatting, and word timestamps alongside overall transcription accuracy.

More than recognizing words

A call transcript needs to preserve a customer's name or account number, not just the general meaning of a conversation. Universal-2 changes how repeated digits are represented during recognition and upgrades the stage that adds punctuation, capitalization, and written number formats.

The practical focus includes:

  • Keeping repeated digits in phone numbers and identifiers.
  • Recognizing names, brands, and other uncommon words.
  • Formatting dates, prices, and sentences for readable output.

A hosted speech service

The release builds on AssemblyAI's existing hosted transcription pipeline. Its research announcement includes comparisons with Universal-1, rather than a claim that every recording will receive the same improvement. The AssemblyAI profile covers its speech services, deployment route, and usage-based pricing.

Source