Google launches Gemini 3.5 Transcribe to further improve speech-to-text accuracy

📅 2026-08-27

Abstract:

Google recently released Gemini 3.5 Transcribe, the latest speech-to-text model in the Gemini series, calling it the most accurate speech recognition model in the series. This model is open to developers, enterprises, and ordinary users, and focuses on improving the transliteration capabilities in noisy environments, complex professional terms, and natural spoken expressions.

According to reports, Gemini 3.5 Transcribe can better handle background noise and complex vocabulary in professional fields. Google said the new model has brought performance improvements to the Gemini app on the macOS platform and the Rambler function of the Gboard keyboard on the Android platform. Developers can use this model to build applications such as voice agents, real-time subtitle tools, and post-call analysis processes.

Google said that Gemini 3.5 Transcribe is designed to capture users’ natural speaking patterns to more accurately understand semantic intent, identify custom vocabulary, and help users complete tasks through voice.

At the functional level, the new model adds a number of optimization capabilities, including automatic self-correction, deletion of spoken filler words such as "um" and "ah", and automatic formatting of output text. At the same time, it can also delegate more complex tasks such as file analysis and image generation to other Gemini models to form a more complete multi-model collaboration experience.

In terms of accuracy, Google pointed out that the average word error rate of Gemini 3.5 Transcribe is 4.0% in streaming speech recognition scenarios, and can be reduced to 2.6% in non-streaming scenarios. The model has better transcription performance in noisy environments and can more accurately identify key information composed of letters and numbers such as order numbers and postal codes.

In terms of language support, Gemini 3.5 Transcribe currently covers 85 languages ​​and supports regional accents and different dialects. For pre-recorded audio, it can perform time-stamped speech attribution recognition for up to three speakers; in addition, the model can adapt to user-defined vocabularies to identify professional terms, special spellings and industry names.

Gemini 3.5 Transcribe is currently available in public preview through the Gemini API in Google AI Studio and Google Antigravity. Ordinary users can experience related functions on Android and macOS platforms. Google also plans to introduce it to the Chrome browser, where users can enter text directly through voice in the input box of any web page.

Recently, Google has continued to accelerate the advancement of the Gemini product line. The company previously launched Gemini 3.7 Flash to join the increasingly fierce price competition for AI models; Gemini in Chrome for Android has also been opened to all users in the United States, and Gemini Assistant has also entered Waymo's new generation of self-driving taxis.

Related tags

Related articles

Comments

0/500
验证码
No comments yet