Gemini 3.5 Transcribe Targets Faster Multilingual Speech Recognition

Google has released a new transcription model. Gemini 3.5 Transcribe brings lower error rates, faster results, and a collection of practical editing tools to speech-to-text applications. The rollout begins August 26, 2026, with Google placing the model inside products, services, and developer platforms.
Google reports an average word error rate of 2.6% for non-streaming transcription across 85+ languages, while streaming transcription reaches 4.0%. Those figures make Gemini 3.5 Transcribe faster and more accurate than Chirp 3, the previous voice-to-text engine.
The largest speed claim concerns the time to final transcription, which improves 70% over Chirp 3. That distinction matters because a transcript that arrives after the conversation ends is less useful than one that keeps pace with it. Speech recognition has spent years turning “almost live” into a product feature; Google is now trimming the “almost.”
Multilingual transcription with fewer cleanup tasks
Gemini 3.5 Transcribe automatically detects more than 85 languages and can handle code-switching in the middle of a sentence. That gives it a wider operating range than systems that expect speakers to stay inside one language from beginning to end.
The model can remove filler words such as “um” and “uh,” then edit the text on the fly. Users can also provide a customized vocabulary containing up to 1,000 terms, which helps the system recognize specialized jargon instead of forcing every unusual word through the same generic filter.
For pre-recorded audio, Gemini 3.5 Transcribe supports up to three speakers and provides word-level timestamps. Those timestamps connect individual words to their positions in the recording, while speaker support adds structure to conversations that would otherwise arrive as one undifferentiated block of text.
Google describes Gemini 3.5 Transcribe as a major advancement from Chirp 3, with improvements in multilingual performance and word error rates. The numbers are doing most of the promotional work here, which is probably healthier than another announcement built from adjectives.
Google products and developer access
Gemini 3.5 Transcribe is already live in Rambler on Pixel 11 and in the macOS Gemini app. Google plans to bring it to Chrome soon, then expand it to more devices and platforms.
Developers can access the model in public preview through Google AI Studio and Google Antigravity. That gives builders a route to test the transcription features, including streaming and non-streaming output, language detection, custom vocabulary, speaker handling, filler removal, and word-level timestamps.
The rollout adds another layer to Google’s Gemini product line, but the practical details matter more than the branding. A model that recognizes more than 85 languages, handles mid-sentence code-switching, and cuts final transcription time by 70% over Chirp 3 has a clear job: reduce the work between spoken words and usable text.
Google also promised the launch of Gemini 3.5 Pro in June 2026, though the provided rollout details for Gemini 3.5 Transcribe begin on August 26, 2026. For now, the transcription model has the clearer delivery date and the broader list of immediate product appearances.
Gemini 3.5 Transcribe does not just chase a lower error rate. It combines speed, language coverage, custom terminology, speaker support, and live text editing in one speech-to-text model—an unglamorous bundle of improvements that can make transcription far less annoying to use.
Based on
- Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages — marktechpost.com
- Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text – Ars Technica — arstechnica.com
- Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text | Ars OpenForum — arstechnica.com
- Google’s new AI transcription edits out your ‘ums’ and ‘ahs’ | The Verge — theverge.com




