Enterprise Tech

Google Launches Gemini 3.5 Transcribe, a Smarter Speech-to-Text Model

Google launched Gemini 3.5 Transcribe, a speech-to-text model supporting 85+ languages that cleans up filler words and self-corrections, ranking fifth on an independent accuracy leaderboard.

By Daniel Mercer Edited by Maria Konash Published:
Google Launches Gemini 3.5 Transcribe, a Smarter Speech-to-Text Model
Google launched Gemini 3.5 Transcribe, a speech-to-text model supporting over 85 languages that scored 2.6% word error rate on an independent benchmark. Image: Google

Key Notes

  • Google launched Gemini 3.5 Transcribe (plus a Live streaming variant), supporting 85+ languages with auto-detection and mid-conversation language switching.
  • On Artificial Analysis's benchmark it scored 2.6% WER (non-streaming) and 4.0% WER (streaming), placing 5th on the leaderboard.
  • It replaces Chirp 3, with Google citing a 70% latency improvement and sub-second (0.40s) streaming response.

Google launched Gemini 3.5 Transcribe on August 26, a new speech-to-text model the company describes as its most precise yet for intelligent voice interactions. Rather than simply converting audio to raw text, the model is designed to interpret how people actually speak, handling self-corrections like “let’s meet Tuesday, no, Wednesday” by keeping the intended correction, stripping filler words like “um” and “ah,” and auto-formatting the output into clean, readable text.

The model supports more than 85 languages with automatic detection, including mid-conversation language switching and regional accents, and accepts custom vocabulary so it can correctly render company names, product codes and specialized jargon. For pre-recorded audio, it attributes speech to individual speakers with word-level timestamps, officially supporting up to three speakers, with performance beyond that considered experimental.

Google made the case for the model’s accuracy using figures from the independent benchmarking firm Artificial Analysis, citing an average word error rate of 4.0% for streaming transcription and 2.6% for non-streaming use, and a 70% improvement in time to final transcription over Google’s previous model, Chirp 3. On Artificial Analysis’s own public leaderboard, Gemini 3.5 Transcribe’s non-streaming version placed fifth overall with that 2.6% error rate, a strong and independently verifiable result, though a distinct claim from an outright “most accurate” title. On the multilingual FLEURS benchmark, Google reported 5.50% WER in streaming mode and 5.04% in non-streaming.

The model ships in two forms: gemini-3.5-transcribe-live, accessed through the Live API, delivers continuous, sub-second streaming transcription with a reported response latency of about 0.40 seconds after speech ends, aimed at interactive voice agents and live captioning; gemini-3.5-transcribe, accessed through the Interactions API, is built for recorded audio like meetings and call logs, adding speaker attribution and timestamps. It is available now in public preview through Google AI Studio and the Gemini Enterprise Agent Platform.

Consumers are already encountering the model without necessarily knowing it. It powers the new Rambler dictation feature in Gboard on Android, voice commands in the Gemini app on macOS, and transcription inside Google’s Antigravity coding tool, with support coming soon to Chrome for dictating directly into any web text field.

Why It Matters

The launch pushes Google directly into a speech-API market already contested by OpenAI, ElevenLabs, Deepgram and AWS, each competing for the same developer and enterprise voice-agent contracts. Because transcription quality and latency directly determine whether a voice assistant or live-captioning tool feels responsive or sluggish, Google’s decision to publish independently sourced benchmark comparisons, rather than only self-reported figures, is a deliberate move to compete on credibility as much as raw capability.

Folding transcription into the same API surface as the rest of the Gemini line, with speaker attribution and timestamps built into the base model rather than sold as a separate add-on, also signals Google’s broader ambition to position voice as a primary interface for its AI products, not merely another input method layered onto text.

A Caveat Worth Keeping

It is worth noting the benchmark figures Google is citing, while run by an outside firm, were selected and published inside Google’s own announcement of a model that had been publicly callable for only a day at the time of launch. That does not make the numbers unreliable, but it is a different standard of verification than a fully independent, post-launch third-party audit, and everything described here, including model IDs, pricing and benchmark scores, remains part of a public preview that could change before general availability.

Real-world performance on poor microphones, overlapping speakers and unpredictable background noise, the conditions most transcription tools actually struggle with, will be the more meaningful test over time.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Enterprise Tech, News