Gemini 3.5 Transcribe: Google's New Speech-to-Text Model Hits 2.6% WER
Google introduces Gemini 3.5 Transcribe, a speech-to-text model with 2.6% WER, smart disfluency cleanup, and function calling, available via Live and Interactions APIs.

Google has quietly shipped a new speech-to-text model that's worth a look if you're building voice agents or transcription pipelines. Gemini 3.5 Transcribe, announced today, is positioned as a direct upgrade to Chirp 3, with better word error rates, 70% faster time-to-final-transcription, and a feature set that goes beyond raw transcription.
The model is available through two APIs: the Live API for real-time streaming with sub-second latency, and the Interactions API for pre-recorded audio with speaker attribution and word-level timestamps. That split is smart—it lets you pick the right tool for interactive voice apps versus post-call analytics.
What's new under the hood
The headline numbers come from Artificial Analysis: 4.0% WER for streaming, 2.6% for non-streaming. On the FLEURS multilingual benchmark, it hits 5.50% WER streaming and 5.04% non-streaming. Those are solid improvements over Chirp 3, especially for noisy environments and alphanumeric entities like postal codes.
But the real differentiator is smart transcription. The model handles self-corrections ("let's meet Tuesday—no, Wednesday"), strips filler words, and auto-formats text. It also supports custom vocabulary for specialized jargon, detects over 85 languages, and can attribute speech to up to three speakers (with experimental support for more).
Function calling and product integration
Gemini 3.5 Transcribe isn't just a transcription model—it can delegate tasks to other Gemini models via function calls. That's already live in the macOS Gemini app, where you can summarize files or generate images with voice commands. The model is also baked into Gboard's Rambler feature, Google Antigravity, and will hit Chrome soon for talk-to-type in any web field.
For developers, the model is accessible in Google AI Studio and the Gemini Enterprise Agent Platform. The Live API supports continuous bidirectional streaming, which is key for building voice agents that need to respond in real time.
Bottom line
If you're in the voice AI space, this is a meaningful upgrade. The WER numbers are competitive, the latency improvement is significant, and the function calling capability opens up new patterns for voice-driven workflows. The multi-speaker limitation (3 speakers, experimental beyond) might be a constraint for some use cases, but for most transcription needs, this looks like a strong option.
Gemini 3.5 Transcribe isn't just a transcription model—it can delegate tasks to other Gemini models via function calls, turning voice into a control plane for AI workflows.
| Metric | Gemini 3.5 Transcribe | Chirp 3 |
|---|---|---|
| WER (streaming) | 4.0% | ~5.5% (estimated) |
| WER (non-streaming) | 2.6% | ~5.0% (estimated) |
| FLEURS WER (streaming) | 5.50% | Higher |
| FLEURS WER (non-streaming) | 5.04% | Higher |
| Time to final transcription | 70% faster | Baseline |
Source: Google
Discussion
0 Comments
Be the first to start the discussion.