The contest over voice interfaces is shifting from "how well does it hear you" to "how cleanly does it write you down." Announcing Gemini 3.5 Transcribe on Aug. 26, Google said that unlike conventional speech recognition that struggles with background noise, jargon, and disfluency, the model "converts raw audio directly into accurate, polished, formatted text." It already powers first-party products including Rambler on Android's Gboard and the Gemini app on macOS.
Why Two Separate APIs
The model comes not as one endpoint but as two distinct APIs, because the jobs differ.
The real-time gemini-3.5-transcribe-live runs continuous, bidirectional streaming with sub-second latency through the Live API — aimed at voice agents and live captioning where instant response matters. The pre-recorded gemini-3.5-transcribe, by contrast, uses the Interactions API to transcribe meetings, call logs, and more with speaker attribution and word-level timestamps — suited to batch work like post-call analytics pipelines.
Streaming WER 4.0%
Languages 85+ auto-detected
Speaker attribution Up to 3 in pre-recorded audio (3+ experimental)
Real-time latency Sub-second
What Changed Versus Chirp 3
Google calls it a "major advancement" over 2025's Chirp 3. By the numbers, time to final transcription improves by 70%. On the multilingual FLEURS benchmark, across a set of top languages and locales, it posts 5.50% WER in streaming and 5.04% non-streaming, beating Chirp 3.
Beyond accuracy, the standout is task delegation. Through function calling, the model can hand complex jobs — image generation, file analysis — to other Gemini models. In the macOS Gemini app, that means summarizing a local file or generating an image at the cursor using only your voice.
| Metric | Gemini 3.5 Transcribe | vs. prior (Chirp 3) |
|---|---|---|
| Non-streaming WER | 2.6% (Artificial Analysis) | Improved |
| Streaming WER | 4.0% | Improved |
| FLEURS streaming WER | 5.50% | Beats Chirp 3 |
| Time to final transcription | — | 70% faster |
| Speaker attribution | Up to 3 (with timestamps) | Newly strengthened |
Where You Can Use It
Availability splits across developers, enterprises, and everyone else. Developers can access it in public preview through the Gemini API in Google AI Studio and in Antigravity. Enterprises get preview access via the Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon. Consumers can already try it in the Gemini app on macOS (English) and Rambler on Android, with a Chrome feature — talk to type in any web field — arriving next.
· Google — Intelligent transcription with Gemini 3.5 Transcribe (official announcement)
· Google — Gemini API transcription documentation
· 9to5Google — Google launches Gemini 3.5 Transcribe
· Engadget — Google's latest transcription model turns ramblings into structured text
- Google launched Gemini 3.5 Transcribe, its latest speech-to-text model, on Aug. 26
- 2.6% non-streaming / 4.0% streaming WER (Artificial Analysis); auto-detects 85+ languages
- Split into real-time (`-live`) and pre-recorded APIs; pre-recorded adds up to 3-speaker attribution with timestamps
- 70% faster time to final transcription than Chirp 3; function calling delegates work to other Gemini models
- Public preview for developers and enterprises; ships in Gboard Rambler, macOS app, and Antigravity, with Chrome coming