Google on Wednesday introduced Gemini 3.5 Transcribe, a speech-to-text model designed for real-time transcription with improved accuracy and latency.
The model achieves a 4.0% word error rate in streaming applications and 2.6% in non-streaming scenarios, according to Artificial Analysis benchmarks. Compared with the prior Chirp 3 model, it reduces the time to final transcription by 70%. On the FLEURS benchmark, the model records a 5.50% word error rate in streaming mode and 5.04% in non-streaming use cases.
Gemini 3.5 Transcribe offers two operational modes. The Live API provides sub-second latency for real-time streaming via the identifier `gemini-3.5-transcribe-live`, while the Interactions API processes pre-recorded audio such as meetings and call logs with speaker attribution using the identifier `gemini-3.5-transcribe`.
The model supports automatic transcription in over 85 languages and can identify up to three speakers in pre-recorded content with timestamps. Additional features include background noise suppression, technical terminology handling, filler word removal, automatic text formatting and custom vocabulary recognition for specialized terms.
Product integration includes the Gemini app on macOS, the Rambler feature on Android via Gboard and Google AI Studio. The model enters public preview for developers through Google AI Studio and for enterprises via the Gemini Enterprise Agent Platform. Chrome is slated to receive dictation capabilities in a future update.












