Alphabet’s Google on Wednesday launched Gemini 3.5 Transcribe, a voice-to-text model designed for real-time transcription with sub-second latency and improved accuracy over its predecessor, Chirp 3.
The new model delivers a 70% reduction in time to final transcription compared with Chirp 3, while maintaining a streaming word error rate of 4.0% and a non-streaming rate of 2.6%, according to Artificial Analysis. In benchmark testing on the FLEURS dataset, it recorded a streaming error rate of 5.50% and a non-streaming rate of 5.04%.
Gemini 3.5 Transcribe is available through two API interfaces: the Live API for real-time streaming with latency of less than one second, and the Interactions API for pre-recorded audio such as meetings and call logs. The latter includes speaker identification for up to three voices with time-stamped transcripts.
The model supports automatic language detection across more than 85 languages and incorporates features including filler-word removal, automatic text formatting, and custom vocabulary recognition for specialized terminology. It is entering public preview for developers via Google AI Studio and for enterprises through the Gemini Enterprise Agent Platform.
Google plans to integrate the technology into the Gemini app on macOS, the Rambler feature on Android via Gboard, and a future update to Chrome’s built-in dictation tool. The company did not disclose pricing or commercial availability timelines.












