Google Launches Gemini 3.5 Transcribe for Speech-to-Text
Google has released Gemini 3.5 Transcribe, a dedicated speech-to-text model built on its Gemini architecture. The model is positioned as an upgrade for developers who need to convert audio into text, whether for captioning, meeting notes, voice assistants, or accessibility tools.
Google says the new model improves on earlier transcription tools with better handling of accents, background noise, and multiple languages, while also being available through the Gemini API for developers to plug into their own apps.
The release fits into a broader trend of AI labs breaking out specialized models optimized for a single task, rather than relying solely on general-purpose chat models to handle transcription as a side feature.
Why it matters: Speech-to-text is a foundational building block for voice products, accessibility features, and content workflows, so incremental accuracy gains here ripple out to a lot of downstream apps. It also signals Google's strategy of unbundling Gemini into task-specific models to compete directly with specialized transcription vendors like Whisper-based startups.