Google Unveils Gemini 3.5 Transcribe AI for Real-Time, Accurate Speech-to-Text
Google has introduced Gemini 3.5 Transcribe, a new AI-powered speech-to-text model designed to convert spoken language into polished, readable text in real time. Unlike traditional transcription systems, this model automatically removes filler words and self-corrections, understands speaker intent, and supports over 85 languages and dialects, including full Hebrew support. It can also distinguish up to three different speakers and attach precise timestamps to each word, making it ideal for meetings and interviews.
The model achieves a low word error rate of 2.6% for file processing and 4% for live transcription, with a 70% faster response time compared to its predecessor. A notable feature is "Vibe Coding," which allows developers to write complete applications and code using voice commands within Google AI Studio. Gemini 3.5 Transcribe also integrates with Google Antigravity, cross-referencing speech with on-screen content and documents to ensure transcription accuracy.
In addition to transcription, the system can learn specialized jargon and technical terms, benefiting sectors like medicine, law, and finance. On Android devices, it powers the Rambler feature in the Gboard keyboard, enabling users to create and edit documents by voice. On Mac computers, it understands screen context and allows voice control of the mouse cursor, including scanning and summarizing open files and generating images based on cursor location.
Google has made the model immediately available to developers via Google AI Studio and Google Antigravity, as well as to enterprise customers through Gemini Enterprise APIs for live and file-based transcription. For individual users, the feature is currently accessible in English on Mac and Gboard for Android in select countries, with plans to expand to the Chrome browser for voice dictation across web text fields.