Article
AI in Your Apps AI News

Google rolls out Gemini 3.5 Transcribe, an AI that cleans up your spoken text

The dedicated speech-to-text model that removes disfluencies, edits on the fly and is being pushed into Google apps and developer tools.

by Whatsnew Newsroom
The image depicts a pink Sony voice recorder placed on lined notebook paper. The device features buttons on its side, indicating its functionality for recording audio. — Credit: Photo by Adhitya Sibikumar on Unsplash c Photo by Adhitya Sibikumar on Unsplash

Google has released Gemini 3.5 Transcribe, a speech-to-text model that aims to turn messy spoken input into polished, editable text rather than a verbatim dump.

Google says the new model should be about 70% faster from voice to final transcribed text and that its live-speech error rate is 5.5%, compared with Chirp 3 at 7.32%. The model also removes fillers, the “ums” and “uhs”, and can rewrite on the fly when you self-correct, using a provided custom vocabulary to handle specialised terms.

That editing is the point: Gemini 3.5 Transcribe moves transcription from verbatim capture toward a comprehension-driven result. It helps for short blocks of dictated copy and can be faster to clean up than manual correction, but it does mean the AI may change your exact wording, a drawback where an authoritative, word-for-word record is required.

The model supports 85 languages and handles up to three speakers in pre-recorded audio, according to the announcement. It is already powering Gboard’s Rambler feature on the Pixel 11, and Google says Rambler will reach more Gemini Intelligence devices later this year. Desktop users get an immediate benefit: the Gemini app on macOS has the model today.

Developers and partners will see Transcribe surface quickly. The Antigravity app will expose the model with full access to screen context and chat history with permission, AI Studio’s build model includes it, and the model is available via the Gemini API. Google also plans to bring the feature to the Chrome browser so voice input can populate any web field.

This release fits the recent pattern of breaking Gemini into task-specific models rather than waiting for a single flagship: Google is shipping audio as a distinct product line and embedding it across consumer apps and developer tools. For users, the takeaway is simple, faster, cleaner dictation is arriving across Google products, but where verbatim accuracy matters, treat the output as edited, not recorded. The rollout continues through the year, with Chrome and broader device support the next milestones to watch.

by Whatsnew Newsroom
whatsnew. APPS · WEB TOOLS · SECURITY · AI

Know what’s new.

The useful side of the internet. Covered properly.

Set as preferred →