Google Gemini 3.5 Transcribe Cleans Up Speech, Removes Filler Words and Handles 85 Languages

Image: The Verge AI
Main Takeaway
Google’s Gemini 3.5 Transcribe turns spoken input into polished text, removing filler words, processing 85 languages and improving speed over the company’s Chirp 3 model.
Jump to Key PointsSummary
Google targets cleaner voice input
Google has introduced Gemini 3.5 Transcribe, a speech-to-text model that turns conversational speech into polished written text by removing filler words such as “um” and “uh,” handling corrections, and formatting output automatically. The model is part of a broader Gemini Audio update announced on August 26, alongside Gemini 3.5 Live and Gemini 3.5 Live Experimental.
The model already powers the Rambler dictation feature on Google’s Pixel 11, and Google is expanding it across more products and developer tools. The Verge described Transcribe as a new addition to the Gemini family, while Ars Technica characterized it as an expansion of technology already used in Gboard.
Faster transcription with fewer errors
Gemini 3.5 Transcribe improves both speed and accuracy compared with Google’s earlier Chirp 3 transcription engine. Google says the new model cuts the time from speech to final text by about 70 percent. Ars Technica reported a live-speech error rate of 5.5 percent for Transcribe, compared with 7.32 percent for Chirp 3.
That improvement is modest in percentage terms, but it addresses a familiar frustration with voice input: correcting small transcription errors can take longer than typing the original phrase. The model is also designed to handle background noise, interrupted speech and specialized terminology more effectively. These gains matter most when users dictate longer passages, technical language or conversational notes rather than short commands.
Editing speech while users talk
Transcribe treats dictation as an editable stream rather than a word-for-word audio record. Users can correct themselves by voice, while the model removes verbal stumbles and produces formatted text. A custom vocabulary feature lets people supply spellings and technical terms so the system can preserve specialized jargon instead of forcing manual cleanup.
The system supports more than 85 languages and can identify up to 3 speakers in pre-recorded audio. It also provides word-level timestamps, extending its use beyond casual dictation to meeting records, interviews and other audio workflows. The Verge emphasized the model’s multilingual performance and wording accuracy, while Ars Technica highlighted its ability to infer intended wording from corrections and unfinished thoughts.
The tradeoff behind polished text
Cleaner prose comes with a basic editorial risk: the output doesn't always represent exactly what a speaker said. Removing “ums” and correcting a verbal restart can make a note easier to read, but it also gives the model room to alter phrasing and interpret intent.
Ars Technica found the approach effective for short blocks of speech during testing with Rambler, while warning that changing wording may be inappropriate in some situations. That distinction matters for legal, journalistic, medical and research transcription, where pauses, corrections and exact wording can carry meaning. Users working with sensitive recordings will need to distinguish polished dictation from a verbatim transcript.
Gemini Audio expands across Google
Gemini 3.5 Transcribe is launching with Gemini 3.5 Live, which improves language recognition, mid-sentence interruptions and live visual processing. Gemini 3.5 Live Experimental adds real-time narration of its reasoning progress during more complex tasks, according to Google’s product description as relayed by The Verge.
The updates initially roll out in English to macOS users of the Gemini app, while Rambler is available on Android in selected countries and languages. Developers can access Transcribe in public preview through the Gemini API, AI Studio and Antigravity. Chrome support is planned next, extending the model from mobile dictation and voice chat into browser-based workflows.
What happens next for developers
The immediate opportunity for developers is access to a speech model that combines recognition, cleanup, speaker attribution, timestamps and custom terminology. That combination can reduce the engineering work required to turn raw audio into usable notes, transcripts or commands, especially for multilingual applications.
The bigger test is trust. A model that edits speech needs clear controls for switching between polished and verbatim output, preserving corrections and auditing changes. Google’s rollout places Transcribe inside consumer apps, Android dictation, Gemini voice features and the API, giving developers several ways to test those boundaries before Chrome support broadens its reach.
Key Points
Google Gemini 3.5 Transcribe removes filler words and corrects spoken errors during dictation.
Gemini 3.5 Transcribe supports more than 85 languages and identifies up to 3 speakers.
Google says Transcribe processes speech to final text about 70 percent faster than Chirp 3.
Gemini 3.5 Transcribe delivers a 5.5 percent live-speech error rate, down from 7.32 percent.
Google is expanding Transcribe from Pixel 11 Rambler dictation into Gemini apps and APIs.
Questions Answered
Google Gemini 3.5 Transcribe is an AI speech-to-text model that converts spoken language into polished text. It removes filler words, handles corrections, formats output and supports custom terminology.
Google Gemini 3.5 Transcribe produces final transcribed text about 70 percent faster than Chirp 3, according to Google. Its reported live-speech error rate is 5.5 percent, compared with 7.32 percent for Chirp 3.
Google Gemini 3.5 Transcribe automatically removes filler words such as “um” and “uh” from spoken input. It can also clean up corrections and format the resulting text.
Google Gemini 3.5 Transcribe supports more than 85 languages. It also accepts custom vocabulary so specialized jargon and unusual spellings are handled more accurately.
Google Gemini 3.5 Transcribe is rolling out through the Gemini macOS app, Android Rambler dictation in selected markets, and the Gemini API in public preview. Google also plans to bring the feature to Chrome.
Google Gemini 3.5 Transcribe isn't always suitable for verbatim records because it changes wording and removes verbal stumbles. Users handling legal, medical, journalistic or research audio should preserve the original recording and verify edited text.
Source Reliability
100% of sources are highly trusted · Avg reliability: 85
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems