Google unveiled Gemini 3.5 Transcribe as the latest addition to its Gemini Audio family, marking a significant step forward in AI‑driven transcription. The model—described by the company as a major upgrade from its earlier Chirp 3 system—offers multilingual support for more than 85 languages and can intelligently filter out filler sounds such as “um” and “uh.” It also formats text on the fly, allowing users to edit documents using only their voice.
One of the standout features is the ability to ingest a customized vocabulary. Organizations that rely on industry‑specific terminology can upload their own word lists, ensuring the model retains proper spelling and avoids unwanted edits. In practice, this means legal firms, medical researchers, and technical teams can dictate without having to constantly correct jargon after the fact.
Beyond simple dictation, Gemini 3.5 Transcribe provides word‑level timestamps and can attribute speech to up to three distinct speakers in a pre‑recorded clip. That capability opens the door for faster podcast editing, meeting minutes generation, and multi‑speaker interview transcription without manual speaker labeling.
The rollout begins today for English‑language users of the macOS Gemini app and the Rambler dictation feature on Android in selected markets. Developers interested in early access can explore the model through a public preview on the Gemini API, available via AI Studio and Antigravity. Google also hinted that Chrome integration is on the horizon, though no date was supplied.
Google originally hinted at a broader Gemini 3.5 suite, including Live and Live Experimental models that would enhance real‑time translation and visual processing. After publication, the company clarified that only the Transcribe model is being launched at this time, postponing the other variants.
Industry observers note that the new transcription tool could tighten Google’s grip on the enterprise speech‑recognition market, where competitors like Microsoft and Amazon have long offered niche solutions. By combining robust multilingual performance with speaker attribution and custom vocabularies, Gemini 3.5 Transcribe positions itself as a versatile option for both consumer and professional users.
Google’s announcement arrives as businesses increasingly rely on voice input to streamline workflows and cut down on manual typing. With the ability to produce cleaner, more accurate transcripts out of the box, the Gemini 3.5 Transcribe model may reduce the time spent on post‑processing and improve overall productivity.
While the company has not disclosed pricing, the feature is currently available at no extra cost to existing Gemini app users. Future updates, including the promised Chrome support, could broaden its reach to a wider audience of web‑based users.
This article was written with the assistance of AI.
News Factory APP - agentic news to boost your SEO & AEO.