Google is introducing two new models that bring advancements in near real-time reasoning to more effectively enable voice agents and make conversing with AI feel more intuitive and intelligent. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, with increased intelligence and multi-step reasoning.
For developers and enterprises, these models deliver the building blocks for reliable, production-ready voice agents. They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping users tackle complex tasks using just their voice. Experience more fluid, intelligent conversations with Gemini 3.8 Live Extended Thinking, which provides enterprise-grade task completion and intelligence.
Gemini 3.8 Live Extended Thinking has captured the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index, and leads in agentic task completion. It also provides strong reasoning capabilities, scoring high on Big Bench Audio, while maintaining a highly competitive price point compared to other frontier models. Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena, and remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale.
On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, Google's models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality. Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It automatically detects and transitions between 97 supported languages mid-conversation, and executes tools and API calls in the background while continuing the conversation.
For tasks that require deeper reasoning, 3.8 Live Extended Thinking reasons and speaks simultaneously, delivering increased intelligence for complex workflows while maintaining an uninterrupted conversational flow. Across Google Workspace and Search, the Live models deliver more intuitive, collaborative experiences — especially when tackling complex tasks. Get step-by-step, real-time troubleshooting help powered by Gemini 3.8 Live — right inside Search Live.
By using the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience. Google is also partnering with companies like Salesforce, Genspark, and Lumeris, who are excited about 3.8 Live and 3.8 Live Extended Thinking, highlighting its impressive latency, fluidity, and tool-calling capabilities.
All audio generated by Google's AI products is watermarked with SynthID, an imperceptible watermark woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on Google's approach to safety and responsibility, review the model card.
Este artigo foi escrito com a assistência de IA.
News Factory APP - notícias agênticas para impulsionar seu SEO e AEO.
