Google Launches Gemini 3.8 Live and 3.8 Live Extended Thinking for Voice AI
Google DeepMind has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new models designed to bring near real-time reasoning and fluid voice interaction to developers, enterprises, and everyday users. These models set new benchmarks in voice agent quality, multi-step reasoning, and conversational continuity. They are available today through the Gemini API, Google AI Studio, Google Workspace, and the Gemini app.
Key points
- Gemini 3.8 Live Extended Thinking claims the number one overall position on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, and leads agentic task completion with 68.6% on the tau-Voice benchmark and 35.1% on Sierra's tau-Voice-banking benchmark.
- Gemini 3.8 Live secures second place in the Speech Agent Arena and is positioned as a highly cost-effective model built for scale, making it accessible for enterprises deploying voice agents at volume.
- Both models support 97 languages with automatic mid-conversation detection and switching, enabling truly multilingual voice experiences without manual configuration.
- Gemini 3.8 Live executes tools and API calls in the background while the conversation continues uninterrupted, allowing the model to acknowledge user requests and keep chatting while tasks complete asynchronously.
- Gemini 3.8 Live Extended Thinking reasons and speaks simultaneously, using verbal cues such as 'Let me check that' to signal processing and providing live progress narration during multi-step background tasks.
- All audio generated by these models is watermarked using SynthID, an imperceptible watermark woven into the audio output to ensure AI-generated content remains detectable and to help prevent misinformation.
Analysis
The launch of Gemini 3.8 Live and 3.8 Live Extended Thinking represents a meaningful shift in how AI voice agents handle complexity. Previous live voice models typically struggled to maintain conversational flow while performing background reasoning or tool calls. These new models decouple reasoning from speaking, allowing the system to narrate progress, acknowledge prompts, and continue the dialogue while complex multi-step tasks execute in the background. This architectural approach makes voice agents feel significantly more human and reduces the awkward silences that have historically degraded user experience in voice-first interfaces.
From a marketing and brand visibility standpoint, the integration of these models into Google Workspace products such as Docs Live, Gmail Live, and Keep Live, as well as Search Live, signals a clear intent to make voice a primary interaction layer for information retrieval and task execution. For SEO professionals, this matters because user queries processed through voice agents are likely to become more complex, contextual, and multi-turn, rather than the short keyword queries historically associated with voice search. Content strategies will need to account for conversational depth and the ability to satisfy multi-step informational needs.
The benchmark performance of Gemini 3.8 Live Extended Thinking is notable for enterprise marketers evaluating AI infrastructure for customer-facing voice applications. Leading the ServiceNow EVA-Bench Pareto Frontier by balancing accuracy with conversational quality directly addresses one of the core tensions in voice agent design. Enterprises can now consider voice agents as viable replacements or complements for structured customer interaction workflows, from onboarding to banking and troubleshooting, with quantified performance benchmarks to support business cases.
The ecosystem of developer platforms and enterprise partners cited in the announcement, including Agora, LiveKit, Pipecat, Vercel, Salesforce, Genspark, and Lumeris, indicates that Google is actively building a production-grade deployment network around these models. For agencies advising clients on conversational AI strategy, this means the tooling layer is maturing rapidly. Teams no longer need to manage complex real-time media streaming infrastructure themselves, as these platforms abstract that complexity and allow focus on user experience design and content architecture.
The introduction of SynthID watermarking across all audio outputs is a strategic transparency measure with implications for brand trust and regulatory compliance. As AI-generated voice content proliferates in customer service, marketing, and search experiences, the ability to verify whether audio is AI-generated will become increasingly important for compliance teams and for maintaining audience trust. Brands deploying these models in customer-facing contexts should factor this watermarking capability into their disclosure strategies and documentation.
What to do
- Audit your current voice search and voice agent strategy to assess whether your content and workflows are structured to handle multi-turn, multi-step conversational queries, since Gemini 3.8 Live models are designed to power exactly this type of interaction at scale in Search and Workspace.
- If you manage enterprise customer experience or contact center workflows, evaluate the private preview of Gemini 3.8 Live Extended Thinking in Gemini Enterprise for Customer Experience, using the published benchmarks (tau-Voice at 68.6% and EVA-Bench Pareto results) as baseline performance expectations for your internal business cases.
- Engage with the Gemini Live API through established developer platforms such as LiveKit, Pipecat, or LangChain to prototype voice agent experiences for your brand, taking advantage of the abstracted media streaming infrastructure to accelerate time to deployment.
- Develop a multilingual voice content strategy that accounts for the 97-language mid-conversation switching capability of Gemini 3.8 Live, particularly if your audience spans multiple language markets, since this feature removes a key barrier to deploying unified global voice agents.
- Incorporate SynthID disclosure into your AI content governance documentation when deploying Gemini-powered voice outputs in customer-facing contexts, ensuring your legal and compliance teams understand how the watermarking works and what it signals to regulators and users.
- Update your content briefs and editorial frameworks to include conversational depth and multi-step answer structures, anticipating that AI voice agents powered by extended thinking models will favor content sources capable of satisfying complex, layered user intents rather than simple keyword-matched responses.
As voice-based search and AI-powered assistants become more central to how users find information, these models directly influence how voice agents handle complex queries, potentially reshaping the dynamics of voice SEO and conversational search visibility. Marketers and brands building voice-driven experiences through partners like LangChain, LiveKit, or Vercel can now deploy more capable and production-ready voice interfaces that handle richer, more contextual user interactions.