The Spectrum Dispatch News

technology

Google Launches Gemini 3.8 Live Models for Real-Time Voice AI Agents

Two new models aim to improve conversational AI through near real-time reasoning, with extended thinking version for complex tasks.

Google Launches Gemini 3.8 Live Models for Real-Time Voice AI Agents

Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new AI models designed to enhance voice-based interactions and agent capabilities.

Google Launches Gemini 3.8 Live Models for Real-Time Voice AI Agents

Gemini 3.8 Live is optimized for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. The model processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while maintaining ongoing conversation flow.

Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, delivering increased intelligence and multi-step reasoning. According to Google, it reasons and speaks simultaneously, using verbal cues like “Let me check that…” to acknowledge prompts naturally while providing live progress narration for background tasks.

Both models are designed to enable production-ready voice agents for developers and enterprises. Google reports that Gemini 3.8 Live Extended Thinking captured the top overall spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, and leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also scored 97.7% on Big Bench Audio while maintaining competitive pricing compared to other frontier models.

Gemini 3.8 Live secured second place in the Speech Agent Arena and demonstrates strong cost-effectiveness for developers and enterprises. On ServiceNow’s EVA-Bench, both models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality.

All audio generated by these models includes SynthID watermarking, an imperceptible watermark embedded directly into audio output to help detect AI-generated content and prevent misinformation.

Google is rolling out both models across multiple platforms. Developers can access 3.8 Live through the Gemini API and Google AI Studio, while enterprises can use it in private preview through Gemini Enterprise. Both models are available in Search Live for general users. 3.8 Live Extended Thinking is available to developers via Gemini API and Google AI Studio, and to enterprises in private preview with coming availability through Gemini Enterprise and Google Workspace for business customers.

Developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents are integrating the Gemini Live API to help developers build voice-driven interfaces. Companies like Salesforce, Genspark, and Lumeris are partnering with Google to leverage these new models.

Key facts

  • Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index, ranking first overall
  • Gemini 3.8 Live ranked second in the Speech Agent Arena
  • The models support 97 languages with automatic mid-conversation detection and transitions
  • 3.8 Live Extended Thinking scored 97.7% on Big Bench Audio reasoning benchmark
  • All AI-generated audio includes SynthID watermarking for detection and transparency

Sources

← All posts