Simon Willison published a browser-based UI for testing the new Gemini Live models
Google DeepMind Launches Gemini 3.8 Live Real-Time Voice Models
The release adds two production-grade real-time voice models to the developer market, with one claiming a top benchmark ranking for speech quality and agentic task completion. Audio watermarking via SynthID is built into all output to keep AI-generated audio detectable.
The full picture
Google DeepMind released two speech-to-speech models: Gemini 3.8 Live, positioned as a cost-effective conversational AI, and Gemini 3.8 Live Extended Thinking, designed for complex multi-step reasoning. Both models support real-time visual input, automatic language detection across 97 languages mid-conversation, and audio watermarking via SynthID. Gemini 3.8 Live Extended Thinking ranked first on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6 and leads agentic task completion benchmarks, according to Google DeepMind. Gemini 3.8 Live secured second place in the Speech Agent Arena. Extended Thinking reasons and speaks simultaneously, narrating its process with verbal cues while executing tools and API calls in the background. The models are available to developers via Google AI Studio and target enterprises and consumer products, with named partners including ServiceNow, Agora, LangChain, LiveKit, Pipecat, Vercel, Salesforce, Genspark, and Lumeris.
How it developed
Sources
Related
- Grew out ofOpenAI's Astra at the Critical cyber tier
- Grew out ofOpenAI's GPT-6 Astra model
Want this in your inbox?
I send one email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free