The Information Machine
Following·New·first covered 15 Sep 2026·2 sources

Google DeepMind Launches Gemini 3.8 Live Real-Time Voice Models

The gist

The release adds two production-grade real-time voice models to the developer market, with one claiming a top benchmark ranking for speech quality and agentic task completion. Audio watermarking via SynthID is built into all output to keep AI-generated audio detectable.

The full picture

Google DeepMind released two speech-to-speech models: Gemini 3.8 Live, positioned as a cost-effective conversational AI, and Gemini 3.8 Live Extended Thinking, designed for complex multi-step reasoning. Both models support real-time visual input, automatic language detection across 97 languages mid-conversation, and audio watermarking via SynthID. Gemini 3.8 Live Extended Thinking ranked first on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6 and leads agentic task completion benchmarks, according to Google DeepMind. Gemini 3.8 Live secured second place in the Speech Agent Arena. Extended Thinking reasons and speaks simultaneously, narrating its process with verbal cues while executing tools and API calls in the background. The models are available to developers via Google AI Studio and target enterprises and consumer products, with named partners including ServiceNow, Agora, LangChain, LiveKit, Pipecat, Vercel, Salesforce, Genspark, and Lumeris.

How it developed
15 September 2026

Simon Willison published a browser-based UI for testing the new Gemini Live models

Sources
The daily email

Want this in your inbox?

I send one email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free