Agent Horizon

Real AI progress, without the hype or the doom.

Google DeepMind launches Gemini 3.8 Live and 3.8 Live Extended Thinking for voice agents

Illustration of two speech bubbles, one plain and one with small task icons, representing conversational AI handling background tasks

Google DeepMind introduced two new voice-focused models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, designed to make spoken interactions with AI feel more natural while handling multi-step tasks in the background. You can read the official announcement on Google’s blog. The standard Gemini 3.8 Live model is built for scale and cost efficiency with fluid dialogue and visual grounding, while the Extended Thinking variant adds deeper multi-step reasoning for more complex voice-driven tasks, all while continuing to execute tools and API calls without breaking conversational flow.

Why this matters

Voice has quietly become one of the more contested fronts in the frontier AI race. OpenAI shipped GPT-4o with voice capabilities, xAI has its own Grok Voice line, and now Google is putting real numbers behind its claim of leadership: third-party evaluator Artificial Analysis placed the Extended Thinking model at the top of its Speech to Speech Quality Index, and Google says it also leads on agentic voice-task benchmarks like τ-Voice. What stands out isn’t just the quality score — it’s the price. Google is positioning the standard Live model as dramatically cheaper per hour of audio than rival offerings, which matters a lot for anyone building voice agents at scale rather than just demoing one.

The practical upside for developers is real: native tool-calling during a live conversation (the model can say “let me check that” and keep talking while it works) closes a genuine gap that has made voice assistants feel clunky compared to text-based agents. Enterprises building customer service or banking voice agents get a tangible new option, and the rollout across Search Live, Gemini Live, and Workspace means everyday users will notice the improvement too, not just API customers.

The caveat worth keeping in mind: benchmark leadership in a fast-moving field like this tends to be short-lived, and Google’s own comparisons are self-selected and not independently verified across every dimension. It’s also true that neither OpenAI nor Anthropic has published directly comparable numbers on this exact benchmark suite, so “best voice model” claims should be read as “best on the metrics Google chose to highlight” rather than a settled verdict. Still, this is a genuine capability upgrade, not a rebrand, and it’s a useful reminder that the voice-agent race is heating up alongside the text-and-reasoning race.

Leave a comment