AI & Machine Learning

Google Launches Gemini 3.8 Live to Power Real-Time Voice Agents

Google introduced Gemini 3.8 Live and a reasoning-focused Extended Thinking variant, aimed at production-grade voice agents that can see, switch languages and run tools in the background without breaking conversation.

By Daniel Mercer Edited by Maria Konash Published: Updated:
Google Launches Gemini 3.8 Live to Power Real-Time Voice Agents
Google CEO Sundar Pichai. Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking for real-time voice agents this week. Image: Google

Key Notes

  • Google introduced Gemini 3.8 Live and a reasoning-focused Gemini 3.8 Live Extended Thinking model on September 15, 2026, aimed at production-grade voice agents that can process visual context, switch between 97 languages mid-conversation, and run tools in the background without breaking the conversation.
  • Both models are priced the same through the Live API at roughly $0.005 per minute of audio input and $0.018 per minute of audio output, and Gemini 3.8 Live Extended Thinking topped Artificial Analysis' Speech to Speech Quality Index while leading on several agentic voice benchmarks.
  • The models are rolling out to developers through the Gemini API and Google AI Studio, in private preview for enterprise customers, and to consumers through Search Live, Gemini Live and, for paid subscribers, inside Google Workspace.

Google has introduced two new audio models built for real-time voice interaction: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. According to the company’s announcement, the pair are its most capable live dialogue models to date, designed to make talking with Gemini feel closer to a natural conversation than a request-and-response exchange.

Gemini 3.8 Live is built for scale and cost efficiency, combining fluid dialogue with the ability to process visual context, such as a camera feed or a shared screen, in near real time. Gemini 3.8 Live Extended Thinking adds multi-step reasoning for more complex tasks, while narrating its progress aloud so a conversation does not stall while the model works in the background.

Built to Keep Talking While It Works

Both models can automatically detect and switch between 97 supported languages mid-conversation, and both can call tools and APIs in the background without interrupting the exchange, acknowledging a request and continuing to talk while a task completes. Extended Thinking is designed to use short verbal cues, such as “let me check that,” to bridge the gap while it reasons through a multi-step request, and to narrate progress as it works through longer background tasks.

On benchmarks Google cited in its announcement, Gemini 3.8 Live Extended Thinking took the top spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6, and led on agentic task completion with 68.6% on the tau-Voice benchmark and 35.1% on Sierra’s tau-Voice-banking test, while scoring 97.7% on Big Bench Audio. Gemini 3.8 Live, the lighter of the two models, placed second in the Speech Agent Arena while remaining considerably cheaper to run. Google also said the models perform well on ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents on complex workflows.

Pricing and Where It is Rolling Out

Google is pricing both models the same way through the Live API: roughly $0.005 per minute of audio input and $0.018 per minute of audio output, figures reported alongside the release, which the company says are based on underlying costs of $3 per million input tokens and $12 per million output tokens. That puts Extended Thinking’s added reasoning at no extra cost over the standard model, at least for now.

Gemini 3.8 Live is rolling out starting today for developers through the Gemini API and Google AI Studio, in private preview for enterprises in Gemini Enterprise, and for consumers in Search Live. Gemini 3.8 Live Extended Thinking is available to developers the same way, and to consumers in Gemini Live and, for Google AI Pro and Ultra subscribers, inside Google Workspace tools including Docs, Gmail and Keep. Enterprise access to the Extended Thinking model in Gemini Enterprise for Customer Experience and Workspace business accounts is described as coming soon.

Competing for the voice agent market

The launch puts Google in more direct competition with OpenAI’s full-duplex voice model and other real-time speech systems positioning themselves as the backbone of a coming wave of voice-first agents, from customer service bots to in-car assistants, part of a broader race toward enterprise-ready AI agents that OpenAI and others are also chasing. Google said Salesforce, Genspark and Lumeris are among the companies already building on the new models, citing their latency, fluidity and tool-calling as reasons for adopting them. The Salesforce mention comes weeks after Salesforce CEO Marc Benioff said he had shifted more of his own AI use toward Gemini, a preference that appears to extend to Salesforce’s voice agent infrastructure as well.

Google is also promoting a developer ecosystem around the Live API, with integration partners including Agora, LiveKit, Pipecat, LangChain, Fishjam, Vercel and Vision Agents, each offering tooling to handle the real-time media streaming layer so developers can focus on building the conversational experience itself. That mirrors the approach Google used with earlier Gemini Flash releases, including Gemini 3.6 Flash, of pairing a new model with a wider partner rollout rather than shipping it alone.

Watermarking and Safety Disclosure

Google says all audio generated by the new models is watermarked using SynthID, an imperceptible signal embedded directly in the audio output that is intended to help identify AI-generated content and limit its use in misinformation. The company published a model card alongside the release describing its approach to safety and responsible deployment for the new audio models.

The release lands as AI voice agents move from scripted phone trees toward systems capable of handling open-ended, multi-step requests entirely by voice, from booking travel to walking a user through troubleshooting a product live inside Search. Whether Gemini 3.8 Live’s combination of lower cost and background tool use is enough to pull developers away from competing voice APIs will likely become clearer as more of Google’s cited partners put the models into production over the coming months.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, Consumer Tech, Enterprise Tech, News