AI

Google ships Gemini 3.8 Live for real-time voice agents

2026-09-16 - ABikram Mondal

Google ships Gemini 3.8 Live for real-time voice agents

Google adds two new audio models to its Gemini lineup

Google DeepMind made Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking generally available through the Gemini API on September 15. The pair targets real-time voice applications and live dialogue without the delays common in earlier generations.

The standard Gemini 3.8 Live serves as the default choice for most low-latency voice agent experiences. It supports interleaved reasoning along with default asynchronous function calling and full session client content updates.

The Extended Thinking variant adds background reasoning during ongoing audio sessions. Developers can select it when a session needs deeper processing while the conversation continues.

Both models use an audio-to-audio approach inside the Live API. This setup keeps the interaction closer to natural speech patterns than text-mediated systems.

Release notes from the Gemini API changelog list the exact availability date and the two model identifiers: gemini-3.8-live and gemini-3.8-live-extended-thinking.

Technical details that matter for builders

These models focus on voice agent workflows rather than general text or image tasks. The design prioritizes low latency while allowing function calls to run in the background.

Full session content updates mean the model can maintain context across a conversation without resetting on every turn. Asynchronous calling lets the system handle external tools without interrupting the spoken flow.

Google positions the Extended Thinking version for cases that benefit from extra reasoning time. It runs higher background computation while the live audio channel stays open.

No context window size or token pricing appears in the public release notes for these specific variants. Earlier Flash models carried introductory rates, but those details do not extend to the new Live pair in the documented sources.

The models sit alongside prior Gemini releases such as the September 2 Flash variants, showing Google continues to iterate on the 3.8 family at a steady clip.

Who gains the most from this update

Developers building voice-first agents gain immediate access through the standard Live API. The low-latency profile suits customer support bots, personal assistants, and hands-free interfaces.

Teams already using Google tools can slot the new models into existing agent frameworks without major rewrites. The asynchronous function calling aligns with common patterns for tool use during calls.

Indian startups working on regional language voice products may find the timing useful if they target the Gemini ecosystem. Real-time performance often determines whether an app feels responsive in noisy environments.

ABikram Mondal builds automation for exactly this kind of problem at https://abikrammondal.com/services/automation.

Enterprises testing voice agents inside Google Workspace or Android Studio environments can evaluate the models through the same channels used for earlier Gemini releases.

What the models still cannot do

Public documentation does not report new benchmark scores for these audio variants. Claims about surpassing specific rivals rest on earlier Flash releases rather than the September 15 launch.

Extended Thinking adds background compute but does not turn the model into a full long-horizon agent that runs independently for hours. It remains tied to the live audio session.

Access stays within the Gemini API and Live API surface. No mention appears of direct integration into consumer apps like the Gemini mobile experience at launch.

Cybersecurity or specialized tuning variants from the same week do not carry over to these Live models. The focus stays on general voice dialogue.

Google has not confirmed whether the models support the full range of multimodal inputs available in text or image Gemini versions.

Context inside the current model race

Google released these audio models one day after OpenAI and Anthropic statements on coordinated safety work. The timing shows continued product shipping even as policy conversations continue.

The 3.8 Live pair follows Gemini 3.8 Flash and Cyber releases from September 2. That cadence keeps Google competitive on speed and agentic features without waiting for a full Pro update.

Other labs shipped agent-focused models earlier in the month. Google counters with a narrower but concrete improvement in live voice handling.

Developers who need immediate voice agent capability can test the new models now. Those waiting for broader capability jumps across all modalities may see less immediate impact.

The release keeps the emphasis on practical deployment rather than headline benchmark numbers.

Next steps for teams evaluating the models

Start with the default Gemini 3.8 Live identifier for standard low-latency tests. Switch to the Extended Thinking variant only when background reasoning adds measurable value in your workflow.

Measure end-to-end latency on your target hardware and network conditions. Voice agents succeed or fail on perceived responsiveness more than raw intelligence scores.

Integrate function calling early to confirm the asynchronous behavior matches your tool stack. Session content updates should preserve context across multi-turn spoken exchanges.

Track the Gemini API changelog for any follow-up pricing or availability changes. Introductory rates have applied to prior Flash models but remain unconfirmed here.

Teams that already run voice prototypes on earlier Gemini versions can upgrade with minimal code changes and compare results directly.

The short version. Google released two new Gemini 3.8 Live audio models on September 15 that improve real-time voice agent performance through low-latency dialogue and background reasoning options.

Sources

Reported from the sources above on 2026-09-16. Figures are as published at the time of writing. If something here has moved on, the linked source is the one to trust.

From the desk of ABikram Mondal

If you got here because you are actually thinking about putting models like this to work inside a real business, wired into the tools a team already uses, that is the work I do. I build for founders and small teams who want the thing to exist and work, not a deck about it.

Automation services  ·  everything I do  ·  talk to me

A new version is available.

Install on iPhone

Add ABikram to your Home Screen — it opens fullscreen like an app and stays updated automatically.

  1. Tap the Share icon in Safari's bottom bar.
  2. Scroll and tap Add to Home Screen.
  3. Tap Add — the trishul icon appears on your phone.

Get the ABikram App

Install the app for blogs, stories, Kundli and every tool — with live updates.