Google DeepMind Launches Gemini 3.8 Live Audio Models
The near-real-time model and its Extended Thinking variant handle voice, visual and text inputs, with different task emphasis and distribution paths.
Edited by Tyronne Panaino
Google DeepMind published its Gemini 3.8 Audio model card on September 15, introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both models process continuous audio, video and text and return spoken responses in near real time, targeting users, developers and enterprises building conversational experiences.
The practical split is task emphasis. Google positions Gemini 3.8 Live for high-volume, latency-sensitive voice interfaces, while Extended Thinking is aimed at complex reasoning and difficult multi-step work. The release expands Google's live-audio range without making the two variants interchangeable for every workload.
One multimodal input envelope, two task profiles
The model card says both variants are based on Gemini 3 Pro. They accept audio, images, video and text with a context window of up to 128K tokens, and produce audio and text with output of up to 64K tokens. Those are supported limits in Google's documentation, not evidence that every application will use the full window effectively.
Google's product page describes Live as optimized for cost-effective scale and near-real-time reasoning. It describes Extended Thinking as narrating progress while working through high-complexity tasks. Teams choosing between them should measure response time, interruption handling, tool behavior and answer quality in the actual workflow rather than infer superiority from the product labels.
Availability differs by variant
Google lists Gemini 3.8 Live across the Gemini API, Gemini app, Google AI Studio, Google Cloud or Vertex AI, and Google Search Live. Extended Thinking is listed for the Gemini app, API, AI Studio, Google Cloud or Vertex AI, and Google Workspace surfaces including Gmail, Docs and Keep. Product access may still depend on plan, region and the terms attached to each channel.
The model card keeps important limits visible
Google warns that both models can show general foundation-model limitations such as hallucinations. It also notes possible slowness or timeouts and gives January 2025 as the knowledge cutoff. These constraints are especially relevant to live interfaces, where a fluent spoken response can make an unsupported answer feel more authoritative than it is.
This coverage has not independently tested latency, reasoning quality, safety behavior or channel availability. A useful evaluation should use representative audio and visual conditions, measure recovery from interruptions and timeouts, and verify which model and settings actually served each session.
Status
Confirmed model release; medium internal confidence. Two official Google DeepMind pages support the model descriptions and distribution, but they come from the same interested actor and provide no independent performance validation here.
Sources
Update note: Last reviewed September 16, 2026. We will revise this post if availability, model limits or independent evaluation evidence changes.
Sources
- Google DeepMind — Gemini 3.8 Audio model card — official
- Google DeepMind — Gemini Audio product page — official
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.