
Gemini 3.8 live introduces real-time visual avatars for enterprise applications
Google has announced Gemini 3.8 Live with Live Avatar, adding real-time video generation and speech to its conversational AI models. The feature supports natural turn-taking, background tool execution, and multilingual synchronization across 97 languages.
Published by Jin · 2 min read · 25 SEPT 2026
- Gemini 3.8 Live
- 97 languages
- US and EU endpoints
| Metric | Model | Specialization |
|---|---|---|
| Gemini 3.5 Transcribe | Specialized, modular audio capabilities | — |
| Gemini 3.5 Live Translate | Specialized, modular audio capabilities | — |
| Gemini 3.8 Flash-Lite TTS | Specialized, modular audio capabilities | — |
| Gemini 3.8 Flash TTS | Specialized, modular audio capabilities | — |

Google has introduced Gemini 3.8 Live with Live Avatar, a new feature that brings near real-time visual presence to conversational artificial intelligence. The system natively couples live dialogue capabilities with low-latency streaming video, creating a visual persona that listens, sees, and speaks with natural expressions and precise lip-syncing.
Multimodal conversation
Human conversation relies on multiple senses at once, including listening, looking, speaking, and observing facial expressions. The Live Avatar feature processes visual and audio inputs simultaneously to support more comprehensive and intuitive interactions. Enterprises can use the tool for customer service or interactive walkthroughs.

Asynchronous tools
The feature is backed by advanced reasoning capabilities that support asynchronous tool execution. This allows the model to trigger tool calls and fetch data in the background while continuing an active dialogue, ensuring an uninterrupted conversational flow during complex tasks such as hotel guest check-ins.
Global scale
Source — Original announcement ↗
Worth a read?
Comments · 0