Back to articles
Multimodal

Gemini 3.8 Live Adds Real-Time Avatars for Enterprise Conversations

3 min read

Introduction

Google DeepMind has introduced Gemini 3.8 Live with Live Avatar, an extension of its live conversational model that adds a near-real-time visual presence. Instead of limiting an assistant to audio responses, the system combines live dialogue with low-latency video generation, allowing an enterprise agent to listen, see, speak, and appear as a responsive digital character.

Key points

  • Multimodal interaction: Live Avatar processes audio and visual inputs together and responds with synchronized speech, video, facial expressions, and lip movements. Google presents this as a way to make customer-facing conversations feel less like interactions with a conventional voice bot.
  • Background tool execution: The agent can call tools asynchronously while maintaining an active conversation. In the company’s example, an avatar handles a hotel check-in workflow while retrieving information in the background, reducing silent pauses during task execution.
  • Multilingual speech-to-speech: The system can move between 97 languages during a conversation. Lip synchronization and facial expressions are adjusted as the language changes, with Google aiming to preserve video quality and avoid visible drift.
  • Brand-specific characters: Organizations can choose from a preset avatar library or create a customized character from a high-quality reference image. The generated avatar is intended to preserve the reference likeness, brand styling, or character identity. Custom avatar creation is currently available only through enterprise allowlisting.
  • Built-in provenance signals: Google says AI-generated audio and video output contains an imperceptible SynthID watermark. The watermark is designed to help detect synthetic media and reduce misinformation or misattribution.

Why it matters

The more important shift is not simply that enterprise agents can now have faces. Live Avatar treats visual presence as part of the conversational loop. In customer support, guided walkthroughs, training, and product demonstrations, expressions and synchronized mouth movements may make explanations easier to follow and make longer exchanges feel more immediate.

Asynchronous tool use is particularly relevant for business deployment. Enterprise assistants often need to query internal systems or trigger workflows before they can provide a complete answer. Keeping the dialogue active while those operations run could improve the waiting experience, but it also raises practical questions about permissions, error handling, and how clearly the system communicates uncertainty.

A visual persona also introduces risks. Users may overestimate the agent’s understanding, mistake a synthetic character for a real person, or associate a generated avatar too closely with a real identity. SynthID can support provenance, but it is not a substitute for clear disclosure, access controls, and deployment policies.

Gemini 3.8 Live with Live Avatar is now available in Gemini Enterprise, with developers directed to the API documentation for integration. The launch points toward a broader enterprise AI pattern: agents that remain visibly present while reasoning, retrieving information, and completing tasks in real time.

Source: Google DeepMind

Comments

Checking sign-in status...

Loading comments...

Related articles