Google releases Gemini 3.8 Live Live Avatar digital human interaction function, giving AI real-time video face

📅 2026-09-25

Abstract:

Google announced that it has officially launched the "Live Avatar" (real-time digital human) function in the Gemini Enterprise enterprise platform. This technology is deeply integrated with Google's latest Gemini 3.8 Live large model architecture. Based on the original near-zero delay and high-smooth voice conversation, it introduces for the first time near-real-time generative video generation capabilities, enabling conversational artificial intelligence to have a dynamic visual image that can "listen, watch, and speak at the same time."

According to the official introduction, Gemini 3.8 Live Live Avatar is designed to create a more intuitive and immersive virtual service assistant for enterprise customers. By accurately synchronizing the audio stream with the underlying video generation model, the system can achieve high-precision lip-syncing, natural facial expression changes, and smooth cross-language dialogue conversion. Currently, the digital human system natively supports up to 97 languages, and there will be no video distortion or facial drift when switching across languages ​​in mid-conversation.

In addition to visual breakthroughs, the platform is also equipped with powerful asynchronous tool calling capabilities. While maintaining a natural and coherent conversation with users, digital humans can silently call various external APIs and enterprise data interfaces in the background to handle complex multi-step businesses such as hotel check-in registration and insurance claim document filling, completely saying goodbye to the "crash" or long silent wait phenomenon that occurred when AI processed background tasks in the past. At the same time, combined with the camera and screen sharing function of the terminal device, the digital human can also provide real-time visual understanding and interactive guidance of the physical objects or software interfaces displayed by the user.

In terms of security and compliance, in order to prevent the risk of artificial intelligence deepfakes and identity theft, Google has set up multiple security lines of defense for Live Avatar. By default, the system only opens to corporate users the digital human image library that has been preset by official review, while the permission to customize exclusive brand images or specific faces is placed under a strict whitelist review mechanism. In addition, all real-time voice and dynamic video streams generated by the system have been embedded with invisible and non-tamperable SynthID digital watermarks to ensure the transparency and traceability of AI-generated content.

Currently, Gemini 3.8 Live and Live Avatar functions have been officially opened to enterprise-level subscription customers through US and European data center service nodes, and developer API access support is simultaneously provided. Industry analysts pointed out that integrating generative video technology directly into the underlying speech model not only eliminates the high latency caused by traditional video rendering suites, but also marks that AI virtual customer service and intelligent interactive terminals are entering a new era of "real-person and real-time visual dialogue".

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet