OpenAI today officially released a new ChatGPT voice experience based on GPT‑Live, which is an important architectural upgrade of its voice technology route. Compared with the previous Advanced Voice Mode, which will be launched to more users in 2024, the new version of the voice mode no longer relies on the traditional voice system, but turns to a real-time audio model specially built for intelligent voice agents.

Advanced Voice Mode was once regarded as an advancement in the voice assistant experience with its more natural dialogue, improved response speed, and support for interruptions. However, the subsequent generation of real-time speech models launched by OpenAI quickly exposed the limitations of this old technology.

As early as May this year, OpenAI announced its most advanced real-time speech model to date, GPT‑Realtime‑2, bringing GPT‑5 level reasoning capabilities to real-time audio interaction scenarios. Compared with past speech systems that only emphasized "speak fast and answer well", GPT‑Realtime‑2 can better handle complex requests, track long conversation contexts more stably, and complete multi-step tasks while maintaining natural speaking speed and tone. On this basis, OpenAI has now released a new generation of ChatGPT voice experience based on the GPT‑Live architecture, further increasing the interaction limit of voice assistants.

The core of this upgrade lies in GPT‑Live, a full-duplex voice model architecture, which allows ChatGPT to "speak, listen and think at the same time", unlike traditional voice assistants that require you to finish speaking and then respond. In actual interactions, GPT‑Live can naturally insert slight tone feedback such as “okay” and “hmm” when the user pauses or slows down his speech, instead of roughly interrupting the complete sentence. When the user needs to think for a moment, the model can also consciously choose to remain quiet to avoid being disturbed.

In terms of interruption control, users can interrupt or change the direction of the question at any time, and the model will instantly adjust the response logic to smoothly connect to the new conversation target. OpenAI says the new voice experience can also handle questions that require networked search, deeper reasoning, or complex task planning because its latest "cutting-edge models" are called behind to generate answers. Currently, GPT‑Live uses GPT‑5.5 by default in the background. This means that what users gain in a voice scenario is not only a dialogue rhythm that is closer to humans, but also a "rethinking" ability that is similar to text mode.

OpenAI also simultaneously announced that it will launch two versions of the GPT‑Live model to global ChatGPT users: GPT‑Live‑1 and GPT‑Live‑1 mini. Among them, users of paid tiers such as ChatGPT Go, Plus and Pro will use the more comprehensive GPT‑Live‑1 by default, while free users will get a voice experience powered by GPT‑Live‑1 mini. For developers and enterprise customers, these two models will also be open to access through APIs, making it easier for them to build a new generation of voice agents and vertical applications.

From a rhythm perspective, OpenAI is accelerating the upgrade of the traditional "walkie-talkie" voice assistant form to a full-duplex intelligent agent mode that is closer to real-person conversations. With the implementation of models such as GPT‑Realtime‑2 and GPT‑Live, the ChatGPT voice mode is transforming from a simple voice question and answer tool to an "intelligent conversation partner" that can perform complex tasks and maintain long-term interactions in a voice environment.