ChatGPT mobile terminal adds voice intelligent agent function

📅 2026-09-24

Abstract:

OpenAI announced that the ChatGPT mobile application has begun to add voice-based intelligent agent functions. In the future, users will not only be able to chat with ChatGPT through voice, but also directly command it to perform more complex practical work using voice. For example, users can ask ChatGPT to create a document, draft an email, summarize a Slack message, or even perform a series of tasks that require multiple steps to complete.

The core of this update is to bring the Agentic Work capabilities that ChatGPT has gradually added to the desktop to mobile phones. For Plus and Pro users, the Work tab on the mobile phone can now be used directly to create documents, draft emails, and summarize Slack messages. It can also perform more complex workflows such as creating websites, making presentations, using cloud browsers, and handling financial-related tasks.

For Free and Go users, OpenAI focuses on the use of plug-ins and connected applications. This means that users at different subscription levels can make ChatGPT work with external tools and applications, but the specific range of functions available differs.

This update also changes the way ChatGPT voice mode itself behaves. OpenAI said that voice conversations can now provide richer text output, and users can see more complete text information at the same time when communicating with ChatGPT. Plus users can also more easily switch between text and voice modes without having to end the current conversation and start over again.

Continuous working across devices is also an important part of this update. Users can use their mobile phone to tell ChatGPT to start processing a task by voice when they are out, and then return to the computer to continue viewing and completing the task. In other words, the mobile phone is mainly responsible for issuing instructions and starting tasks at any time, while the desktop can continue to handle more complex workflows.

This change is related to OpenAI’s continued investment in voice interaction in recent years. In July this year, OpenAI launched a new GPT-Live speech model, focusing on improving natural dialogue, interruption processing, and continuous communication capabilities. Subsequently, OpenAI brought this voice capability to desktop applications, allowing users to control the intelligent agent in ChatGPT Work through voice, or complete software development tasks in Codex.

Previous desktop versions already allowed users to issue complex multi-step commands via voice. For example, developers can directly tell ChatGPT to create a new code thread, submit a Pull Request, and find the root cause of a program vulnerability without having to step through each operation. This mobile update basically follows this idea, but further extends voice control and intelligent agent capabilities to mobile phones.

OpenAI believes that voice is becoming an important way for people to interact with AI to complete complex tasks. Instead of having to stop and type text, users can directly tell the AI ​​what needs to be done through voice while walking, commuting, or doing other things, and then let the intelligent agent continue to execute it in the background.

This trend is also consistent with OpenAI’s previous positioning of voice technology. The company has stated that it hopes that as model capabilities continue to improve, voice will eventually become a major computer interaction method, and users can continue to command AI to complete increasingly complex and longer-lasting tasks through natural language.

OpenAI also demonstrated this development direction when it launched GPT-Live in July this year. The new speech model can not only handle user interruptions more naturally, but also call more advanced text models for search, reasoning and agentic tasks during continuous voice communication. In this way, the voice mode is no longer just an interface responsible for "speaking", but has gradually become the entrance to control the entire AI working system.

This mobile update also means that OpenAI is gradually changing the positioning of the ChatGPT mobile application. In the past, the mobile version of ChatGPT was mainly responsible for tasks such as question and answer, search, writing and voice chat. Now users can start a complete workflow directly from their mobile phone and let AI complete the actual operation in connected applications and tools.

At the same time, OpenAI's approach also competes with the direction its competitors are taking. Anthropic has also recently been strengthening Claude's mobile and desktop collaboration capabilities, and further integrating the chat interfaces of Cowork and Claude, allowing users to initiate tasks from mobile devices and then transfer them to the desktop to continue processing. However, the two companies still have different choices in product structure. OpenAI still separates the ordinary chat interface from work spaces such as Work.

Judging from OpenAI’s current product roadmap, voice, intelligent agents, and cross-device work are gradually integrating. In the future, users will not necessarily need to open a specific application, find specific buttons, and then operate the software step by step. Instead, they can directly state their goals, and ChatGPT will be responsible for understanding the task, calling tools, and executing multiple steps.

This mobile update does not just add a voice chat function, but also allows ChatGPT on the mobile phone to begin to have the ability to "voice command work". As OpenAI continues to improve products such as GPT-Live and Work, ChatGPT's voice mode is gradually transforming from a simple conversation tool to a portal connecting users with AI agents, external applications, and actual digital workflows.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet