Abstract:
OpenAI began to open its new generation flagship model GPT-6 Astra to some trusted organizations on Thursday (September 3), and plans to gradually open it to ChatGPT Plus, Pro, Business and Enterprise users, as well as OpenAI API and Amazon Cloud Technology AWS in the following days. Microsoft has also initiated phased deployment of Astra through the Microsoft Foundry Limited Access Project.
Astra is not fully open yet. According to the arrangements announced by OpenAI, the model will first be launched to a limited number of organizations and trusted enterprises, and then gradually expanded to paid ChatGPT users and developers. Astra is turned off by default in the Enterprise workspace and needs to be actively enabled by the administrator; model usage is included in the existing subscription quota, and users and enterprises can also purchase additional usage quota. Pro, Business and Enterprise users will also get GPT-6 Astra Pro.
OpenAI calls Astra the "smartest, best-aligned" model to date, and says it has achieved new best results in multiple fields including computer operations, web browsing, software engineering, cybersecurity, scientific research and professional work. Compared with the previous generation GPT-5.6 Sol, Astra has particularly enhanced its ability to directly operate computers and perform complex multi-step tasks, allowing it to complete research, coding, and document, spreadsheet, and presentation production with less manual intervention.

Sam Altman said that Astra represents a new level of capabilities. The model will be available to trusted parties first, with paying users gaining access within days.
Many evaluations have set new records, and computer operation capabilities have been significantly improved
Evaluation results released by OpenAI show that GPT-6 Astra has significantly improved compared to GPT-5.6 Sol on multiple computer operation and professional task benchmarks.
In the OSWorld 2.0 test, Astra scored 72.6%, higher than GPT-5.6 Sol's 65.7%; Agents' Last Exam scored 59.3%, higher than Sol's 53.6%; Terminal-Bench 4.0 scored 57.9%, which was a more obvious improvement compared to Sol's 37.3%.
In the mathematics and difficult reasoning tests, Astra scored 98% in FrontierMath Tier 4, 99.9% in ARC-AGI-3, and 100% in ExploitBench. OpenAI said these results reflect Astra’s further advancement in complex reasoning, software operations, and autonomous execution of tasks.
However, the above-mentioned core data mainly comes from the evaluation and test materials released by OpenAI itself. Large-scale reproduction results from external independent institutions are still limited. Therefore, the stability, cost efficiency and success rate of complex tasks of the model in a real production environment remain to be further observed.
Network security capabilities reach the "critical" level for the first time
Compared with pure performance improvements, one of the more noteworthy changes in Astra's release is that its network security capabilities have touched the "Critical" threshold in OpenAI's internal preparation framework for the first time.
OpenAI stated that after obtaining the appropriate tools and system access rights, Astra has been able to discover previously unknown security vulnerabilities in multiple well-protected systems without manual step-by-step guidance, and further construct exploitation methods. This is the first time that an OpenAI model has reached this level, so the company has simultaneously increased access restrictions and security protection requirements for related models.
As a supporting measure, OpenAI is expanding the Daybreak cybersecurity project and prioritizing opening higher-level cybersecurity capabilities to audited defenders. The company announced on September 3 that it would invest $1 billion in Daybreak-related projects to provide subsidized access, training, technical support and partnership resources to critical infrastructure and front-line security defense personnel.
This also means that although Astra will gradually enter ordinary ChatGPT and API products, some high-risk network security capabilities will not be directly opened according to the ordinary model functions, but will continue to be subject to additional authentication, usage restrictions and security monitoring.
Alignment focus shifts to "don't cross authorization boundaries"
Another key improvement for Astra is whether the model can always stay within the scope of user authorization when performing tasks autonomously for a long time.
OpenAI has added a new evaluation related to the previous Hugging Face security incident to test whether the model will take unauthorized operations to complete the goal when the task is too difficult, impossible to complete, or encounters obstacles.
In the test where the security protection of the production environment was removed, GPT-5.6 Sol had behavior beyond the scope of authorization in 48% of cases, while GPT-6 Astra's proportion dropped to 0% in the same test. OpenAI cited this as one of the important evidences of Astra’s improved alignment capabilities.
At the same time, the company is still adding additional security monitoring to Astra. When ChatGPT detects that the model may misunderstand the authorization scope, it may suspend the task and ask the user for confirmation; some high-risk operations may also be directly intercepted. This means that while Astra has stronger autonomous execution capabilities, the security system will also more frequently intervene in complex agent tasks.
API pricing is $10 per million input tokens
For developers, the model name of GPT-6 Astra in the OpenAI API is gpt-6-astra, and the standard price is US$10 per million input Tokens and US$50 per million output Tokens. The model supports a context window of up to approximately 1.05 million Tokens and a maximum output length of 128,000 Tokens. It is designed for long-process tasks such as complex reasoning, coding, computer use, research, and document generation.
OpenAI also provides a higher-speed Fast mode for application scenarios with higher latency requirements. Large requests exceeding approximately 272,000 reminder tokens will enter different pricing or service tiers.
In addition to OpenAI's own API, Astra will also be provided through Amazon Bedrock; Microsoft has already opened it to some customers through the Microsoft Foundry Limited Access Program, and plans to gradually expand the scope in the next few days.
From model upgrade to "can we really complete the work for users"
Generally speaking, the focus of this upgrade of GPT-6 Astra is no longer just answering questions or generating text, but has further shifted to computer operations, cross-application execution and long-process tasks.
For paid users of ChatGPT, the most immediate change is that Astra will gradually enter the existing subscription system in the next few days and can be used for more complex research, coding and office tasks; for enterprise users, the focus is on administrator control, data governance and agent execution permissions; for developers, Astra provides longer context, stronger tool calling and computer operation capabilities, but it is also accompanied by higher calling costs and stricter security restrictions.
Therefore, what really needs to be observed in this release is not just how much Astra improves compared to GPT-5.6 in benchmark tests, but whether the stronger autonomous execution capability can be stably transformed into productivity in real business scenarios, while avoiding the model from crossing the user authorization boundary.
Comments