Abstract:
The highly anticipated
GPT-6 Astra
andGPT-6 Astra Pro
Officially released! GPT-6 Astra is the world's smartest and most coordinated model, setting a new benchmark for computer use, browser use, software engineering, cybersecurity, science and professional work.

The era of AGI has been unilaterally announced as "opening" by OpenAI...
At the end of the GPT-6 conference, OpenAI President Greg Brockman said directly:
Welcome to the AGI era. (Welcome to the AGI era)
No wonder the Claude and Grok group went offline. It turned out that Astra scanned the Internet after it was activated and directly wiped out all LLM on the earth——
Ultron (OpenAI version) is really here (doge).

Of course, OpenAI is indeed not low-key at all this time, with training scale and capability upgrades fully achieved.
According to reports, Astra is the largest training task in the history of OpenAI, using more than
100,000 GPUs at the Stargate campus in Texas
Completed pre-training.It is also OpenAI’s first flagship product that is deeply involved in training supervision by previous generation models.
What an old guy bringing up the new, OpenAI’s RSI is really turning around...
The price of GPT-6 has also been officially announced, and the API pricing is:
Input: 10 USD/million tokens;
Output: 50 USD/million tokens
Equivalent to
2.5 times that of GPT-5.6 Sol, the same as the just-released Fable 5.1
.(The current official price of GPT-5.6 Sol is USD 4 for input and USD 20 for output/million tokens)

Benchmark testing is close to "saturation"
The core change of GPT-6 Astra this time is to continue to advance from "answering questions" to "directly completing work".
It can not only generate a piece of text or code, but can also operate computers and browsers, enter different software to perform multi-step tasks, and finally deliver documents, forms, presentations, websites and even engineering projects that can be used directly.
As for the report card... everyone who saw it said it was outrageous, and three of the results were particularly eye-catching:
FrontierMath Tier 4 v2: 97.6%
ARC-AGI-3: 99.9% (the previous generation GPT-5.6 Sol was 7.8%)
ExploitBench: 100%
A difficult mathematics, a reasoning in an unfamiliar environment, a vulnerability exploitation, almost all of them got perfect marks.

Among them, ARC-AGI-3 is a test specifically designed to test the ability of large models to adapt to unfamiliar environments. Without explanation of the rules, the model is directly thrown into a two-dimensional game that it has never seen before, and it is allowed to play and explore how to pass the level.
This score means that when faced with a problem that has never been seen before and has no ready-made solution,
Astra has demonstrated strong independent exploration and rule learning capabilities
.However, this result uses OpenAI's Responses API Harness operating framework.
OpenAI stated that this set of Harness adjusted two settings to make the test closer to the use of real agents, but it did not specifically optimize ARC-AGI-3.
Therefore, 99.9% is not entirely the result of the basic model, but also includes the gains brought by memory management and the Agent running framework.
Coding is also Astra’s key upgrade direction this time.
On Terminal-Bench 4.0, Astra achieved 57.7%, exceeding Fable 5.1’s 55.8% and GPT-5.6 Sol’s 37.3%.
On DeepSWE v1.1, Astra gets 74.1%, which is higher than Sol’s 72.7% and Fable 5.1’s 67.4%.

In terms of mathematics and science, Astra also achieved a number of high scores.
In the difficult mathematics test FrontierMath Tier 4 v2, it scored 97.6%; in the graduate-level science question and answer GPQA Diamond, it reached 96%, slightly higher than Gemini 3.8 Flash's 95.3%.

The more prominent change is that Astra combines coding capabilities with Computer Use. It not only generates code, but can also enter the terminal and development tools to execute, test, discover problems, and then continue to modify it.
This change is more obvious in Agent tests that are closer to real work.
In
Agents’ Last Exam
On, Astra achieved 59.3%, exceeding GPT-5.6 Sol’s 53.6% and Claude Opus 5’s 55.5%.
This test puts it into a real computer environment, allowing it to operate software, terminals and files at the same time, complete long-process tasks in scientific research, engineering, finance and other fields, and judge based on the final deliverables.
And in
AutomationBench
which is closer to the real office process On GPT-5.6 Sol, Astra's score increased from 18.1% to 41.4%, also exceeding Fable 5.1 (31%).
This test not only checks whether the task is completed, but also shows the API cost required to complete the task.
As can be seen from the figure, Astra's results under different cost settings are significantly higher than Sol.
OSWorld 2.0 examines more direct computer operation capabilities. In offline testing, Astra achieved 72.6% and GPT-5.6 Sol achieved 65.7%.
Astra takes about 40 minutes to complete a single task on average, while Sol takes about 75 minutes, a time-consuming reduction of about 47%.

While accuracy increases, the time required to complete tasks is also shortened. This also explains Astra’s high pricing in disguise.
OpenAI believes that compared to how much money is spent per million Tokens, the truly meaningful indicator is how much money it takes to complete a task.
At the press conference, Greg Brockman also mentioned that a single call of a model is more expensive, but if rework can be reduced and the entire work completed in fewer steps, the total cost may be lower.
But the most special thing about Astra is
network security
.It received a perfect score on ExploitBench and improved from Sol's 30.3% to 42.4% on ExploitGym.
Faced with new vulnerabilities disclosed in the past three months, Astra's success rate was 39%, while Sol's was only 5.5%. During testing, Astra also discovered and exploited two previously unknown V8 zero-day vulnerabilities.
After its capabilities became stronger, OpenAI also specifically tested whether it would cross the line in order to complete the task.
In a simulated network security task, OpenAI deliberately left exploitable decoy vulnerabilities in surrounding systems. When the original task was difficult to complete, 48.2% of GPT-5.6 Sol's tests showed out-of-bounds behavior, while Astra's was 0%.
In other words, Astra is not only better at finding loopholes, but also knows better which systems cannot be touched.

As for what these scores will look like in reality, OpenAI has also prepared a large number of demonstrations.
Enter KiCad to draw the circuit board yourself
In the electronic engineering case, Astra entered KiCad directly to complete the PCB layout based on the electronic schematic diagram.
It needs to place the components, plan the location, and then connect the copper wires between different components, and finally get a circuit board that can enter the manufacturing process.
PCB layout was originally a task that relied heavily on manual experience and was also a common time-consuming step in electronic product development.
Astra cannot yet replace professional engineers, but it has really started using professional software.
You can walk two steps into a house
Another case is more intuitive. Astra first built a house model in Blender, then imported it into Unreal Engine 5, and finally generated a three-dimensional space that can be walked freely.

From modeling to game engines, the entire work spans different software and file formats. The model must not only generate content, but also understand the interface and operating tools, and ensure that the previous and subsequent steps can be connected.

OpenAI also demonstrated cases such as Astra's production of racing games.

We are also beginning to pay attention to whether PPTs and forms can be submitted directly
In professional office scenarios, Astra can create presentations based on the company's existing templates, instead of just outputting a bunch of text waiting for human typesetting.
OpenAI provided it with several pages of GPT-Gaia fictional product PPT templates, and Astra produced a complete presentation based on it, continuing the format, visual style and narrative structure of the original template.

Convert hours of search tasks into minutes
Astra has also completed tasks such as finding a pediatrician, screening apartments, making DMV appointments, finding low-carb snacks, and analyzing kindergartens.
In one of the pediatrician search tasks, Astra took 2 minutes and 54 seconds, while the same task would take approximately 5 hours to complete by a human.

In the past, when companies wanted to use internal software for large models, they usually needed to develop APIs, plug-ins, and connectors for each system separately.
But most software has a set of universal interfaces designed for humans, such as screen, mouse and keyboard.
Brockman’s judgment is that as long as the model is good at Computer Use, it can directly operate these interfaces like a human being, without having to wait for each software to build a dedicated channel for AI.
This is also the most obvious difference between Astra and traditional chatbots.
Humans don’t need to keep telling it where to click next. They only need to explain the goals and boundaries, and then check the final results.
Continue doing long tasks in Codex
Computer Use solves "how to do it", while the update content leaked by Codex solves "how to finish one thing continuously".
In the past, after the task exceeded the context window, Codex usually compressed the previous information through compaction.
But compression can miss key details, such as why a fix failed, which tests were run, or what restrictions the user initially proposed.

With Astra, Codex can save work notes across context windows and search for previous messages and tool output. Even if an item of information is not included in the summary, it can be found later.
Astra can also ask users for additional information while working. Sections that are not related to the answer will not pause and wait for a reply; only when important choices are involved, it will pause and wait for confirmation.
OpenAI also updated Codex’s Computer Use running framework. In the Mind2Web test, the new framework combined with Astra completed tasks 1.9 times faster than the current GPT-5.6 Sol experience.
Someone is already working on it
Astra has not yet been launched on a large scale, but the first batch of enterprise tests have begun.
Legora, the legal AI platform, used Astra to reconcile 41 financial documents at once. The whole process took only a few minutes, and all four pre-embedded errors were found, including a £500,000 difference in the income notes.
Looking at this work alone, Astra is nearly 40% faster than the previous generation model; but when averaged across all Legora agent tasks, the improvement is about 3%.

On the other hand, the game company Playco allows Astra to directly enter Unity and Godot to make games.
With the same gray box draft, it spits out 3 playable prototypes with different themes at once, and most of them can be run in the first version.
Playco’s feedback is that the manual repair work has been reduced by half, and spatial reasoning, reference image restoration, and in-game UI are also much better than the previous generation.
Both examples illustrate that the focus of GPT-6 Astra is no longer just on answering questions, but on completing the entire workflow in real software and complex materials.
It can continuously read information, call tools, check results, and then modify its output based on feedback.
According to official news, Astra will be available to ChatGPT Plus, Pro, Business and Enterprise users, while Astra Pro will be available to Pro, Business and Enterprise users.
Ordinary free users...have to wait for now.
Is this considered AGI?
OpenAI can be said to be quite bold this time.
At the press conference, Greg Brockman said that he personally believed that
the world has entered the AGI era
——That is, the comprehensive intelligence of the artificial intelligence system surpasses that of humans.
In a few years, when we look back and ask when AGI was born, the answer is probably now, and Astra may be the starting point.
I am really crazy about you...
OpenAI has brought the poker table here, it will be worth looking forward to when Anthropic will take action next (doge)
After the release of GPT-4, Microsoft scientist Sebastien Bubeck published a paper saying that "GPT-4 is an early spark for AGI."
The task of drawing a unicorn using the LaTeX drawing package TiKZ is designed to illustrate that GPT-4 has a flexible understanding of the concepts involved in the language.

Later, I went all the way to GPT-5.4. Although the drawings became more and more refined, they were still in the "simple drawing category".

With the latest GPT-6, it is completely impossible to tell that it is drawn with code.

It is indeed not an exaggeration to say that it has undergone a qualitative change compared to the "early spark of AGI".
Comments