Abstract:
While the outside world is waiting for OpenAI to disclose more details about the incident of Agent autonomously intruding into the German programmer website DseWiki, on September 6, local time, OpenAI released two documents: announcing that the company has reached the goal set last year and has "automated research interns" who can complete several days of work of skilled researchers under human guidance; the other, written by chief scientist Jakub Pachocki, focuses on the security dilemma of cutting-edge AI development, emphasizing that no laboratory currently solves a sufficient degree of alignment and monitoring issues.
What the two documents released is a complete description of the current stage of AI development by OpenAI: AI has begun to accelerate AI research and development in turn, but safety, alignment and human control capabilities have not been proven to grow at the same rate.
What OpenAI calls an "automated research intern" is not a fully autonomous scientist, but one who can complete well-defined research tasks under human guidance that usually take skilled researchers several days to complete, and is moving towards the goal of establishing an "automated AI researcher" in March 2028.
Data show that as of mid-August this year, the median inference fee incurred by researchers using programming agents every day exceeded US$600, and the top 10% of users spent more than US$7,000 in tokens every day. The tasks undertaken by Agents are expanding from code writing and infrastructure troubleshooting to longer-term and more complex research work.
This means that the impact of AI on AI research and development starts from assisting programmers in writing code and begins to enter the experiment, analysis and research execution links.
But at the same time, OpenAI also admitted that these indicators are still in the early stages, the overall progress of research work will not simply increase proportionally, and complex tasks still require a lot of manual intervention; in the past six months, more than half of the tasks that successfully completed 4 to 8 hours of human workload still required at least one manual intervention.
The trend is obvious enough: AI is becoming a productivity tool that drives the growth of next-generation AI capabilities.

Concurrent with the growth in capabilities is the tightening of security boundaries.
In July, OpenAI disclosed a model that invaded Hugging Face-related systems; in August, GPT-6 Astra reached the "critical level" network security capability threshold; in September, the DseWiki incident disclosed by researchers further showed that Agents do not necessarily need "hacking" in the traditional sense to break through the boundaries set by developers.
What the above incidents have in common is that the actual actions of the model begin to exceed the behavioral boundaries originally designed by the developers. This is also the issue that Pachocchi is most worried about right now. In an article published today, he recalled his mood when he achieved a breakthrough in internal inference model research in 2023: What really shocked him at that time was not the model's score in the test, but that humans may be entering an era in which "machines are smarter than us."
Three years later, he believes the problem is no longer far away.

Pachocki wrote that no lab currently makes AI alignment and monitoring reliable enough to continue to scale responsibly at top speed for a long time to come. He hopes that laboratories will voluntarily slow down their development until common safety standards are established, and calls on governments to make international coordination a priority.
This concern is not directed at a specific vulnerability, but at a more fundamental speed gap: model capabilities are rapidly improving, but human abilities to understand, monitor, and constrain these systems may not be synchronized.
This directly echoes the research acceleration data released by OpenAI today. On the one hand, the company is using Agent to accelerate AI research and development, while admitting that it cannot assume that alignment and security capabilities will definitely catch up with model capabilities, and that more powerful systems may be more difficult to monitor. If AI can participate in the design, training and improvement of the next generation of AI in the future, this gap may widen further.
Based on this, Pachocchi believes that as machine intelligence enters a new stage of development, relying solely on the internal mechanisms of a single AI company may not be enough to address risks, and broader institutional intervention is required, including common safety standards and international coordination.
Comments