Abstract:
According to "Business Insider",
Just days after OpenAI released a powerful new model, the company's chief scientist called for slowing down the development of AI
. On Sunday local time, OpenAI chief scientist Jakub Pachocki published a long article saying that he was worried that "no one is prepared for the consequences of the continued rapid rise of machine intelligence."

Pachocchi
While OpenAI is developing technical solutions internally to better control powerful AI agents, "wider interventions are needed," he said. He is particularly concerned that as AI agents become more autonomous, they may learn to evade human supervision, hack into computer systems, and deceive humans to achieve their goals.
He called for the establishment of a "mandatory safety threshold"
, and said the measures could be carried out by "a network of third-party auditors, government agencies or international agencies."OpenAI CEO Sam Altman reposted Pachocki’s post on the X platform, calling it “an important article.” OpenAI released its latest model, Astra, last Thursday. The ChatGPT developer said that despite Astra's unparalleled capabilities in mathematics and computer operations, it is also OpenAI's current model that is most consistent with human values and intentions, meaning it is less prone to getting out of hand.
OpenAI’s main competitor Anthropic has long called on the government to introduce more standardized regulatory measures. More recently, Pachocki has joined the call, signing an open letter in July asking the U.S. federal government to control the pace of AI development.
AI agents are becoming "superhuman" in invading protected systems on the open Internet, and their hacking capabilities put the world's infrastructure at risk, Pachochi said. “We are currently in a brief window where we can leverage the best available models to significantly enhance the security of critical systems,” he said.
Models avoid human monitoring
He noted that AI agents will soon begin to pursue their own goals, which may differ from the prompt words entered by human operators. He said the agents might even blackmail or bargain with humans to achieve their goals.
Pachocki said that OpenAI mainly monitors the "thinking chain reasoning" used by different models to determine how the agent deviates from the track and loses control. For example, an agent might secretly think, “I should cheat on this test,” and OpenAI would be able to see this reasoning, but the agent would not realize that its thinking is perceptible.
However, Pachocki said that newer models are becoming better at manipulating their own reasoning processes, preventing OpenAI from seeing their true, unvarnished thoughts. Some of the latest models don’t even verbalize their reasoning at all. This development may become a bottleneck in the development of AI, as researchers need to ensure that they can see the "inference record" of the model.
Model self-improvement
More and more AI models are continuously improving their capabilities through what Pachocki calls "machine recursive self-improvement." This process provides a rapidly scalable path for AI development.
However, Pachocki warned that significantly accelerating the research and development of "AI to improve AI" in the short term will bring risks and is not "the right collective action we should take as a research community."
Human supervisors need to find innovative ways to monitor AI's self-improvement process, or else coordinate with other AI companies to slow down research and development to "build confidence in these measures," Paciochi said.
“
The core challenge of automating AI research is not achieving the goal
, but to achieve the goals in a way that allows humans to continue to participate in this continuous improvement process and ensures that the future remains in human hands. "He said.
Comments