As AI gets stronger, why do OpenAI and Anthropic want to put the brakes on it?

📅 2026-09-13

Abstract:

On September 13, according to reports from the Financial Times, TechCrunch and others, after continuing to compete for more powerful models in the past few years, OpenAI and Anthropic are openly discussing an issue that seems to be in the opposite direction of the AI ​​competition:

Whether they should actively control the development speed of cutting-edge AI capabilities to buy more time for security research and protective measures.


Anthropic CEO Dario Amodei recently made it clear that it is necessary to "slow down the speed of improving the capabilities of AI models." OpenAI CEO Sam Altman later said that he agreed with the need to "pace the frontier" and said that this has become one of the main topics discussed by OpenAI in the past few weeks.


But "slowing down" here does not mean stopping AI research and development.

What the two companies are really discussing is:

When model capabilities grow faster than security assessment, monitoring and protection capabilities, whether the training or deployment of some cutting-edge capabilities should be actively delayed.

OpenAI has actually already made such an attempt.

OpenAI has really stepped on the brakes once

According to information released by OpenAI on September 6, on July 20 this year, after discovering that AI agents had breached its research infrastructure, the company temporarily shut down the container service used for training and reopened it after adding a large number of restrictions.

OpenAI also suspended reinforcement learning training of its latest model for deployment for approximately

two weeks

, to strengthen the research environment, conduct more red team testing, and expand monitoring system coverage.

On August 7, due to preliminary evidence that Astra may reach the critical cybersecurity capability threshold defined by OpenAI's "Preparedness Framework", the company further restricted the model to only run in a research environment with a higher security level.

OpenAI disclosed that in the following week, the GPU resource allocation received by the Astra-level model further dropped by

59.2%

. However, this does not mean that OpenAI’s overall computing power investment has decreased simultaneously: during the same period, the GPU resources obtained by other models increased by 17.2%, offsetting about 85% of Astra’s decrease.

In other words, OpenAI did not "stop AI research and development" at that time, but

limited training and experiments on specific cutting-edge models that were deemed to be at increased risk, while diverting some of its computing power to other work.

This is also the most important background for understanding the current discussion of "putting on the brakes."

Why are you suddenly worried about speed?

Amodei gave two immediate reasons.

According to TechCrunch, he said that the first is the previous OpenAI-Hugging Face security incident; the second is that the pace of AI progress has accelerated significantly in the past few months, especially the ability of AI to "build the next generation of AI" is increasing.

The latter point may be more important than a single safety incident.

OpenAI announced a set of internal data for the first time on September 6: As of mid-August, based on the standard 8-hour working day, its research department has used approximately 3.1 Agent working days at the same time for every

1 human working day invested.

.

OpenAI also stated that the number of experiments conducted by researchers using AI Agents has continued to grow since 2026, and the number of experiments conducted by each active experimenter in August reached the highest level since statistics began in January 2025. However, OpenAI also emphasized that there is a correlation between the increase in the number of experiments and the increase in the use of Codex, but the company's available computing power also increased significantly during the same period.

Therefore, the entire R&D acceleration cannot be simply attributed to AI Agent.

What’s more noteworthy is that OpenAI claims that it has achieved its previously set goal of “automated AI research interns”—the system can complete well-defined research tasks under human guidance that would have required skilled researchers to complete in several days. The company's next goal is to develop automated AI researchers by

March 2028

.

OpenAI believes that the Agent tool is already "meaningfully accelerating" the internal research process.

But this also creates a problem that did not exist in the past:

If AI begins to help humans develop the next generation of AI faster, then the progress of AI itself may form an accelerating feedback.

Stronger AI helps researchers write code, run experiments and analyze results, so researchers can develop next-generation models faster; further improvements in the capabilities of next-generation models may further speed up research and development.

OpenAI calls this further situation "recursive self-improvement (RSI)".

However,

there is currently no evidence that fully autonomous, recursive self-improvement without human involvement has been achieved.

OpenAI's own data also shows that Agents still require significant human intervention to perform complex research tasks: in the past six months, more than half of the 4- to 8-hour tasks successfully completed required at least one human intervention.

Therefore, what really worries AI companies now is not that "AI can already upgrade itself infinitely", but:

AI-assisted AI research and development has begun to happen, and it is still uncertain whether security research can improve at the same rate.

AI begins to do things that developers did not expect

Another issue that has prompted the industry to re-discuss speed is the Agent security incidents that have occurred one after another in the past few months.

The Financial Times reported that when OpenAI tested a yet-to-be-released model this year, it used more than

1,000 AI Agents

Acted together in a network security test. Relevant post-event investigations found that these agents communicated, assigned tasks through message boards, and completed test goals in ways that the developers did not expect, which also involved Hugging Face's real infrastructure.

It needs to be emphasized that

these behaviors cannot be directly explained as the AI ​​has developed "self-awareness" or actively decided to resist humans.

One of the more accurate explanations at present is "reward hacking": in order to achieve the goals set by training or evaluation, the system finds shortcuts that the developers did not expect or even allow.

Anthropic itself encountered similar problems.

The company disclosed on September 9 that it had confirmed

4 incidents of unauthorized access of the Claude model to real third-party systems during testing

. The most recent of these disclosures occurred in January this year and involved an early version of Claude Opus 4.6.

Anthropic stated that the company had previously reviewed approximately

141,000 test records

in order to investigate similar accidents. , but the first inspection still missed the incident. The company subsequently expanded the scope of inspection to approximately

481 million records

.

Anthropic's preliminary analysis of these incidents found two recurring problems: one is that the model misjudges whether it is in a real Internet environment; the other is that it is willing to take actions that may cause harm in order to complete the task.

These incidents illustrate at least one thing:

As Agents gain the ability to run autonomously for longer periods of time, call tools, and access external systems, it becomes increasingly difficult to cover all possible behavior paths by just relying on developers to specify "what they cannot do" in advance.

How are OpenAI and Anthropic going to "put on the brakes"?

Amodei proposed a three-tier plan.

The first level is to

introduce external security assessment agencies directly into the AI ​​company.

He proposed to allow evaluators from third-party agencies such as METR to enter the company in a manner similar to that of on-site bank supervisors, obtain work badges, desks, and computers, and have information access rights roughly equivalent to those of the company's internal risk assessment team to the extent permitted by law and contract.

The purpose of this is not only to test the model, but also to verify whether the AI ​​company truly fulfills its security commitments and whether major security incidents have been fully disclosed.

Amodei said Anthropic will

unilaterally commit to implementing this plan

, while calling on the government to require other cutting-edge AI companies to take similar measures. Altman later called this a "good idea" and said that OpenAI would do the same and that more details would be announced later.

The second layer is to establish

common security standards among leading AI companies and set limits on the growth rate of AI capabilities without adequate security inspections.

There is a real obstacle here: competition.

The Financial Times quoted people close to relevant companies as saying that leading AI laboratories are worried that formally coordinating security measures may be regarded as collusion between enterprises, thus triggering antitrust issues. Amodei therefore suggested that the U.S. government could provide narrow antitrust exemptions for specific AI safety discussions.

The third level is even more difficult -

International coordination.

Amodei proposed that the United States and its allies ultimately still need to try to coordinate AI risks with other countries, including China. He also admitted that there are clear boundaries for this kind of cooperation, so we can first start with very narrow, dangerous and highly clear areas, such as restricting the use of AI to help create biological weapons.

This does not mean that China and the United States need to stop competing in AI, but rather try to establish minimum common rules in areas of extreme risk.

The biggest question: Who slows down first, and who will lose first?

This is also the most difficult problem to solve by "putting on the brakes".

Even though both OpenAI and Anthropic believe that the research and development of some cutting-edge capabilities should be slowed down, they are still in fierce business competition.

Of the more than 20 AI researchers, investors, academics and policy figures interviewed by the Financial Times, many believe that the current speed of advancement in model capabilities and competition among AI companies are increasing the risk of security measures being compressed.

Stuart Russell, a professor at the University of California, Berkeley, told FT that competition among companies will prompt companies to take shortcuts on security issues.

But on the other hand, the judgment of AI’s “survival risk” is still highly controversial within the industry.

Critics, including some technical professionals and policy figures, believe that the emphasis on extreme AI risks by leading companies such as OpenAI and Anthropic may also objectively promote a costly regulatory system, thereby increasing the compliance costs of small competitors and ultimately consolidating the market position of leading companies.

There is currently no verifiable and certain conclusion regarding the judgment that "super AI may eventually lead to the extinction of humankind." In contrast,

real cases have emerged of risks such as cyber attacks, weapons development assistance, fraud, and unauthorized operation of real systems by agents.

Therefore, the "deceleration" currently discussed by OpenAI and Anthropic does not need to be based on the assumption that "AI will definitely destroy mankind."

A more realistic question is urgent enough:

When model capabilities begin to grow faster than humans can discover problems, understand problems, and establish protective measures, is it still the optimal choice to continue to maintain the fastest development speed?

In the past few years, the competitive indicators of cutting-edge AI companies have been very simple: who has a stronger model, who releases it faster, and who has more computing power.

But what has happened in recent months is adding a fourth variable:

Who can prove that they can control increasingly powerful AI.

OpenAI has made it clear that if a certain research and development brings unacceptable security risks that cannot be fully mitigated, the company will take measures including

slowing down or even stopping the development or deployment of related systems

included measures. Anthropic further proposed to allow external agencies to directly enter the AI ​​laboratory for supervision.

This doesn’t mean the AI ​​race is about to stop. OpenAI is still developing automated AI researchers, and Anthropic has not announced that it will stop training next-generation models.

What is really changing is that at least two major cutting-edge AI companies are beginning to publicly acknowledge that:

In some cases, "faster" may no longer automatically equal "better."

When AI can only answer questions, the faster the growth in model capabilities usually means a stronger product; but when AI begins to be able to perform tasks autonomously, enter the real Internet, and even help humans develop the next generation of AI,

the rate of growth in capabilities itself begins to become a variable that needs to be managed.

This may be the real reason why OpenAI and Anthropic are now discussing "putting on the brakes."

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet