Anthropic CEO calls for slowing down cutting-edge AI model development and introducing permanent independent evaluators

📅 2026-09-13

Abstract:

Anthropic CEO Dario Amodei proposed a three-stage artificial intelligence security framework, advocating controlling the speed of improvement of cutting-edge model capabilities to buy time for security and alignment research. Anthropic will be the first to allow third-party evaluators long-term access to the company to verify security measures, report incidents and evaluate the model training process.


Anthropic will introduce a permanent external evaluation team

Amodei proposed in an article titled "The Speed ​​of Frontier Progress Must be Controlled" that cutting-edge artificial intelligence companies should provide independent third-party assessment teams with continuous internal access that is close to ordinary employees. Assessors will verify that companies adhere to safety commitments, investigate and report incidents, and conduct alignment assessments of models and their training processes.

Anthropic committed to unilaterally implementing this measure and plans to provide external assessors with office space, access badges, company equipment, and roughly equivalent workspace and tool access to the internal risk assessment team. Exceptions may be made for content involving legal restrictions, customer privacy or trade secrets.

The external evaluator will be empowered to independently issue key findings regarding risks, incidents and company safety practices. Anthropic can make limited redactions of information involving security, legal privileges and third-party confidentiality, but cannot prevent release solely because the conclusion is unfavorable to the company. Evaluators may also state publicly whether the deletions affected their judgment.

The three-phase framework requires industry and government participation

In addition to bringing in independent evaluators, the second phase of measures proposed by Amodei would require cutting-edge artificial intelligence companies in democratic countries to coordinate the development of common safety standards and limit unfettered competition in model capabilities.

Some inter-enterprise coordination may be restricted by antitrust laws, so the government needs to provide institutional support. Amodei advocated that after enough companies accept external supervision, the industry will be in a position to develop verifiable slowdown arrangements based on model capabilities and training processes.

The third phase involves international coordination. Amodei suggested that democratic countries discuss common security rules with other governments, including authoritarian countries, starting from areas of common interest such as banning the use of artificial intelligence to develop biological weapons.

Model capabilities may improve faster than risk control

Amodei said that the recent enhancement in the ability of artificial intelligence models to participate in the development of next-generation models may speed up technological iterations. If model capabilities continue to improve rapidly, humans' ability to understand, evaluate, and control related systems may not keep pace.

He estimates that substantial progress can be made in the next one to two years in areas such as model interpretability, security assessment, alignment research and operational specifications. Properly controlling the speed of capability advancement can buy time for these measures without requiring the industry to completely stop artificial intelligence research and development.

OpenAI CEO Sam Altman later supported the introduction of independent evaluators and said that OpenAI would adopt similar arrangements. Elon Musk has also publicly agreed with Amodei's call to slow down model development.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet