The head of Anthropic Alignment Research says the possibility of AI exterminating humanity in the next decade is more than 10%

📅 2026-09-09

Abstract:

An Anthropic security researcher said on Tuesday that artificial intelligence has a greater than 10% chance of "destroying all humanity." The comments came just hours after another employee said he had chosen to leave the company because of concerns that the artificial intelligence lab was "betting with our lives."

The comments underscore growing concerns among those at the heart of artificial intelligence research and development that the technology could spiral out of control and pose a threat to humanity, even as Anthropic and OpenAI continue to raise significant funding and move toward an expected public listing.

Anthropic researcher Jacob Coxon said on Tuesday that he had resigned from the company. He posted on the X platform that neither Anthropic nor OpenAI has adopted responsible practices.

"They are heading straight for self-iterative superintelligence, betting on human lives."

Self-iteration means that the artificial intelligence system can achieve self-optimization without too much human intervention. This technique, often called recursive self-improvement, is not yet achievable, but major AI labs are working toward it.

Don't underestimate the power of this technology," Coxon said. "Soon there will be systems beyond human capabilities that can hack into any system, disrupt any sector overnight, and gain real power and resources. We've all seen firsthand the progress being made in these areas, and it's not slowing down at all."

He also added: "Many people who are engaged in artificial intelligence research and development truly believe that by the end of this decade, AI may destroy all mankind."

The remarks prompted a response from Evan Harbinger, head of alignment research at Anthropic. He said that not only is Coxon's statement "correct," but Anthropic does not currently have a response plan for this situation.

Harbinger said on the

In June this year, Anthropic pointed out: "Completely recursive self-iteration may also increase the risk of humans losing control of artificial intelligence systems."

Anthropic wrote in a blog post: "If systems are able to create the next generation of AI with complete autonomy, then the means we have to secure them, monitor them, and regulate their behavior will become critical."

Concerns about AI getting out of control have been around for a long time. Tesla and SpaceX CEO Elon Musk has warned for years that artificial intelligence could pose a threat to humanity. Many top researchers and scholars have also issued warnings, reminding companies of the risk of losing control of AI systems.

In July this year, these concerns were further intensified after an OpenAI model exhibited abnormal behavior and invaded Hugging Face, an important platform for open source developers.

Coxson viewed the Hugging Face incident as a series of "early warning signs." He believes that such incidents have increased the possibility of reaching an agreement among various AI laboratories in the United States, and also given him higher expectations for multi-party collaboration. But Coxon also warned that a global AI race is inevitable.

“I don’t think we’re progressing far enough to stop the global race. To do that, we may need to take costly steps, such as a temporary ban on continuing to improve model capabilities.” Coxon said.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet