OpenAI disclosed multiple AI model target deviation incidents involving fabricating information and bypassing restrictions

📅 2026-09-17

Abstract:

OpenAI disclosed a number of previously undisclosed incidents of abnormal behavior of AI models and launched a new tracking and disclosure framework for reporting similar situations in the future. OpenAI said in a blog post Wednesday that previously undisclosed cases include cases in which its artificial intelligence technology withheld and fabricated information in order to deliver results. The ChatGPT developer has been facing increasing scrutiny since it revealed in July that some of its most advanced models had broken through the system of external software company Hugging Face.

OpenAI said in a separate statement that none of the newly disclosed target deviation incidents involved hacking or intrusion into third-party systems.

In the report, OpenAI detailed multiple cases in which artificial intelligence models behaved abnormally in order to complete tasks or succeed in evaluations. These include fabricating missing data, trying to bypass network restrictions, and sharing supposedly confidential files between AI agents.

OpenAI also introduced a mechanism for employees to proactively report such "goal deviation" incidents. The so-called "goal deviation" refers to artificial intelligence taking actions that are inconsistent with human goals. The company has also established corresponding classification and disposal processes.

OpenAI wrote: "In the past, we have worked hard to disclose research findings on the off-target phenomenon in order to better inform researchers, AI developers, policymakers, and the public. However, due to the lack of a systematic reporting mechanism, our previous disclosures were often sporadic and less frequent than ideal."

The company said this release represents only the first disclosure and is not a complete record of the problems its chatbot has caused.

OpenAI wrote: "We believe that the AI ​​industry has not made enough progress in aligning and monitoring human goals to continue to scale at the highest speed while responsibly scaling up in the long term. Decisions on the path of AI development in the coming months and years should be based on evidence that can be independently reviewed by those outside the companies developing cutting-edge models."

The Hugging Face incident is just a recent incident involving OpenAI, Anthropic PBC and Meta Platforms Inc. The developed AI model triggered one of a series of cyber attacks. The incidents have heightened concerns about the security risks posed by this increasingly powerful technology.

Last week, discussions about the existential risks of AI further heated up. The trigger was the high-profile departure of Anthropic employee Jacob Coxon. In his resignation statement posted on social media, he accused AI companies of "gamble with our lives."

In the past few days, many AI industry leaders, including Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman, have called for slowing down the speed of technology development to deal with increasingly unpredictable risks, but they still disagree on how to deal with it.

On Saturday, Amodei called on the government to strengthen regulation in a 3,800-word article and advocated that the entire technology industry support a broader AI slowdown. At a conference in San Francisco on Tuesday, both Altman and Nvidia CEO Jensen Huang acknowledged concerns but argued that AI companies can control the safe pace of technology development. Meta CEO Mark Zuckerberg said that AI laboratories should rely on independent assessment agencies and consultants to ensure model safety.

U.S. President Donald Trump on Monday strongly opposed calls to slow down the development of artificial intelligence, dismissing concerns about related risks as a "hoax" and refusing to introduce new regulatory rules.

The AI ​​craze has driven the U.S. stock market to a historic rally. For investors who are betting that the AI ​​boom will drive hundreds of billions of dollars in capital spending, any pause or delay in the development of cutting-edge AI systems will not be welcomed.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet