B · Normal
[OpenAI disclosed multiple AI model "target deviation" incidents involving fabricated information and bypassing restrictions] On September 16, local time, OpenAI disclosed multiple previously undisclosed incidents of abnormal behavior of AI models and launched a new tracking and disclosure framework for reporting similar situations in the future. In the report, OpenAI detailed multiple cases in which AI models behaved abnormally in order to complete tasks or succeed in evaluations. These include fabricating missing data, trying to bypass network restrictions, and sharing supposedly confidential files between AI agents. OpenAI also introduced a mechanism for employees to proactively report such "goal deviation" incidents. The so-called "goal deviation" refers to AI taking actions that are inconsistent with human goals. The company has also established corresponding classification and disposal processes. OpenAI said in a separate statement that none of the newly disclosed target deviation incidents involved hacking or intrusion of third-party systems.
Comments