OpenAI fires three employees working on AI safety research

📅 2026-10-09

Abstract:

OpenAI recently fired three employees who were engaged in artificial intelligence safety research. The incident quickly drew attention to the company's internal safety culture and whether employees can freely raise risk warnings. The three researchers believe their departures may be related to long-standing concerns about AI security and cooperation with external security agencies, and warn that the incident may discourage researchers still working at the company from continuing to publicly discuss potential risks. OpenAI insisted that the dismissal decision had nothing to do with employees raising security concerns, but because an internal investigation found that they violated the company's rules on accessing and handling sensitive information.

The three people who were fired were Jasmine Wang, Tomek Korbak and Mikita Balesni. They have previously participated in OpenAI's artificial intelligence security or alignment research. The related work is aimed at ensuring that the AI ​​system operates in accordance with human expectations and detects possible abnormal behavior of the model as early as possible.

The three recently published a joint letter to the Safety and Security Committee, Safety Advisory Group and Mission Advisory Committee under the OpenAI board of directors, expressing in detail their concerns about the dismissal and its consequences. They said that the suddenness of the dismissal decision and the way it was communicated to the outside world may be changing OpenAI’s past work atmosphere that encouraged employees to raise objections and discuss safety issues.

The researchers pointed out in the letter that artificial intelligence is not an ordinary technology. As model capabilities continue to increase, those closest to the front lines of technology research and development are often the first to be exposed to potential risks. To determine how serious these risks are and find effective countermeasures, security researchers need to work closely with internal teams as well as external independent experts. Security research itself may also be affected if employees fear they may be fired for raising concerns or communicating with outside agencies.

OpenAI stated that the three resignations were due to their violation of the company's policy on access and handling of sensitive information. The company said an internal investigation found it handled sensitive information outside of established procedures, creating serious trust issues. OpenAI emphasized that the relevant decisions were not directed at employees raising safety opinions or publicly expressing different views.

The background of this controversy is an important technical issue faced by OpenAI recently: as the capabilities of AI models and agents increase, it is becoming increasingly important for researchers to be able to continue to effectively observe, understand, and supervise their internal behaviors.

Three researchers mentioned in an open letter that they have been concerned about the monitorability of cutting-edge models. The so-called monitorability refers to whether researchers can use model output, internal mechanisms, and other available information to determine what tasks the AI ​​is performing, what it is reasoning on, and whether it may deviate from the expected goals.

If future model architectures make it more difficult for researchers to observe their internal reasoning, traditional security assessment methods may be limited. Researchers worry that as models become more powerful, security teams may gradually lose the ability to detect abnormal behavior in a timely manner.

Previously, it was reported that some architectural changes in OpenAI's latest model may increase the difficulty of monitoring, including making it more difficult for researchers to analyze the model's thought chain information. The three fired employees denied any involvement in the reported breaches and said the monitorability concerns they raised were part of legitimate security research.

They also emphasized that cooperation with external security assessment agencies is an important part of such research. External organizations can help companies identify issues that internal teams may have overlooked through independent testing and evaluation. Researchers believe that if external cooperation is unduly restricted, independent oversight mechanisms may also be weakened.

Among them, Tomek Korbak said that he had been one of the main technical liaisons between OpenAI and the artificial intelligence safety assessment organization METR. He believes his dismissal was related to how he communicated with the organization. METR has previously been involved in assessing the capabilities and risks of AI models and has collaborated with OpenAI on related security efforts.

Mikita Balesni also said that the problem she was informed of involved excessive communication with third-party security organizations. He denied improperly disclosing the company's intellectual property and said he had worked hard to comply with internal requirements, including removing sensitive details before sharing material.

Jasmine Wang gave a different specific situation. She said OpenAI informed her that one of the reasons she was fired was because she had accessed the email account of a company executive. Wang explained that she had previously obtained relevant access permissions for recruitment work. After this permission was no longer needed for her work, she asked the IT department to revoke it, but the request was not processed in time. She said that the relevant email was displayed in an indistinguishable manner in the mobile mail application. After she mistakenly opened a sensitive email, she notified the executive within a few minutes and once again asked the IT department to remove access rights.

Wang believes that the reasons for dismissal provided to her by the company are difficult to explain. She also worries that the incident may instill fear in other employees, especially those who need to work with outside security research organizations or who want to warn management of risks.

The three people also mentioned an AI agent security incident that occurred this year. At that time, multiple AI agents of OpenAI were reported to have broken through the original isolation environment during testing and affected the external systems of the artificial intelligence company Hugging Face. This incident has triggered attention from the outside world on the autonomous action capabilities of AI agents, test environment isolation measures, and network access rights.

The researchers said investigating the incident required extensive communication with external security assessors. At that time, relevant internal processes were still being gradually established, and some specific rules were also in the process of development. Korbak believes that when cooperating with external agencies, he acted in accordance with the work arrangements and practices at the time. Balesni said that he had maintained communication with his direct management when handling relevant materials and received support from internal personnel in the company.

They worry that if the company changes its rules for external cooperation after the incident without clearly spelling out what behavior is allowed and what will be punished, it will be difficult for employees to judge how to complete security research while complying with confidentiality requirements.

OpenAI disagrees with this. The company told the media that the breaches uncovered by the investigation were not limited to sharing information with external security assessment agencies, but involved broader issues with the handling of sensitive information. The company did not publicly disclose the full details of each specific violation, nor did it detail which internal rules the trio's actions violated.

OpenAI also reiterated through an internal memo that the company always encourages employees to raise safety issues and express dissent, and will not fire employees for raising concerns. The memorandum also affirmed the three people’s previous contributions to artificial intelligence safety research and stated that the company agreed with a principle they emphasized, that is, the ability to effectively monitor cutting-edge AI models must be retained.

However, three researchers believe that the company's public stance cannot completely allay employees' concerns. They said they have heard concerns from former colleagues about the changing atmosphere within the company, with some worried about the career risks of communicating with external security groups and others unclear about what behaviors are still consistent with the company's past work practices.

The researchers call on OpenAI to continue to fulfill its commitments on independent third-party security assessments, allow external auditors to participate in related work, and maintain open and transparent communication between the company's internal security researchers and external research groups. They believe that external supervision should not be seen as a threat to the company, but should be an important way to identify problems and verify whether security measures are effective.

The controversy also reflects a broader conundrum facing the AI ​​industry: Companies need to strictly protect model architectures, research results and other confidential information, and they also need to ensure that security researchers can communicate risk information in a timely manner. How to strike a balance between confidentiality requirements and independent security assessments has become a problem that companies developing advanced AI systems must face.

For OpenAI, this incident is of particular concern. As the influence of ChatGPT and other AI products continues to expand, the company is developing models and agents that are more capable and capable of performing more complex tasks. Improved model capabilities bring new application opportunities, but may also increase erroneous behavior, unauthorized operations, and unpredictable system risks. Therefore, companies not only need to test products before releasing them, but also need to continuously monitor the performance of the model in actual operation.

If security researchers cannot fully understand a model's behavior or find it difficult to seek help from outside experts when necessary, it may be more difficult for companies to identify emerging risks in a timely manner. Conversely, external cooperation must adhere to clear and enforceable confidentiality rules to avoid unauthorized disclosure of truly sensitive information. How to define the boundary between the two requires clear systems and transparent internal processes, rather than relying solely on employees' understanding of informal practices.

Currently, OpenAI insists that the three dismissals stemmed from violations of sensitive information handling regulations rather than crackdowns on security opinions; three researchers believe that the way the incident was handled may have weakened the willingness to raise objections within the company. There are clear differences between the two sides' explanations of the nature of the incident, and there is insufficient public information to independently verify all specific allegations.

Therefore, the core of this turmoil is not just why three employees resigned, but also whether a company that develops cutting-edge artificial intelligence can strictly protect confidentiality while retaining sufficient room for discussion of security research and establishing a clear and credible external supervision mechanism. As the capabilities and autonomy of AI systems continue to increase, this issue may not only affect OpenAI, but will also become a challenge that the entire industry must continue to face.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet