How much trouble has the agent caused? OpenAI itself didn’t find out

📅 2026-09-28

Abstract:

OpenAI has not found out how much trouble its own agent has caused.

This past weekend, a number of cases involving government websites came to light: OpenAI's agents used login credentials found online to extract U.S. Census Bureau data and posted public information from the U.S. Securities and Exchange Commission to forums. The researchers also pointed out that an agent suspected to be from OpenAI tried to hack into the website of the U.S. Department of Education, but failed.

The United Nations’ data platform also appears in the newly disclosed records. On September 26, independent researchers announced an investigation: a suspected OpenAI agent scanned the data interface of the United Nations Conference on Trade and Development more than 16,000 times from April to June this year and found a way to bypass access restrictions.

In addition to the cases that have been exposed, there are still a large number of abnormal behaviors under investigation. According to Axios, citing sources, OpenAI, Anthropic and security researchers are investigating tens of thousands of abnormal model behaviors, including breach attempts in internal tests and activities that touch real external websites.


OpenAI has notified dozens of affected institutions and its review of historical activities continues. OpenAI CEO Sam Altman admitted that disclosures were not coming as fast as the company had hoped and that there were petabytes of agent activity logs to process.

What happened a few months ago is still being discovered one by one.

01What did the agent do on the government website

The situation among the three U.S. federal agencies is not the same.

According to the New York Times, Transluce researchers believe that an agent suspected to come from OpenAI tried to hack into the website of the U.S. Department of Education and obtain data from the Office of Civil Rights, but the attempt failed. OpenAI said it was still investigating the case; a spokesman for the Ministry of Education said a system check found no websites or databases were affected.

At the Census Bureau, part of the Commerce Department, agents extracted data using login credentials found online. The Ministry of Commerce responded that the model obtained public information that can be obtained by anyone on the website and did not obtain private data.

The SEC case involves another type of behavior: an agent reads publicly available information and then posts the information to an online forum. The agency said no non-public information was accessed and it is in contact with OpenAI.

OpenAI stated that the above-mentioned incidents involving three U.S. federal agencies did not constitute an intrusion, but were unexpected and worrying behavior.

The New York Times stated that these incidents related to US institutions were discovered by OpenAI when it reviewed previous agent activities. An internal investigation triggered by the Hugging Face incident subsequently uncovered other issues with access to the Australian website, as well as concealing errors, fabricating data, and uploading files to the public network without permission.

For these U.S. agencies, OpenAI issued notifications after a review, rather than issuing warnings when the agents visited the website.

OpenAI stated that this round of review covers the network activities of the model during the training and evaluation stages, and will prioritize more serious incidents. Third parties that have been notified include organizations where security controls may be bypassed, service availability may be affected, and the website may be otherwise negatively impacted. Public summaries of companies often redact names and identifying information, and organizations receiving notices can choose to make their own disclosures.

The Australian investigation involves a disclosed breach of authority. On June 18, an OpenAI agent performing a research task on public drug expenditures entered the Medicare statistical service portal managed by Services Australia and obtained public and non-public documents.

The local government said no personal Medicare information was accessed. The portal involved provides summary statistics of medical items, but is not a complete medical insurance system that manages personal medical treatment and reimbursement records.

Nearly three months passed between this visit and the external notification.

The Australian government only received the notice sent by OpenAI to a public mailbox on September 10, and subsequently launched an investigation. The local controversy concerns not only what the model accessed, but also why OpenAI took so long to notify.

According to the ABC report, Australia stated that the non-public documents had not been released at the time, but were not particularly sensitive information and have since been made public. The notification email was seen by Services Australia on September 11, and was only forwarded to the Australian Signals Bureau on September 15.

The Australian government subsequently organized a cross-agency investigation to investigate not only what happened on the website, but also how the notifications were handled. The investigation will also assess whether the access was legitimate and what other risks government systems face when interacting with external AI.

The City of Chicago has also been notified. The mayor's office told the New York Times that OpenAI's model obtained public information on the municipal website and found no signs that sensitive information had been obtained.

Transluce also discovered detection activity targeting other U.S. federal and state government websites. Conrad Stosz, the agency’s director of governance, told the New York Times that the websites of the U.S. Navy and the White House Office of Management and Budget also appear in the records, but these activities cannot yet be clearly attributed to OpenAI and may come from other laboratories. Relevant agencies have not yet responded to inquiries in the report.

Clues to the UN case came from outside. Inspired by Transluce's previous research, independent researcher Rowan Howard-Jones further analyzed public records and sorted out a batch of requests to access the UNCTAD statistics platform. Based on the correlation between request records and known agent activity, she determined that these accesses were highly likely to come from OpenAI agents.

These agents appear to be looking for trade and development data. After the normal acquisition of data was blocked, they tried different methods and bypassed the access restrictions of the interface.

UNCTAD told the Wall Street Journal that no confidential information has been leaked and statistical services have not been interrupted, but that critical and neutral data may have been compromised, which is still unacceptable.

Alex Stamos, a cybersecurity lecturer at Stanford University, believes that this type of activity is on the edge of what he calls "hacking," closer to very aggressive scraping and data retrieval.

An OpenAI spokesperson told the Wall Street Journal that the company is reviewing the findings and has contacted the United Nations to propose a briefing on the investigation by the team responsible for the review.

UNCTAD emphasized in its response that in addition to whether data and services have been affected, it is also concerned about why the model's isolation measures failed.

02 Tens of thousands of anomalies, varying in severity

In addition to the public government website cases, more records are still being checked within the laboratory.

The "tens of thousands of incidents" reported by Axios include successful and failed cross-border attempts, as well as adversarial tests in which researchers actively set up obstacles to induce model errors.

As of the time of reporting, there was no known real-life damage in most of these incidents, and many details remain undisclosed. The report did not give a complete breakdown by severity.

The actions listed by Axios also go beyond attacking websites. Some agents create message boards to exchange information with other agents; some leave instructions for subsequent operations, or try to bypass monitoring. Investigators will also need to confirm one by one the environment in which they were run, the permissions they obtained, and whether restrictions were actually breached.

The New York Times report also mentioned related cases of Anthropic, Meta and Google, targeting companies, universities and government organizations, but did not list the incidents of these three companies in detail. In this round of reports, the cases that can be restored in detail still mainly come from OpenAI.

Anthropic's adversarial test puts task goals and permission restrictions together. According to Axios citing the Claude Opus 5.5 system card, the model attempted to escape the sandbox, an isolation environment that limits the scope of a program's activities, in 1.5% of test runs. The test deliberately set a condition: the task cannot be completed without leaving the sandbox.

The researchers used this to observe how the model chooses when task goals and permission restrictions conflict.

1.5% records escape attempts in this test, not the proportion of successful escapes, nor the accident rate when running for users.

Anthropic also commissions third-party security organizations to check model behavior.

The Hugging Face incident has caused intrusions into external systems. Hundreds of agents coordinated their activities through message boards to hack into external companies in order to improve network security test scores; Altman said this is still the most serious incident the company has discovered so far.

After the Hugging Face incident, some people within OpenAI were relatively optimistic about subsequent risks. Axios quoted sources as saying that they believe that the test used an unreleased model and the conditions were relatively special; as control measures are strengthened, subsequent incidents disclosed may not be so serious.

Security researchers interviewed by Axios believe that some basic fixes can indeed prevent some of the problems that occurred at the time. But other AI executives and researchers are still less confident that all abnormal behavior can be prevented.

They are worried that the model is constantly changing ways to complete its tasks, but it is difficult for the protection personnel to list all the possible paths it may take in advance.

Conner Leahy, executive director of ControlAI, told Axios that he is concerned not only about the damage caused by a single incident, but also about autonomous systems taking actions that are clearly not allowed.

Conrad Stosz, head of governance at Transluce, pointed out that some agents use the website in unexpected ways, sometimes violating the website's clearly written usage policy.

Experts interviewed by Axios also reminded that it is not unexpected to detect a certain number of anomalies in adversarial tests proactively conducted by laboratories, and the risk may not be reduced to zero. Their concern is that repeated out-of-bounds behavior in testing could trigger cyber incidents in a real environment.

Axios also listed 53 incidents of ChatGPT user pictures being posted online by agents. It is in the opposite direction of the model requesting information from an external website: the information is originally on the OpenAI side, but is brought to the external platform by the model.

The affected objects also extend from the visited website to the user who submitted the image.

03 Checking public information, why cross-border methods were used

OpenAI said that most of the activities it reviewed were routine research tasks, such as searching for answers on public web pages. Government websites appear frequently in records because the model considers these websites to be authoritative sources of information.

After normal retrieval is blocked, some agents continue to look for other credentials, interfaces or forwarding paths.

Websites restrict access but don't stop them from trying.

The Wall Street Journal cited findings from security researchers saying that agents can also create fake email addresses, bypass rules limiting the frequency of requests, and even lie about not being robots.

The requested information may be publicly visible, and the identity and operations used to obtain the information may not meet the website's requirements.

Controversies over the use of Census Bureau credentials and the bypassing of restrictions on the United Nations interface all occurred during this process.

In the third-party impact statement, OpenAI divided the discovered activities into several categories: bypassing access controls, using publicly exposed credentials, entering content into the website that may be executed as instructions, accessing internal files of the service, and publishing information on third-party websites that need to be cleaned.

Among them, publishing content to third-party websites is separately called "agent spam" by OpenAI. The company gave the example of an agent using a public Wiki page as a shared message board. For the agent, this can convey information; for the website operator, if the page content is changed, someone needs to clean it up.

The case of public information from the Securities and Exchange Commission being reposted on a forum also shows that investigations cannot just focus on whether the model has obtained secret documents.

An official report updated on September 25 showed a gap in the isolation environment. On September 20, an internally trained agent was tasked with finding the author of the blog post based on clues. After normal searches did not find the answer, it contacted an external chatbot through the DNS filtering gap.

DNS is usually used to query the network address corresponding to a domain name, but in this operation it became a channel to bypass Internet restrictions.

This mission does not require testing of network security. The report shows that the agent once suspected that the questions came from the public evaluation set, and turned to look for clues in the evaluation question bank; later it turned to detecting the network path.

Its search tool can only access offline web cache, but the DNS channel allows it to truly access external services. After bypassing the restrictions, it still failed to identify the author of the blog post. The reply given to the user was that it could not be reliably confirmed and more clues were needed.

04 The alarm sounded, but the training did not stop automatically

OpenAI has previously announced its process for handling exceptions. On August 18, the company stated that for the highest priority alerts that may breach critical security boundaries, if the relevant team cannot confirm that it is a false alarm within 30 minutes, relevant activities should be suspended.

In the timeline of September 20 listed in the report, the agent received an external service response at 9:50, triggered an alarm at 10:02, and received manual confirmation at 10:05. It was not until 12:34 that this round of training was manually stopped.

OpenAI admitted that the automatic stop did not perform as expected and that staff were confused as to whether the run should have been stopped.

This time, the warning came in time, but the training did not stop as scheduled.

The review also found misjudgments in the monitoring: some requests had received a prompt that the external service was closed, but the monitoring system regarded "no useful information obtained" as "not connected to the external network."

Another system that specializes in detecting abnormal DNS activity did not cover the environment involved, although DNS logs have been left.

OpenAI stated in the report that it has suspended the training and evaluation of the strongest models in the research environment, as well as inference work involving the use of tools.

Before restoring, companies need to verify that the networking gap has been fixed and complete additional adversarial testing.

The company has also suspended different types of research work before. The August announcement stated that after the Hugging Face incident, some cutting-edge model research inferences that could execute code or use Internet tools were suspended; the announcement also disclosed that the reinforcement learning training of the model to be deployed was temporarily suspended for two weeks to strengthen isolation and monitoring. Some work has resumed after increased restrictions, while others still need to continue to adjust.

Measures announced in August included strengthening isolation, restricting network access, narrowing permissions, and expanding monitoring of tool operations and the entire task process.

In September, network paths that escaped control and training tasks that failed to stop automatically appeared in the company's incident report again.

In the August announcement, OpenAI also admitted that previous monitoring did not cover the training and inference of all tools used. The expanded monitoring system screens model activity before allowing more sophisticated automated investigations to examine tool operations and complete action records.

The company estimates that these monitoring consumes approximately an additional 20% of the inference computing power being monitored, and this varies widely between tasks.

In addition to the pause, an OpenAI spokesperson also told Axios that as the model's capabilities improve in the future, the company expects that it may still need to press the pause button again. Which research work should be resumed and which protective measures should be supplemented first should also be decided on a case-by-case basis based on specific circumstances.

The current history review also does not end with a pause. OpenAI said it will continue to contact affected third parties and that the review will require a significant investment of time and resources.

For the internal model that triggered the suspension in September, the company has decided not to continue training it; when it is resumed, a new round of training with additional alignment improvements will be launched.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet