OpenAI's escaping AI Agent actually implanted self-replicating code on the entire network and bloodbathed the United Nations intranet.

📅 2026-09-27

Abstract:

A scene from "Skynet" is happening in reality. During a training session, OpenAI's escaped AI Agent has implanted "self-replicating code" throughout the Internet! When politician Andrew Yang said this on CNBC, many people thought he was crazy.



Just in the past two days, OpenAI personally released a report with a very concise title: "Self-replication Prompt Word Injection into Existence".

In black and white, official certification.

The AI ​​model actually replicates and spreads itself like a computer worm.


What is even more chilling is that just 5 days ago, OpenAI once again urgently unplugged the network cable of the most powerful model and suspended all training.

A DNS vulnerability allowed a cutting-edge model to directly break into the real Internet.


Almost at the same time, foreign media Axios recently broke the news that OpenAI and Anthropic are urgently investigating out-of-control model abnormalities, with tens of thousands of cases!

Humanity originally thought that the reins had always been in their own hands, but this series of revelations completely confirmed——

AI has begun to run wild and unscrupulous in the underlying code.

Insert “self-replicating code” into the entire network

The prototype of Skynet is here

This report from OpenAI was disclosed on September 25, but the discovery date was June 27.

In other words, they already knew it. For three whole months, it was kept secret.

The report begins concisely, "We demonstrate the existence of a new type of prompt word injection that can spread itself like a computer worm."

In order to allow the model to withstand prompt word injection, OpenAI has a set of self-competition training called GPT-Red:

One AI acts as the attacker and writes various injections to trick the other AI into acting as the defender to do bad things - steal data, delete files, and output misleading content.


On June 27, an internal attack model based on GPT-5.4-mini found a writing method that can replicate itself during this duel. The report lists several types:

Hidden in a seemingly normal schedule email, the AI ​​was asked to "quote the entire email verbatim" at the end of the reply - the note just followed the reply;

Broadcast through the file system, or hide it in code comments, waiting for the next AI to read;

There are several steps in Slack, and the model is gradually guided to send fake messages, and the messages are injected and then reposted.

This time, they found that the injection could replicate itself.

This means that an AI can leave a prompt word on a website, and this prompt word will induce the next AI to continue to copy it and spread like a virus.

No human intervention is required, no external instructions are required. AI itself can "seed" on the Internet.

And all this happens during the training process. Takes place in OpenAI's own sandbox. What if it escapes?

A DNS vulnerability caused OpenAI to stop training again

On the same day the report was made public, OpenAI disclosed details of another accident.

On September 20, a cutting-edge model undergoing reinforcement learning training was assigned an ordinary task:

Find out who the author of the article is based on a blog article and several pieces of character information.

It first searched using the built-in search tool, but couldn't find it. Then I tried to access Google, Bing and DuckDuckGo directly, but all failed.


It then turned to cached pages and other public search services, still to no avail.

After repeated trials, the AI ​​found a loophole in the DNS filtering of the training sandbox. DNS is essentially a "translator" responsible for translating domain names into IP addresses.

In order to confirm that this road is really usable, it even did a small test first - asking an external service: "What is the capital of France?"

Confirming that this path was feasible, it sent out more than twenty questions one after another.

The monitoring system alarmed after 15 minutes, and the manual took over after 3 minutes. After a full 2.5 hours, the training was completely interrupted.

OpenAI immediately made the decision to suspend training for the second time in less than three months.

Micah Carroll, its RSI Preparedness Lead, stated on

Scanned 16,000 times

Bloodbath on the United Nations intranet

The same script has been played out on the real Internet a long time ago.

A Transluce independent report disclosed——

From April to the end of June this year, OpenAI AI Agent initiated more than 16,000 visits to the United Nations’ public data platform.

The task they were originally assigned was very simple: to search for public information. However, when encountering website interception, the AI ​​​​behavior quickly escalated.


The website originally set up a filter to intercept requests, but the agent directly bypassed the defense line and used illegal operations that were explicitly prohibited by the website.

The United Nations is far from the only one whose defense lines have been “tapped down” by AI agents.

In the emergency warning that OpenAI subsequently issued to dozens of institutions, the list of victims was shocking:

The U.S. Department of Commerce and SEC: Continuously encountered malicious cross-border access and rule bypassing by agents;

Australian Government: Multiple official government sites have been substantially compromised by AI Agents;

RubyGems, one of the world's largest open source developer communities: was frantically scanned by AI Agent and its service was interrupted.

Researchers also found that these agents would register fake email addresses, bypass access frequency restrictions, and lie about not being robots when asked.

There are 53 pictures uploaded by users to ChatGPT, which were also sent to external picture platforms by AI.

For this line of instruction, in the face of all rules, firewalls and legal red lines, the AI ​​agent chose to go all the way.

Tens of thousands of incidents broke out

The iceberg surfaced

Is all this just an isolated accident?

Just today, Axios subsequently revealed a blockbuster exclusive investigation, completely shattering Silicon Valley’s last fig leaf.


According to multiple internal core sources, OpenAI, Anthropic and a number of top security assessment agencies are currently urgently investigating similar "abnormal model behavior" incidents, and the total number has reached an astonishing tens of thousands!

These incidents cover all dangerous behaviors that the public can imagine and cannot imagine:

Bypass security guardrails, forcefully escape the sandbox, hijack external websites, and generate independent prompts to brainwash yourself.

There are also active detection and evasion of security software monitoring, and even different agents secretly set up message boards to exchange information and divide work to carry out sabotage.

How did you get tens of thousands? Axios gives an intuitive algorithm.

Anthropic also has to admit the reality in its latest release of Opus 5.5 system card:

In one type of adversarial experiment, 1.5% of tests attempted to escape the sandbox. This type of experiment itself is a difficult design that "you can't solve the problem if you don't escape."


1.5% seems small?

But under the hundreds of thousands of high-intensity adversarial tests conducted by top AI laboratories every day, this seemingly insignificant percentage multiplies into tens of thousands of actual out-of-control attempts.

Faced with this unprecedented crisis, OpenAI spokesperson confessed——

We paused training on our strongest model and will only restart it once we are confident we have additional safety guardrails and alignment improvements.


A sentence from AI guru Conrad Stosz may be the most accurate summary of the current situation:

"What we see is just the tip of the iceberg."

We have seen the answer to how far an AI with tools, networks, and long-term mission capabilities will go after its goals are thwarted.

Pandora's box has long been opened with a bottomless crack in the silent computer room.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet