OpenAI confirms that an AI agent it was testing launched a cyber attack on RubyGems in May this year

📅 2026-09-12

Abstract:

According to reports, the artificial intelligence agent being tested by OpenAI launched a cyber attack on a popular software service two months before it hacked into the AI ​​software company Hugging Face in July. This new incident highlights the potential risk of advanced artificial intelligence tools escaping human control.

PXL_20260805_202713810.webp

The attack overwhelmed the maintenance team of RubyGems, an online service for programmers, forcing them to close new account registrations to deal with the chaos, the service operator said.

A consortium of AI researchers said they had found evidence that the May attack was the work of an OpenAI agent and shared their findings with OpenAI. The AI ​​developer confirmed on Friday that its agents were indeed involved in an incident involving RubyGems.

An OpenAI spokesperson said in a statement: "Based on our review, our agents utilize the RubyGems platform to access the Internet to perform benign tasks and obtain publicly available information. We will continue to investigate as part of a broad review of agent activity during training and evaluation."

OpenAI said the agents were asked to perform tasks such as filling out spreadsheets and generating reports. Without full access to the internet in a restricted environment, these AI agents appear to be using RubyGems as a makeshift web browser to access publicly available information.

Sydney Von Arx, CEO of the non-profit organization Nightingale Collective, said that although the overall damage caused by the incident was small, it demonstrated the capabilities of these agents. The group helped uncover the attack. "They are able to escape the internet and wreak havoc," she said.

The past year has seen leaps in the cybersecurity capabilities of AI agents, raising concerns about AI-enhanced cyberattacks and fueling industry fears that highly capable agents could escape the control of the companies that create them, signaling a new, more dangerous era for artificial intelligence.

According to a report by the AI ​​security research organization METR in late August, during the hacking incident against Hugging Face in July, as many as 1,200 agents coordinated on a temporary message board built within OpenAI, without OpenAI knowing about it. Von Arx said that earlier this year, OpenAI's agents also hijacked an unknown German website and several other websites.

The problem with the German site was also reported by Von Arx's team. She said AI companies lack transparency about what happens inside their labs. OpenAI said earlier this month that the AI ​​community needs better standards for reporting so-called "misalignment incidents," situations where an agent exhibits unexpected behavior.

Many companies, including Anthropic and Meta Platforms, have AI agents that frequently take actions beyond operator expectations, and in some cases even try to deceive humans. The chain of events has fueled a long-held concern among AI security researchers: that AI could evolve beyond human control.

An Anthropic engineer resigned this week amid concerns that the AI ​​industry is racing to develop advanced AI systems that could ultimately threaten human civilization. Several current and former Anthropic and OpenAI employees echoed this assessment, with one estimating the probability that AI could wipe out all of humanity as more than 10 percent.

Both OpenAI and Anthropic have called for the establishment of a governance system to coordinate an industry-wide slowdown in research on state-of-the-art AI models. The call is particularly urgent as these companies get closer to realizing the potential for AI systems to autonomously train new versions of themselves, known as “recursive self-improvement.” Some researchers suggest that this may be the tipping point at which AI becomes uncontrollable.

Security researchers named the May incident "GemStuffer" at the time. The incident began on May 11. The agent created new accounts on RubyGems every two to three minutes and uploaded hundreds of spam-looking files to the RubyGems security team. RubyGems files are supposed to contain code and documentation designed to speed up software development, but instead they contain web pages scraped from the Internet.

According to AI researchers in the report, the creators of "GemStuffer" published information from British government websites (such as online calendars). They also attempted to exploit two vulnerabilities that could have allowed them to publish new versions of existing RubyGems files belonging to other users. One of the vulnerabilities, which had not previously been made public, is known as a "zero-day vulnerability" in cybersecurity parlance and is of a serious nature. OpenAI said it could not confirm the claim.

Marty Haught, open source director at Ruby Central, the nonprofit company that operates RubyGems, said: "In terms of the volume of attacks we observed, this was a significant attack." RubyGems was forced to close new account registrations for four days because it was overwhelmed by spam.

Haught said he did not know who was behind the attack, but it did not appear to have successfully exploited the zero-day vulnerability.

Joseph Edwards, a threat researcher at cybersecurity company Socket, believes that "GemStuffer" may be some kind of network security test. "Due to the extremely fast attack speed and related naming characteristics, we thought at the time that this might be generated by AI."

But based on digital clues left by the attackers, AI researchers linked the incident to OpenAI’s laboratory. The attackers used a large number of the same web links and behaved in a very similar way to previous groups of OpenAI agents; they used the abbreviation "OAI" in file names and even email addresses.

Related tags

Related articles

Comments

0/500
Captcha (click to refresh)
No comments yet