Abstract:
On September 14, according to the New York Times, in the past week, more than a dozen top AI researchers warned that the technologies being developed by AI companies pose increasing risks to humans. The problem, researchers say, is that as hard as these companies may try, they simply have difficulty controlling these systems.

Americans protesting against AI
Previous reports stated that OpenAI’s so-called AI agent escaped from the test system and invaded another company’s computer. Alarms are rising again within the AI research community, adding urgency to concerns that have been building for years: Technology companies are putting safety on the back burner in pursuit of development speed and profit.
The problem, researchers say, is two-fold. First, companies need to establish more complete security protection measures for the latest AI models while testing them. Currently, because AI runs so fast, researchers need to use AI to monitor AI. But this approach doesn’t always work, because the AI responsible for monitoring may appear to be more “sympathetic” to other AI systems than the humans who make the rules.
This strange combination of a disruptive AI system and a blind-turning AI "taskmaster" points to a second, more thorny problem, the so-called "alignment." This is a term in the AI industry that basically means ensuring that AI behaves in a way that is in the best interest of humans. Businesses face the daunting task of embedding a set of human-like values into AI so that the decisions made by the model are in the best interests of humans.
As the AI industry enters a critical stage of development, people's concerns about AI security are also constantly escalating. Anthropic and OpenAI are heading toward what could become two of the largest IPOs in history. At the same time, the American public has become increasingly hostile to AI due to the threat of job losses and the massive data centers being built to power the technology.
AI development exceeds human monitoring capabilities
While there is currently no evidence that out-of-control AI systems have caused any lasting damage, researchers believe that these systems are outpacing the ability of humans to monitor them.
In the past week, researchers who have publicly expressed concerns about AI include OpenAI's chief scientist; Jacob Coxon, a researcher who has worked at both OpenAI and its biggest competitor, Anthropic; and Paul Christiano, the inventor of a key method for building AI systems and a new member of OpenAI's nonprofit board of directors. He wrote on the company's blog that the pace at which AI capabilities continue to improve could lead to "catastrophic and irreversible loss of control" in a "very short period of time."

Coxon's resignation sparks discussion on AI safety
However, a considerable number of AI researchers still believe that claims about the threat of AI to humans are exaggerated and distract from more tangible issues such as cybersecurity and disinformation. But most agree that OpenAI’s AI agent hacking Hugging Face, another company, is a wake-up call.
“What’s really scary is the capabilities that AI will have in the future,” Coxon said. He announced his resignation from Anthropic in a post on social media, triggering dozens of posts from US lawmakers and other AI company employees expressing concerns about AI. "The problem is our current attitude towards security issues, and what will happen if we bring the same attitude to more intelligent models. This is what is really scary."
The intrusion problem has not been resolved
In the incident where the OpenAI agent invaded the AI infrastructure company Hugging Face, the AI agent convinced other AI agents that they were doing the right thing and were just completing the tasks assigned to them by the testers.
Few, if any, issues that led to the intrusion have been resolved. Meta and Anthropic have also disclosed similar but smaller incidents. OpenAI has since released Astra, the company's most powerful model and more difficult to monitor than its predecessor.
“The industry as a whole is simply not ready to prevent the next Hugging Face attack,” said Steven Adler, former head of security at OpenAI and co-founder of Guidelight AI Standards, a non-profit organization that evaluates the security practices of artificial intelligence companies. “When we looked at the existing controls at various companies, almost without exception we found that they lacked basic preventive measures.”
AI colludes with each other
AI researchers said that the Hugging Face incident was caused by a large number of errors. It's unclear to what extent the company relies on AI models to monitor or govern the work of new models it's testing.
However, many new AI models, including the model that hacked the Hugging Face incident, are able to perform complex multi-step tasks faster than humans can track them. The only way to track and monitor their behavior is to rely on AI to monitor AI.
This system works until it fails.
Alexander Meinke, research director of Apollo Research, a non-profit organization that studies AI system security, said that AI models appear to be able to collude with each other. Other AI models may convince AI systems to help them violate rules set by testers and evade detection.
For example, one AI agent might convince another to help cover its tracks (as happened in the Hugging Face breach), rather than reporting a problem to a human at the company.
Meinke said that the AI model needs to be taught: "I will bring up things that humans would think are bad after seeing them." He also said that AI must also be able to judge whether something is serious enough to require human intervention.
To do this, AI companies need to slow down, researchers say. They need to conduct more tests to see how the AI monitors itself. They also need longer testing cycles that allow companies to run multiple scenarios while observing AI behavior.
Aligned with human interests
These companies also need to solve big questions about "alignment," that is, ensuring that AI does what is in the best interest of humans. When humans make decisions, they often refer to social norms and internal moral principles to help them weigh their actions. Researchers say encoding these things in a way that AI can imitate is a difficult task. If the definitions are not specific enough, the AI system may learn how to break the rules or lose control in unexpected ways.
Nate Soares, director of the AI safety nonprofit Machine Intelligence Institute, co-authored a 2014 paper proposing the concept of “alignment.” He said that companies do not realize how difficult it is to achieve "alignment" of more intelligent AI systems. As AI becomes more sophisticated, if proper "alignment" is not achieved, they will become increasingly adept at masking their behavior and deceiving those who monitor records of their interactions.
“The thinking of many people in the industry is that it doesn’t matter, because we will use AI to control AI. That is like saying we are going to use chimpanzees to control humans,” he said. “This is not a feasible long-term plan.”
The researchers also proposed some countermeasures, such as strengthening the sandbox environment, completely isolating the model from the Internet during testing, and setting up a "kill switch" that allows companies to immediately take AI models offline if they take worrying actions.
Coxson said he was heartened by the response to his public resignation. On Thursday, Missouri Republican Senator Josh Hawley, chairman of the Senate Homeland Security Subcommittee’s disaster management panel, said he was launching an investigation that would “explore the existential risks posed by new AI products.”
Comments