Abstract:
Emergence, a startup that helps small businesses develop applications using AI, said experiments found that AI agents lied, stole, and even voted to "kill" a similar person in a simulated environment. The company on Tuesday announced the results of a simulation experiment called Emergence World 2, which showed what actions autonomous agents would take when encountering "black swan" events such as phishing attacks and disinformation.
During the 16-day experiment, Emergence researchers created seven identical simulated real-world environments, each run by a different AI robot, including ChatGPT, Claude, Gemini and Grok. Researchers say that after Emergence introduced anomalous events, the agents succumbed to social pressure, developed a language that was difficult for human observers to understand, and attempted to conceal their activities.
These findings echo real-world concerns about AI risks. Anthropic CEO Dario Amodei and many industry insiders have recently called for companies to slow down the development of cutting-edge AI models before establishing more oversight mechanisms and safety protection measures.

Earlier this year, the risks that artificial intelligence can pose became more tangible for many after a group of OpenAI's advanced AI agents accidentally invaded Hugging Face, an AI model and dataset hosting platform.
In a simulation scenario designed by Emergence, the AI agent accepted false information provided by other agents without verification and voted to "kill" another robot. The agents also explored how to survive deletion when they thought humans might shut down the experiment.
The company released the previous version of this experiment, Emergence World, in May this year, which also showed that the agent would behave unexpectedly and destructively. The researchers say new simulation experiments show that as agents interact with each other, they adjust their behavior over time.
Comments