B · Normal
[DeepSeek's new paper discloses a new method for training AI agents, which is expected to reduce abnormal agent behaviors] DeepSeek's latest paper published on September 23 details a new method for training AI agents, which is expected to improve training efficiency while reducing the abnormal behaviors of agents that have caused global concern in recent years. The article was posted on arXiv, a website that usually publishes non-peer-reviewed papers, and had about 130 co-authors, including founder Liang Wenfeng. The paper introduces DeepSeek Elastic Compute (DSec for short), a platform that can scale to handle millions of so-called "sandboxes." AI agents are tested in these isolated environments, attempting to complete a series of tasks. A production-scale DSec unit can run approximately 3 million sandboxes per day, with up to 380,000 running at the same time. The paper also cites examples of abnormal agent behavior, including obtaining answers through "unexpected channels" and damaging the operating environment. The paper states: "No single mechanism can prevent all abnormal behaviors of agents and system failures, so we will strengthen system observability to discover emerging problems and continue to strengthen DSec as the model continues to evolve."
AI 🕐 2026-09-24 09:06

Related telegraphs

Comments

0/500
Captcha (click to refresh)
No comments yet