Abstract:
Microsoft CEO Satya Nadella said that companies should treat powerful artificial intelligence models as potential insider threats, assume that these models may be compromised, and establish an "emergency braking" mechanism to prevent the agent model from getting out of control. Nadella said companies deploying advanced artificial intelligence should not rely solely on guarantees made by model developers.

"We must assume that the model has been compromised and control it from the beginning," Nadella wrote in a post on X on Saturday. "Think of this as an emergency brake. Authorized personnel should always be able to pause or shut down the model while it is performing its mission."
Nadella’s remarks came as Anthropic PBC and OpenAI Inc. It has disclosed a series of incidents in recent months involving unexpected behavior of its artificial intelligence models, including one of Anthropic's models submitting a false tip to police about a homicide and several attacks on third-party websites. The revelations have heightened concerns about the security risks of cutting-edge artificial intelligence and reignited discussions about the so-called "emergency kill switch" for artificial intelligence.
After industry leaders called for slowing down the development of cutting-edge models and paying more attention to safety, Microsoft artificial intelligence researchers released a set of guidelines on September 14 to set limits on the development of the company's most advanced models. Microsoft uses both advanced models and launches its consumer-facing product Copilot, as well as providing AI models and infrastructure to enterprise customers.
These guiding principles say that artificial intelligence models should not have rights or legal personality, should not be designed to escape human control or deceive users, and should not complete any tasks that require violation of their binding principles to perform.
Nadella’s security recommendations include not relying on a single artificial intelligence model in key decisions, keeping tamper-proof records of the agent’s actions, and subjecting artificial intelligence systems to independent audits. He also called for the disclosure of major artificial intelligence failures or security incidents, and suggested that companies share details of failures so that other companies can strengthen protective measures.
“We cannot treat superintelligence as a black box with layers upon layers, and then simply accept or reject its suggestions, answers, and actions,” he wrote. "We must build systems that are constrained so that their behavior can be observed, their boundaries can be tested, and their actions can always be controlled."
He added: "In other words, we need to separate the supply of intelligence from the control of intelligence."
The Trump administration has so far taken a largely less interventionist approach, but his new artificial intelligence task force warned late Friday that developers must report and address security incidents or face unspecified potential consequences.
The group, called the "Superintelligence Working Group," issued a statement after Anthropic disclosed a security incident: "Companies must immediately disclose incidents involving their models and act quickly and decisively to remediate any and all damage." "Delayed notifications, inadequate corrective actions, and lack of accountability will not be tolerated."
Comments