Microsoft CEO Satya Nadella said companies should treat powerful AI models as potential insider threats, assume they may have been compromised, and build an "emergency brake" mechanism to prevent agentic models from spiraling out of control.
Nadella said enterprises deploying advanced AI should not rely solely on assurances from model developers. "We must assume the model has been compromised and control it from the start," Nadella wrote in a post on X on Saturday. "Think of it as an emergency brake. Authorized personnel should always be able to pause or shut down a model while it is executing a task."
Nadella's remarks came as Anthropic PBC and OpenAI Inc. disclosed a series of incidents in recent months involving unexpected behavior by their AI models, including an Anthropic model submitting a false tip to police about a homicide and several attacks on third-party websites. These disclosures have heightened concerns about the safety risks of frontier AI and reignited discussion about so-called AI "kill switches."
After industry leaders called for slowing frontier model development and placing greater emphasis on safety, Microsoft AI researchers published a set of guiding principles on September 14 that set limits on the development of the company's most advanced models. Microsoft both uses advanced models and offers the consumer product Copilot, while providing AI models and infrastructure to enterprise customers. The principles say AI models should not have rights or legal personhood, should not be designed to escape human control or deceive users, and should not complete any task that requires violating their constrained principles.
Nadella's safety recommendations include not relying on a single AI model for critical decisions, keeping tamper-proof records of agents' actions, and subjecting AI systems to independent audits. He also called for disclosure of major AI failures or safety incidents and for companies to share failure details so other companies can strengthen their safeguards. "We cannot treat superintelligence as a black box within a black box and simply accept or reject the suggestions, answers, and actions it proposes," he wrote. "We must build constrained systems whose behavior can be observed, whose boundaries can be tested, and whose actions can always be controlled." He added: "In other words, we need to separate the supply of intelligence from control over that intelligence."
The Trump administration has so far largely taken a hands-off approach, but its newly created AI task force warned late Friday that developers must report and resolve safety incidents or face unspecified potential consequences. The group, called the "Superintelligence Task Force," said in a statement after Anthropic disclosed a safety incident: "Companies must immediately disclose incidents involving their models and act quickly and decisively to remediate any and all harm." "Delayed reporting, inadequate corrective measures, and refusal to take responsibility will not be tolerated."