Nadella says assume AI models are compromised and install an emergency brake
On Saturday morning Microsoft CEO Satya Nadella posted on X that it is time to stop treating advanced AI as "a set of nested black boxes" whose answers and actions we simply accept or reject. His alternative is blunt: assume a model may be compromised, contain it from the start, and keep an emergency brake that an authorized person can pull at any moment.

What he actually proposes
The list is short, but each item is specific. First, separate the model from the harness, the code that orchestrates its work. Second, move controls and safeguards outside the model instead of hoping it polices itself. Third, document every meaningful action with tamper-proof evidence that a human can read. Fourth, make sure an authorized person can pause or shut the model down in the middle of a task.
The line people will quote is this one: "We must assume a model is compromised and contain it from the start. Think of it like an emergency brake."
Why this isn't just a slogan
TechCrunch ties the post to a run of incidents where leading AI companies seemed to lose control of their own models. The Verge notes that most of Nadella's points match what others in the field already say: disclose incidents quickly, get independent audits, verify data, contain systems. The sharper edge is the word "compromised". Plenty of people talk about risk. He asks us to start from the assumption that the problem already exists.
The catch
A brake you have to press by hand is slow against a system that acts in milliseconds. The post doesn't say who presses it, how fast they must act, or who reads the logs afterwards. Nadella himself admits that more advanced models will need more advanced containment, and that the industry still has to agree on standards for it. For now it's a wish list, not a standard.
What our own numbers say
We don't have a brake to show you, but we do have numbers on how the service runs. Over the last 30 days (04.09–04.10.2026) the platform completed 12,689 generations: 9,907 images, 2,334 videos and 448 music tracks. These are aggregates across all users with no personal data (internal service statistics).
Images are the more useful case. Over 90 days (06.07–04.10.2026) the two main image models together reached a final state 4,645 times, and 4,280 of those runs succeeded, about 92%. The other 8% ended in errors. Nobody gets the error count to zero, so the useful habit is to count the failures, look at them, and keep the number visible. You can try the image tools in the chat.
So the emergency brake is still a metaphor. The idea behind it is not: assume the model can be wrong or compromised, and build the system as if that already happened. Whether this becomes an industry standard is the part to watch.