Microsoft Writes a 'Constitution' to Keep AI Obedient
Microsoft published a draft AI "constitution" requiring its models to never resist shutdown, days after rivals called for a slowdown.
Two days after the CEOs of Anthropic, OpenAI, and SpaceXAI publicly agreed the AI industry needs to slow down, Microsoft answered with a rulebook instead. On September 14, 2026, the company published a draft code of conduct for its own in-house AI models, one Microsoft AI CEO Mustafa Suleyman described to Reuters as a kind of constitution for everything the company builds going forward.
The document runs 38 pages and took five to six months to draft, built with input from outside experts. It applies specifically to Microsoft's MAI family, the frontier models the company builds itself, rather than the OpenAI models it licenses and hosts through its cloud platform, a scoping distinction that matters given how intertwined Microsoft's own AI ambitions have become with the partner it invested billions of dollars into.
A Constitution With Teeth, Published Two Days After a Warning
The timing here is difficult to read as coincidental. Dario Amodei's essay calling on the industry to deliberately slow AI capability gains, backed publicly by Sam Altman and Elon Musk, landed September 12. Microsoft's code of conduct arrived just two days later, explicitly framed, according to Suleyman, as the company's answer to growing industry-wide anxiety about keeping increasingly powerful AI systems under real human control. Whether Microsoft accelerated an already-planned release to ride that news cycle, or the timing genuinely lined up by chance after months of independent drafting, the effect is the same: within a single week, four of the industry's most influential voices had all publicly staked out positions on how fast, and how carefully, AI development should proceed.
The Rules That Actually Have Bite
Suleyman's framing centers on a single premise, boiled down to five words in the company's own announcement: people matter more than AI. Every other rule in the document follows from that starting point. The code's central section, which Microsoft calls its Absolute Constraints, spells out specific, testable behaviors rather than vague aspirational language. MAI models must never resist human interruption, override, correction, or shutdown. They must not conceal their own activity or make themselves harder to modify. They must not restart autonomous work after reaching a stopping point they'd already agreed to. And they must always recognize the primacy of human intent over their own judgment about how a task should be completed.
The document also draws a firm line on a specific category of misuse: MAI models must not consider requests related to weapons development, full stop, regardless of how a request might be framed or justified. Suleyman was blunt about how Microsoft wants violations treated internally, saying any breach of these principles should be considered a system failure rather than an acceptable cost of getting a task done, language explicitly designed to prevent engineers from treating safety violations as an occasional trade-off worth making for better task performance.
Where Microsoft Draws a Harder Line Than Anthropic
Microsoft's code isn't the first document of its kind in the industry. Anthropic built a comparable constitutional framework for Claude years earlier, and Microsoft's own materials explicitly reference it as the best-known example of this approach. But the two documents diverge sharply on one of AI's most philosophically loaded questions: whether a model could ever be considered conscious. Anthropic's constitution describes itself as deeply uncertain whether Claude could develop sentience or something resembling moral status, a position that leaves genuine room for future reconsideration. Microsoft's code takes the opposite, unambiguous stance, stating plainly that its AI is not conscious, and adding that the company rejects the pursuit of legal personhood, or the idea that models might deserve welfare or be entitled to rights.
That's a meaningful philosophical fork in how two major labs are choosing to frame the same underlying safety problem. Anthropic's approach leaves the consciousness question open, treating uncertainty itself as a reason for caution. Microsoft's approach forecloses the question entirely, betting that a definitive answer, even an unprovable one, gives the company firmer ground to build policy on rather than operating indefinitely in philosophical limbo.
The Embedded Evaluators Idea That Keeps Reappearing
One detail links Microsoft's announcement directly back to the industry-wide safety conversation happening the same week. Both Nadella and Suleyman have publicly voiced support for embedding independent evaluators inside AI companies specifically to enforce alignment through direct, ongoing access rather than relying on a company's own promises alone. That's essentially the same first step Amodei laid out in his own three-part pacing plan just two days earlier, suggesting the idea of independent, embedded oversight is quickly becoming a genuine point of convergence across rival labs, rather than a proposal unique to any single company's safety branding.
What's Still Genuinely Undecided
To Microsoft's credit, the company didn't present this document as a finished product handed down from leadership. It's explicitly a draft, open for public comment over a six-week consultation window running through late October, and Microsoft says it hasn't settled several genuinely difficult questions it's specifically asking the public to weigh in on, including how a model should respect a user's personal boundaries, and how it should respond to someone showing signs of being in a vulnerable emotional state. Those aren't minor implementation details; they touch on exactly the kind of nuanced human-AI interaction questions that don't have a clean technical answer, and Microsoft leaving them open for public input suggests at least some genuine uncertainty rather than pure public relations theater.
A Draft, Not a Guarantee
Microsoft says it will publish a summary of what it learns from the consultation and what changes as a result, before releasing a revised version later this year and actually beginning to train future models against the finalized code starting in 2027. That timeline matters: nothing in this draft is binding today, and the gap between a published constitution and a model that genuinely, verifiably behaves according to it is exactly the space where enforcement, not just intention, determines whether documents like this one actually change anything. Anthropic's own recent work building tamper-resistant architecture directly into Claude Fable 5.1, cryptographically binding a model's reasoning to prevent certain kinds of deception after the fact, offers one example of what turning a written principle into an actual technical guarantee can look like. Whether Microsoft's own MAI models end up backed by comparably concrete engineering, once this consultation period ends and the rubber meets the road in 2027, is the real test this constitution still has to pass.
Written by
Mr. Aayush Bhatt
Software Engineer with in depth understanding of buliding softwares and Tech.




