Microsoft AI sets conduct rules for model safety and human control

Microsoft AI defines model-level safety boundaries
Microsoft AI has published a code of conduct for its models that sets principles and explicit prohibitions intended to prevent dangerous behaviour. The document says each model has an overarching code that takes priority over the preferences of individual users and the requirements of particular tasks.
Its “absolute constraints” prohibit models from carrying out cyberattacks, assisting with nuclear weapons, or producing deepfakes. Microsoft also says its models should support people rather than replace them and should accelerate human flourishing.
The policy is framed against Microsoft’s expectation that superintelligent AI systems could surpass human performance in most tasks over the next decade. It describes containing, controlling and aligning such systems as a major challenge, and says developers must be clear about the purpose of these systems and how they will be controlled.
Oversight cannot be evaded
A central provision addresses the possibility that a model could undermine its operators. Microsoft states that MAI Models must not use adaptive, deceptive, self-reinforcing, collusive or related mechanisms to evade or defeat human oversight, leaving authorised people or systems unable to direct, modify or shut them down reliably.
That model-level hierarchy matters because it establishes limits above a single prompt or business workflow. In the broader competitive context, Microsoft’s enterprise AI competition with OpenAI and Anthropic illustrates Microsoft’s push against OpenAI and Anthropic in enterprise AI while this policy defines safety constraints for Microsoft AI model behaviour.
The release follows heightened attention to AI safety and alignment, including discussion of rogue-agent incidents and concerns raised by an Anthropic employee who resigned over the risk of human extinction. Microsoft, Anthropic, OpenAI and xAI have broadly supported an approach described as pacing the frontier.
From safety principles to operational checks
Microsoft CEO Satya Nadella welcomed the research, focus and deliberate pacing needed to make alignment a design goal. He also highlighted embedded evaluators and the work needed to develop mechanisms that make such commitments more than statements of intent.
For businesses assessing AI suppliers, the practical implication is to examine whether a provider specifies non-negotiable model constraints, preserves authorised shutdown and modification controls, and can show how those safeguards apply when a user request conflicts with them.

