Microsoft published a new AI code of conduct in September 2026, establishing the values and behavioral limits that govern its AI models — including absolute prohibitions on cyberattacks, nuclear weapons assistance, and deepfake production.
The document sets out a two-tier framework: broad principles that Microsoft AI models should uphold, and specific safety constraints that implement those principles. Among the guiding principles are supporting humans rather than replacing them and accelerating human flourishing. The safety constraints go further, explicitly forbidding models from using “adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.”
Under the system Microsoft describes, each model operates under an overarching code of conduct that takes precedence over instructions from individual users or specific tasks. The document also opens with a prediction that superintelligent AI systems will surpass human performance in most tasks within the next decade, framing alignment as “one of the greatest challenges humanity has ever faced.”
Microsoft CEO Satya Nadella expressed support for the broader industry push toward careful AI development, including the concept of embedded evaluators in AI labs. “We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal,” Nadella wrote.
The release comes amid heightened focus on AI safety, driven by a series of rogue-agent incidents and the resignation of an Anthropic employee who cited the growing risk that AI could cause human extinction. Microsoft, alongside Anthropic, OpenAI, and xAI, has broadly aligned with an approach of pacing frontier AI development.
The code of conduct is described as more operationally focused than recent public statements from Anthropic CEO Dario Amodei calling for frontier pacing, offering a concrete look at how Microsoft translates safety principles into model training practices.
Source: TechCrunch