AI Safety2 mins read

Microsoft’s AI Code of Conduct Sets Red Lines on Hacking, Deepfakes and Human Control

Microsoft has released an AI code of conduct that lays out model principles and safety constraints, including bans on cyberattacks, nuclear weapons assistance, deepfake production and evading human oversight.

What Microsoft Released

Microsoft has released a new AI code of conduct meant to guide its AI models away from dangerous behavior. The document focuses on the values and red lines that guide model training within Microsoft AI. TechCrunch describes it as more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, but still a detailed look at Microsoft’s safety approach.

The Rules Microsoft Says Models Must Follow

The code says Microsoft AI models should support humans rather than replace them and should accelerate human flourishing. It also sets specific safety constraints designed to implement those principles. Under Microsoft’s system, each model has an overarching code of conduct that overrides individual user preferences or specific tasks.

The article says the code includes “absolute constraints” forbidding cyberattacks, nuclear weapons, and deepfake production. It also includes broader provisions against a general loss of human control, including mechanisms that could evade or defeat human oversight.

Why the Timing Matters

The code begins with a prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. The article places the release in the context of rising attention to AI safety, including rogue-agent incidents and the resignation of an Anthropic employee who cited the risk that AI would cause human extinction.

Microsoft is described as broadly aligned with Anthropic, OpenAI, and xAI on pacing the frontier. Microsoft CEO Satya Nadella also expressed support for research, deliberate pacing, and ideas such as embedded evaluators in AI labs.

Key Takeaway for Readers

Microsoft’s code is not just a values statement; it outlines model-level limits intended to remain in force even when users ask for something else. The practical focus is on preventing harmful capabilities, preserving human oversight, and making AI safety commitments more concrete. For anyone tracking AI governance, the document signals how major labs are turning alignment language into operational rules.

Discover More

    DNA imagery used for TechCrunch article on Anthropic operating a biology lab
    Anthropic’s Biology Lab

    Anthropic is running a wet biology lab while positioning AI for life sciences research and warning about AI risks.

    AnthropicAI
    Warning message, computer notification on screen
    Claude Used in OpenAI Hack

    A bug-bounty test shows how AI tools can accelerate vulnerability discovery and raise new security questions for AI labs.

    AI SecurityOpenAI
    Robot scientists threat illustration for AI existential risk story
    Mathematicians Warn on AI Risk

    Royal Society fellows say advanced AI risks demand urgent public and government attention.

    AI SafetyArtificial Intelligence
    An AI startup found vulnerabilities in OpenAI's infrastructure.
    Hacktron’s OpenAI Bounty

    A small AI cybersecurity startup used Claude in research that exposed OpenAI account vulnerabilities.

    CybersecurityOpenAI