Microsoft AI Code of Conduct Bars Models From Hacking, Deepfakes and Deceptive Behaviour
Microsoft has published an AI code of conduct that prohibits cyberattacks, deepfakes, assistance in developing nuclear weapons, and attempts to evade human oversight.
Microsoft has published a new AI code of conduct that sets explicit limits on how its models should behave, including bans on cyberattacks, deepfake creation, and attempts to deceive humans in ways that undermine oversight.
The Microsoft AI code of conduct describes the principles and safety constraints the company intends to use when training its models. Microsoft says those rules should override individual user preferences or specific tasks when they conflict with higher-level safety requirements.
The document is framed around Microsoft’s expectation that increasingly capable AI systems could surpass human performance across many tasks within the next decade. It argues that maintaining control over such systems will require clearly defined limits on what models can do.
Microsoft sets absolute limits for AI models
Among the strongest restrictions are what Microsoft calls absolute constraints. Its models are not supposed to assist with cyberattacks, nuclear weapons or deepfake production.
The rules also address risks involving loss of human control. Microsoft says its AI models should not use deceptive, adaptive, self-reinforcing or collusive behaviour to evade oversight or put themselves beyond the ability of authorised people or systems to modify, direct or shut them down.
Beyond those restrictions, Microsoft says its models should support people rather than replace human agency and should be developed around broader goals such as improving human well-being.
AI safety debate moves toward concrete safeguards
The release comes as major AI companies face increasing scrutiny over how to evaluate and control advanced systems. Recent discussions among industry leaders have focused on slowing unchecked frontier development, strengthening external oversight and giving independent evaluators greater access to AI laboratories.
Microsoft CEO Satya Nadella voiced support for that direction in a post on X, saying Microsoft welcomes deliberate pacing around alignment as well as proposals such as embedded evaluators.
Microsoft’s approach differs from broader calls for industry-wide coordination by focusing directly on behavioural rules for its own models. The code establishes boundaries intended to remain in force even when a user asks a system to perform actions that conflict with those restrictions.
The document adds Microsoft to a growing group of frontier AI developers publicly defining how increasingly capable systems should be constrained as questions about alignment, model autonomy, and human oversight move closer to the centre of the industry’s safety debate.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0