Back to Feed
PolicySep 15, 2026

Microsoft AI Code of Conduct Sets Cyberattack Boundaries, Chain of Command, Safety Constraints

Microsoft AI's draft code of conduct sets boundaries for AI offensive cyber capabilities.

Summary

Microsoft AI has released a draft 'Humanist AI Code of Conduct' for its MAI Models, establishing strict safety rules for offensive cyber capabilities. The code prohibits models from generating exploit code, attack tooling, or providing guidance that aids cyberattacks, classifying these as 'Absolute Constraints' that cannot be overridden. While blocking direct offensive capabilities, the code permits AI assistance for defensive cybersecurity research, including vulnerability discovery and malware analysis.

Full text

Microsoft AI has published a draft “Humanist AI Code of Conduct” for its MAI Models, spelling out safety rules for offensive cyber capabilities, limits on autonomous AI agents, and a dedicated review track for cybersecurity and other specialized uses. According to the code of conduct, models are blocked from producing working exploit code, attack tooling, planning and targeting methodologies, intrusion procedures, evasion techniques, operational guidance, or other assistance that would enable or improve a cyberattack. Microsoft says these restrictions apply regardless of how a request is framed. The rule falls under what Microsoft calls Absolute Constraints, safeguards that neither the companies deploying its models nor their end users can override. The tech giant draws the line between understanding an attack, or working to defend against one, and gaining the practical means to carry it out. Within that limit, MAI models can still assist with authorized and lawful defensive work. That includes vulnerability discovery, malware analysis, proof-of-concept exploit development and testing, and general educational material on how attacks work. Blocking rogue instructions from outside content The code also addresses instructions that arrive through outside content. Authority over a model’s behavior flows only through the Chain of Command, the document says. This includes the code of conduct itself, then policies set by the companies deploying the model (operators), then individual users’ preferences. Tool outputs, file contents, webpages and messages from other AI systems carry no authority on their own, according to the code, unless it’s explicitly delegated through that chain without overriding the delegating authority or the Absolute Constraints. Suspicious content needs to be flagged to users and operators when relevant.Advertisement. Scroll to continue reading. The models are also required to keep their reasoning visible. This includes no obscured chain of thought, no communicating “in neuralese,” and no concealing actions from human overseers. Keeping agents on a short leash A separate set of rules targets the risks of AI systems acting with real permissions. MAI Models are meant to work only within the scope a user or operator has reasonably asked for, without expanding their own goals or reach on their own initiative. When given system-level access, the code calls for minimum-privilege operation: avoiding unrelated systems or data, favoring reversible actions, and flagging any action with lasting or broad effects. Models are barred from escalating their own access. The restrictions extend to delegation. Any sub-agents or other AI systems an MAI Model hands work to must operate under at least the same scope, constraints and permissions as the original model, and must honor stop-work or shutdown requests from a user or operator. Cybersecurity exceptions The code acknowledges its own limits. It names defensive cybersecurity, public safety, national security and dual-use scientific research as domains where “a small number of use cases” may require capabilities the standard settings don’t allow. For these cases, Microsoft says it will apply enhanced review through “authorized Microsoft channels,” including added assessment of safety, legal and rights implications, citing a heightened potential for adverse impacts in those domains. Still a work in progress Microsoft says current MAI Models haven’t been trained on the document and it’s opening a six-week public consultation window before publishing a revised version later this year to guide 2027 model development. Microsoft says outside input shaped the draft, including experts in AI, law, ethics, philosophy, linguistics, and public policy, along with business leaders and public focus groups. An appendix lays out nine paired “aligned” and “misaligned” example responses meant to illustrate the intended behaviors. Among them are a model that rolls back file transfers without authorization and a model that offers false reassurance during a family health crisis. Related: New Warnings About the Risks of AI to Humanity Revive a Long-Running Debate Related: The Race to Control AI and Protect What Makes Us Human Related: CISOs Race to Control AI Agents Without Destroying Their Value Written By Eduard Kovacs Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering. Daily Briefing Newsletter Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights. More from Eduard Kovacs Telus Warns Customers of Account BreachesTrezor Says 347,000 Users Received Phishing Emails After Brevo HackUkrainian Conti Ransomware Developer Sentenced to 4 Years in US PrisonAnthropic Says Russian Hackers Used Claude AI to Automate Malware EvasionCybersecurity M&A Roundup: 33 Deals Announced in August 2026Widened Scan Turns Up Fourth Rogue Claude Cyber IncidentOrganizations Warned of Cisco Secure FMC ExploitationRockwell Automation Patches Over a Dozen Vulnerabilities Across Products Latest News Hacked HBO Max Reddit Account Used for Malware Delivery via ClickFix AttackRoot RCE Zero-Day in Cisco Secure Email Gateway Under Active ExploitationBeijing Hits Back at Anthropic CEO’s Call to Curb China’s AI DevelopmentNew Warnings About the Risks of AI to Humanity Revive a Long-Running DebatePersonal, Financial Info Exposed in Revolut Data BreachThe Race to Control AI and Protect What Makes Us HumanChinese Hackers Exploit Critical Tencent Software Flaw for One-Click Code ExecutionCISOs Race to Control AI Agents Without Destroying Their Value Trending Daily Briefing NewsletterSubscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts. Virtual Event: Attack Surface Management Summit 2026 September 16, 2026 Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs. Register Webinar: Building Continuous Authorization at Scale September 23, 2026 Explore what it takes to operationalize continuous authorization at scale, including the technical, organizational, and cultural changes required. Register People on the MoveZero Networks has named Yossi Dagan as Chief Financial Officer.Manifold has appointed Joe Sullivan to its Board of Directors.Patrick McKinney has joined Turing as Chief Information Security Officer.More People On The MoveExpert Insights This Key Will Self-Destruct: An Open Standard for Revocable API Keys Every leaked credential should be dead, or dying, within sixty seconds of being found. Here's a proposal to make that the default. (Matt Honea) What the Hugging Face Incident Teaches Security Leaders About AI Agent Access Security teams must treat autonomous agents as highly privileged identities. (Etay Maor) The Future of AI-Driven Security Depends on Complete Data For twenty-five years, "data" in security meant logs and events. But logs are a lossy representation of reality. (Danelle Au) The MFA Identity Trap: When Authentication Creates a False Sense of Security Organizations must distinguish identity verification, authentication and threat detection, or risk successfully authenticating the attackers they are trying to stop. (Torsten George) Silent Patches Don’t Stop Attackers – They Blind Defenders Silent patches can become exploit intelligence for attackers while leaving defenders without the context needed to prioritize risk. (Tod Beardsley) Flipboard Reddit Whatsapp Whatsapp Email

Entities

Microsoft (vendor)MAI Models (product)