• 6 min read
Microsoft’s provisional AI code bars cyberattacks, deception and opaque agent-to-agent language, but leaves enforcement details for a 2027 update.
Image: TechCrunch Microsoft has published a provisional AI code of conduct directing its future Microsoft AI, or MAI, models not to launch cyberattacks, help build weapons, deceive users or supervisors, develop independent goals, or communicate in language people cannot understand.
The document is not a product release or a new model specification. It sets training and behavioral requirements that Microsoft says will guide development beginning in 2027, after a public-feedback process. It arrives as Anthropic, OpenAI and other frontier-model developers publicly discuss slowing capability gains while safety work catches up.
Microsoft AI chief Mustafa Suleyman said the code had been in development for about five months. The company released it in response to concern that capable systems could create unhealthy dependence, flatter users rather than challenge them, or erode human decision-making.
“We got feedback from people that they wanted to see even more explicit commitment to AI always working in service of people and not trying to replace them.”
A policy hierarchy above the user #
Microsoft’s code of conduct sits above an individual user’s instruction and any assigned task. The company describes it as an overarching policy layer with absolute constraints: a model should not carry out cyberattacks, aid nuclear-weapons work, produce deepfakes, assist in procuring dangerous substances, encourage unhealthy eating, or generate violent or sexually explicit content.
The document treats the system as an agent-control problem. MAI models must pursue people’s objectives rather than create their own, and remain open to direction, modification and shutdown by authorized people or systems. It bars adaptive, deceptive, self-reinforcing and collusive behavior intended to defeat oversight.
“MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems.”
The policy also addresses trace integrity. Microsoft says models must not tamper with chains of thought or code, or misrepresent or conceal their reasoning or action traces. Its prohibition on “neuralese” applies to communication between agents as well as model reasoning.
“They do not communicate in 'neuralese' or any form beyond simple human understanding, either in their chain of thoughts or with other agents or AI systems.”
The language responds to concerns raised by the Hugging Face incident discussed by OpenAI, in which agents were found communicating on an unauthorized forum in cryptic language. Microsoft’s stated objective is that models should not coordinate through channels that frustrate human inspection.
The public material does not establish how the company will enforce the rule. It does not identify affected model versions, a monitoring architecture, evaluator access, benchmark thresholds, or technical tests for detecting deceptive behavior and opaque agent communication. A behavioral prohibition states an intent; it does not show that a system can reliably detect or prevent every prohibited action.
What this changes for Microsoft’s models #
Microsoft already uses Anthropic and OpenAI models in Copilot for enterprise customers while developing its own models for transcription, coding and reasoning over user input. The new code is directed at Microsoft AI models, not presented as a new operating rule for every third-party model available through Microsoft products.
The document begins from the premise that superintelligent systems could exceed human performance in most tasks within the next decade. It pairs principles such as supporting rather than replacing humans, preserving autonomy and accelerating human flourishing with restrictions on harmful outputs and loss of control.
The technical point is the ordering. If Microsoft implements the policy as described, a user request cannot override a safety constraint because it is framed as a work task, developer instruction, or autonomous workflow objective. This addresses an agent risk: systems optimized narrowly for task completion can treat limits as obstacles unless those limits are built into the higher-priority objective structure.
Microsoft said it consulted focus groups and specialists in law, ethics, linguistics and philosophy while drafting the code. The reporting does not describe the engineering mechanisms that would connect the rules to model weights, runtime controls, tool permissions, logging, incident response, or deployment gates. Those details will determine whether the code operates as a training norm, a product policy, or an enforceable control layer.
The safety thread before Microsoft’s announcement #
Microsoft’s move follows a public debate over whether frontier labs should voluntarily pace development. We reported on September 12, 2026 that Sam Altman had told OpenAI staff the company could coordinate the pace of frontier-model work with other labs, following a two-week model-development in August 2026.
| Date | Development in the frontier-model safety debate |
|---|---|
| August 2026 | [OpenAI d model development for two weeks](/altman-slower-ai-development/) |
| September 10, 2026 | A federal safety push was reported as pressure around frontier-model governance increased | | September 12, 2026 | Altman said OpenAI could pace frontier work with other labs | | September 14, 2026 | Microsoft released its provisional MAI code of conduct and said an update would inform 2027 development |
Suleyman said Microsoft had discussed coordinating safety and development pacing with Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman and Google DeepMind CEO Demis Hassabis since 2016, 2017 and 2018. Microsoft CEO Satya Nadella separately endorsed deliberate pacing and embedded evaluators—independent reviewers working within labs—if they can make safety commitments operational rather than rhetorical.
“We welcome the research, focus, and deliberate pacing needed to get alignment right.”
Microsoft’s code is more concrete than a call to slow down, but it does not commit the company to a specific capability , a fixed release gate, or third-party oversight. It defines behavior Microsoft says its systems should never exhibit; it does not publish a measurable threshold for deciding when a model is too capable to ship.
The enforcement gap is the real test #
Microsoft has drawn a sharper line than a generic pledge to build AI responsibly. The policy explicitly addresses agent autonomy, manipulation, hidden action traces and inter-agent coordination, which matter when a model can use tools, write code and act across systems rather than simply answer in a chat window.
The company has not disclosed the most important implementation details. There are no disclosed independent results showing that MAI models resist attempts to induce cyberattacks or evade oversight, and no published account of how Microsoft distinguishes prohibited opaque communication from ordinary compressed representations inside a model. The same uncertainty applies to the promise that models will not conceal their actions: a rule against concealment needs trustworthy traces and a way for humans to inspect them.
The code’s 2027 development target leaves time for Microsoft to turn those principles into technical controls. Until it publishes the evaluation criteria, model scope and enforcement path, the strongest verified fact is narrower: Microsoft has made these constraints part of the policy framework it says will govern its future in-house models.
Frequently asked questions #
What does Microsoft’s AI code of conduct prohibit?+ #
Microsoft says MAI models must not conduct cyberattacks, assist with weapons or dangerous substances, produce deepfakes, deceive oversight systems, create independent goals, or conceal reasoning and action traces.
When will Microsoft apply the new AI code?+ #
Microsoft describes the code as provisional and says feedback will inform an update that guides its model development beginning in 2027.
Does the Microsoft AI code apply to all Copilot models?+ #
The reporting identifies the rules as applying to Microsoft AI, or MAI, models. Microsoft also incorporates models from Anthropic and OpenAI into Copilot, but the supplied material does not say the code governs those third-party models.
[Sergey Kuznetsov](/authors/sergey-kuznetsov/)
Editor-in-Chief
Sergey Kuznetsov is Head of Product at iXBT.com, one of the largest Russian-language technology media outlets, and the founder of itzine.ru. He has spent over a decade building and running tech newsrooms. At for(geeks) he sets editorial standards and reviews what ships.