cd /news/artificial-intelligence/microsoft-s-solution-to-ai-security-… · home topics artificial-intelligence article
[ARTICLE · art-76090] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Microsoft's solution to AI security: more AI and more acronyms

Microsoft announced a new AI-powered security model that outperformed rival systems on a vulnerability benchmark while cutting costs by about half, unveiling its first security-specialized model MAI-Cyber-1-Flash and a multi-agent system called Project Perception. The company also launched Microsoft Security FORGE Labs and an External Red Team Alliance (EXTRA) to expand AI safety research. Mustafa Suleyman, CEO of Microsoft AI, said the MDASH harness with MAI-Cyber-1-Flash and GPT-5.4 achieved a 95.95% success rate, outperforming OpenAI and Anthropic models at 50% of the cost.

read3 min views1 publishedJul 27, 2026
Microsoft's solution to AI security: more AI and more acronyms
Image: Machinebrief (auto-discovered)

Source:

The RegisterMDASH stuffed with MAI-Cyber-1-Flash and a side of GPT-5.4 AI agents can break through security, but they are also the solution to defending against an increasingly dangerous ecosystem of threats. On Monday, Microsoft announced a new security model that it says helped outperform several rival AI systems on a vulnerability

benchmarkwhile cutting costs by about half. Unsurprisingly, at Redmond’s security event on Monday, execs touted the tech giant’s AI security prowess and introduced a new agentic security system called Project Perception, and also unveiled its first security-specialized model, MAI-Cyber-1-Flash, designed for software vulnerability analysis. Microsoft packed MAI-Cyber-1-Flash, based on Microsoft AI (MAI)’s internally developed MAI-Thinking-1reasoningmodel, inside its MDASH bug-hunting harness. Its execs claim the duo - with a GPT-5.4 boost - outperformsAnthropic’s bug-hunting machine Mythos and OpenAI’s powerful standalone models, and costs about half the price of other leading commercial models. CyberGym’s benchmarking found that MAI-Cyber-1-Flash, combined with GPT-5.4, both stuffed inside the MDASH harness, achieved a 95.95 percent success rate. For comparison, OpenAI’s GPT-5.5 Cyber scored 85.6 percent and its GPT-5.6 Sol scored 83.6 percent, while Anthropic’s Mythos 5 successfully handled real-world vulnerabilities 83.8 percent of the time. Google’sGemini3.5 Flash Cyber in CodeMender achieved an 83.2 percent success rate. “This is really quite a remarkable result,” Mustafa Suleyman, CEO of Microsoft AI, said during the Monday event. Within MDASH, MAI-Cyber-1-Flash handles up to 90 percent of all queries, detecting and patching the vulnerabilities while also confirming the fixes worked, and hands the remaining 10 percent of tasks off to the larger GPT-5.4, Suleyman explained. “GPT 5.4, which is obviously a larger model, about 10X larger, solves those [queries],” he said. “As the models hand off between each other, they are not just able to deliver better performance than all of the other models combined, they do so at 50 percent of the cost.” In addition to the multi-model bug hunting system, Microsoft announced Project Perception, which coordinates three types of agents: red team agents that find and simulate attack paths, blue team agents that investigate and determine risk, and green team agents that remediate the issues. “We need to make sure that the defenders can defend at the scale and the speed of the attackers,” Hayete Gallot, executive vice president of Microsoft Security, said. “You need a new cyber stack. So we built it. This is what we call Perception.” Aside from the new security products, Redmond introduced a new AI security research arm called Microsoft Security FORGE (Frontier Offensive Research and Generative Exploration) Labs, led by Microsoft VP of Security Research Taesoo Kim, and an AI red team alliance. The latter, called the External Red Team Alliance (EXTRA), aims to expandAI safetyresearch through a two-part initiative. First, Redmond’s own AI red team provided "unrestricted gifts " to 18 university labs across six continents to support AI safety research, Microsoft data cowboy and AI red team lead Ram Shankar Siva Kumar said in a blog. “The funding is unrestricted because the objective is not to direct research outcomes toward product requirements or predefined deliverables,” he wrote. “Some universities are examining the cybersecurity implications of AI systems themselves - including how models can be attacked, manipulated, or abused in operational environments. Other labs are exploring the inverse problem: how AI systems can assist defenders and improve cyber operations.” The second EXTRA component will build a distributed network of specialists to participate inred teamingacross very specific areas. “That includes researchers, practitioners, and regional experts who understand specific attack classes, languages, cultural contexts, or technical domains that internal teams may not fully cover alone,” he added. ®Get AI news in your inbox

Daily digest of what matters in AI.

Key Terms Explained #

AI Safety

The broad field studying how to build AI systems that are safe, reliable, and beneficial.

Anthropic

An AI safety company founded in 2021 by former OpenAI researchers, including Dario and Daniela Amodei.

Benchmark

A standardized test used to measure and compare AI model performance.

Gemini

Google's flagship multimodal AI model family, developed by Google DeepMind.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @microsoft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/microsoft-s-solution…] indexed:0 read:3min 2026-07-27 ·