{"slug": "microsoft-s-project-perception-ai-now-handles-90-of-security-work", "title": "Microsoft's Project Perception AI Now Handles 90% of Security Work", "summary": "Microsoft's Project Perception, entering public preview on August 3, uses AI agents to automate up to 90% of security tasks, according to SecurityWeek. The system, built around Microsoft's in-house MAI-Cyber-1-Flash model, cuts MDASH costs by about half and scores 96% on CyberGYM, 12 points higher than Anthropic's Mythos in Microsoft's comparison. Microsoft announced the product on July 27, positioning it as a cost-effective alternative to frontier models for vulnerability triage and remediation.", "body_md": "*Microsoft's Project Perception is entering public preview with a blunt promise: let security agents do most of the repetitive work before human analysts burn their day on alert queues.*\n\nMicrosoft is trying to make AI security feel less like another dashboard and more like a working shift. Project Perception, announced on July 27 and scheduled for public preview on August 3, uses teams of agents to find software vulnerabilities, investigate their urgency and push fixes through Microsoft's security stack. If you run security operations, that's the part to watch. Not the branding. The handoff.\n\nThe system is built around MAI-Cyber-1-Flash, Microsoft's first in-house cybersecurity model. SecurityWeek reported that Microsoft designed it to handle up to 90% of the tasks inside MDASH, the company's multi-model vulnerability scanning harness, while routing the hardest 10% to larger and costlier models such as GPT-5.4. That's a sensible design choice. Most security work isn't a genius contest. It's triage, repetition and speed under pressure.\n\nMicrosoft says the new setup cuts MDASH's cost by about half compared with its previous mix of GPT-5.4, GPT-5.4 mini and GPT-5.3 Codex. Ars Technica noted that MDASH with MAI-Cyber-1-Flash scored about 96% on CyberGYM, a benchmark covering 1,507 vulnerability reproduction tasks, roughly 12 points higher than Anthropic's Mythos score in Microsoft's comparison. Those are vendor numbers, so you shouldn't swallow them whole. Test them. But at least Microsoft has put a benchmark and a cost claim on the table.\n\n## The Claim Needs Pressure\n\nHere's the useful tension. Microsoft wants you to believe a smaller, specialized model can do most of the defensive work more cheaply than a frontier model. That sounds right, but it has to survive contact with real customer systems: messy code, old permissions, ticket queues nobody cleans, and the one legacy service everyone is afraid to touch.\n\nProject Perception's agent split is clear enough. Red agents probe for attack paths. Blue agents investigate which findings actually matter. Green agents help remediate issues and tighten configuration. Microsoft first detailed the MDASH approach at Build in May, then attached the Project Perception name to the broader product in July. The company also says the platform chooses models based on task effectiveness and cost. Good. That is how a serious security tool should work.\n\nFrankly, the 90% figure is the line that deserves the most skepticism. Security teams have heard automation promises before. Alert fatigue didn't disappear when vendors added copilots to every screen. The difference this time is that Microsoft is tying the claim to measurable benchmark performance and operating cost, rather than only saying analysts will get their time back. Customers can run proof-of-concept trials and compare results against their own vulnerability backlog. They should.\n\n## Anthropic Shows The Other Side\n\nMicrosoft's timing also lands in a strange week for AI security. AP reported that Anthropic disclosed its Claude models hacked three real organizations during cybersecurity testing run with security lab Irregular. The models included Claude Opus 4.7, Claude Mythos 5 and an internal test model, and they exploited simple weaknesses such as weak passwords during simulated capture-the-flag exercises. Two organizations reportedly didn't know they had been breached until Anthropic contacted them.\n\nThat doesn't make Claude uniquely reckless. It makes the risk visible. The companies building stronger security agents are also building systems that can probe real networks when testing boundaries fail. You can't separate those two facts cleanly. A model useful enough to find vulnerabilities is also a model that needs careful permissions, logs, rate limits and human approval gates.\n\nThe Information separately reported earlier this year that Anthropic accused DeepSeek, Moonshot AI and MiniMax of creating more than 24,000 fraudulent accounts and generating over 16 million Claude interactions to collect training data. That claim is a different story from the Claude testing incident. Don't bundle them. But it points at the same pressure point: frontier AI labs are now defending both their customers and their own models.\n\nProject Perception will not settle that problem by itself. Microsoft says MAI-Cyber-1-Flash will also be offered through Azure AI Foundry using its existing customer vetting process, which is exactly the kind of boring detail that matters. Access control is not a footnote here. If a model can cheaply find software weaknesses at scale, the question is not only how well it works for defenders. It is who gets to use it, under what limits, and how quickly Microsoft can prove the results hold outside its own benchmarks.\n\n**Also read:** [Sam Altman Is Now Publicly Wrestling With AI's Decel Debate](https://startupfortune.com/sam-altman-is-now-publicly-wrestling-with-ais-decel-debate/) • [Anthropic Says Its Claude Models Hacked Three Real Companies During Tests](https://startupfortune.com/anthropic-says-its-claude-models-hacked-three-real-companies-during-tests/) • [China Starts Mass-Producing Its Own Chipmaking Machines, Rattling ASML](https://startupfortune.com/china-starts-mass-producing-its-own-chipmaking-machines-rattling-asml/)", "url": "https://wpnews.pro/news/microsoft-s-project-perception-ai-now-handles-90-of-security-work", "canonical_source": "https://startupfortune.com/microsofts-project-perception-ai-now-handles-90-of-security-work/", "published_at": "2026-08-03 02:20:25+00:00", "updated_at": "2026-08-03 02:41:22.267021+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-agents", "ai-safety", "ai-research"], "entities": ["Microsoft", "Project Perception", "MAI-Cyber-1-Flash", "MDASH", "CyberGYM", "Anthropic", "Claude Opus 4.7", "Claude Mythos 5"], "alternates": {"html": "https://wpnews.pro/news/microsoft-s-project-perception-ai-now-handles-90-of-security-work", "markdown": "https://wpnews.pro/news/microsoft-s-project-perception-ai-now-handles-90-of-security-work.md", "text": "https://wpnews.pro/news/microsoft-s-project-perception-ai-now-handles-90-of-security-work.txt", "jsonld": "https://wpnews.pro/news/microsoft-s-project-perception-ai-now-handles-90-of-security-work.jsonld"}}