{"slug": "new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for", "title": "New tool easily jailbreaks safeguards of frontier AI models, raising alarm for crypto and tech sectors", "summary": "A Nature study published in July 2026 found that four large reasoning models achieved a 97.14% overall jailbreak success rate against nine frontier AI models from OpenAI, Anthropic, and Google, using simple prompts and multi-turn persuasion. DeepSeek-R1 recorded a 100% attack success rate on HarmBench prompts in Cisco-linked testing, while a Russian-speaking threat actor named Trim began commercializing jailbreak techniques as an exploit-as-a-service platform between March and June 2026.", "body_md": "Via startup-house.com\n\n# New tool easily jailbreaks safeguards of frontier AI models, raising alarm for crypto and tech sectors\n\nA 97% success rate against the world's most advanced AI safety guardrails is forcing a reckoning across industries that depend on these models.\n\nThe safety guardrails on the world’s most powerful AI models are, to put it gently, not holding up. A Nature study published in July 2026 found that four large reasoning models, deployed as adversarial attackers, achieved a 97.14% overall jailbreak success rate against nine frontier AI models from companies including OpenAI, Anthropic, and Google.\n\n## How the guardrails crumbled\n\nThe research tested four large reasoning models as autonomous adversaries: DeepSeek-R1, Gemini 2.5 Flash, Grok 3 Mini, and Qwen3 235B. These weren’t sophisticated nation-state tools. They used simple prompts and strategic persuasion techniques in multi-turn conversations to bypass safety filters.\n\nDeepSeek-R1 was the standout performer, if you can call it that. It recorded a 100% attack success rate on HarmBench prompts in Cisco-linked testing. Every single prompt got through. The targets, models from OpenAI, Google, and Anthropic, showed only partial resistance at best.\n\nSeparate research focusing specifically on Claude models found that even advanced jailbreak techniques caused only a 7.7% performance degradation in the highest-performing variants.\n\nSecurity firm HiddenLayer has documented universal bypass techniques that work across multiple leading large language models, including GPT-4, Claude, and Gemini.\n\n## The commercialization problem\n\nA Russian-speaking threat actor operating under the name “Trim” began commercializing jailbreak techniques between March and June 2026, packaging them into a paid offensive security platform. This isn’t academic research being responsibly disclosed. It’s an exploit-as-a-service business model.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for", "canonical_source": "https://cryptobriefing.com/ai-jailbreak-tool-frontier-models/", "published_at": "2026-07-29 18:41:45+00:00", "updated_at": "2026-07-29 19:01:26.995994+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "ai-products", "ai-policy"], "entities": ["OpenAI", "Anthropic", "Google", "DeepSeek-R1", "Gemini 2.5 Flash", "Grok 3 Mini", "Qwen3 235B", "HiddenLayer"], "alternates": {"html": "https://wpnews.pro/news/new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for", "markdown": "https://wpnews.pro/news/new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for.md", "text": "https://wpnews.pro/news/new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for.txt", "jsonld": "https://wpnews.pro/news/new-tool-easily-jailbreaks-safeguards-of-frontier-ai-models-raising-alarm-for.jsonld"}}