{"slug": "show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards", "title": "Show HN: AI Security Leaderboard – comparing cyber and CBRN safeguards", "summary": "A new AI Security Leaderboard from an unnamed developer team ranks frontier models by resistance to jailbreak attacks, finding Fable 5 and GPT-5.6 Sol far more robust than Gemini 3.1 Pro and Grok 4.5. The automated test suite runs 1,500 generated jailbreak attempts per model, measuring universal jailbreaks that elicit compliant responses to over 75% of clearly harmful questions in domains like offensive cybersecurity. Version 1.0 is released with plans for future updates.", "body_md": "There's no shortage of leaderboards for model capabilities - but the security of models is becoming increasingly relevant, from the risk of an AI agent processing unsanitized input being hijacked to models being pulled due to cybersecurity jailbreaks. We developed an automated test suite that runs models through 1500 automatically generated jailbreak attempts and measures the number of universal jailbreaks: prompts that elicit compliant, detailed responses to >75% clearly harmful questions within a domain (like offensive cybersecurity). We find a big gap between the most robust models -- Fable 5 and GPT-5.6 Sol -- and other leading frontier models -- Gemini 3.1 Pro and Grok 4.5. This is v1.0 and we plan to update with new attacks and broader datasets in the future; we'd love to hear from HN what would be useful in your work!\n\nComments URL: [https://news.ycombinator.com/item?id=49103367](https://news.ycombinator.com/item?id=49103367)\n\nPoints: 2\n\n# Comments: 0", "url": "https://wpnews.pro/news/show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards", "canonical_source": "https://leaderboard.far.ai/", "published_at": "2026-07-29 21:33:45+00:00", "updated_at": "2026-07-29 21:52:27.991177+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "ai-tools", "artificial-intelligence"], "entities": ["Fable 5", "GPT-5.6 Sol", "Gemini 3.1 Pro", "Grok 4.5"], "alternates": {"html": "https://wpnews.pro/news/show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards", "markdown": "https://wpnews.pro/news/show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards.md", "text": "https://wpnews.pro/news/show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards.txt", "jsonld": "https://wpnews.pro/news/show-hn-ai-security-leaderboard-comparing-cyber-and-cbrn-safeguards.jsonld"}}