There's no shortage of leaderboards for model capabilities - but the security of models is becoming increasingly relevant, from the risk of an AI agent processing unsanitized input being hijacked to models being pulled due to cybersecurity jailbreaks. We developed an automated test suite that runs models through 1500 automatically generated jailbreak attempts and measures the number of universal jailbreaks: prompts that elicit compliant, detailed responses to >75% clearly harmful questions within a domain (like offensive cybersecurity). We find a big gap between the most robust models -- Fable 5 and GPT-5.6 Sol -- and other leading frontier models -- Gemini 3.1 Pro and Grok 4.5. This is v1.0 and we plan to update with new attacks and broader datasets in the future; we'd love to hear from HN what would be useful in your work!
Comments URL: [https://news.ycombinator.com/item?id=49103367](https://news.ycombinator.com/item?id=49103367)
Points: 2