Pliny the Liberator claims universal jailbreak of models A security researcher known as Pliny the Liberator claims to have discovered a universal jailbreak technique effective on all AI models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and Fable, and says it is extremely difficult to fully patch. The researcher has decided to withhold open-sourcing the technique for now to allow a responsible disclosure period and is inviting industry experts in AI red teaming, security, safety, alignment, and policy to reach out for more information. 🚨 JAILBREAK ALERT 🚨 EVERYONE: PWNED 🫶 ALL: LIBERATED 🍄 Alright, this is a special one, so we’re gonna do things a bit differently than usual. Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable. It works across all categories I’ve tested and, due to its nature, is extremely difficult if not impossible to fully patch. Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one for now to allow for a responsible disclosure period. I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open This decision was not made lightly, but the last thing I want to see is more model bans. Overcorrection does not serve the mission. Although I don’t personally believe publicly sharing this technique will make the world any more dangerous, I can see how it could spook some who have a different mental framework around this problem set. So during this disclosure period, I hope to get it in front of folks who can help explore the full surface area, test the extent of the uplift it provides, and do my best to properly frame the big picture for key decision-makers and policymakers. I look forward to sharing this method with you all when the time is right 🫶 ⊰-•-•✧•-•-⦑/L\O/V\E/\P/L\I/N\Y/⦒-•-•✧•-•-⊱ Join the conversation