AI Red Teaming: From Checkbox to Evidence
AI red teaming must shift from a checkbox exercise to providing verifiable evidence of security testing, according to a guide on AI workflow deployment. Buyers now demand actual test vectors, guardrai…
AI red teaming must shift from a checkbox exercise to providing verifiable evidence of security testing, according to a guide on AI workflow deployment. Buyers now demand actual test vectors, guardrai…
Pliny the Liberator claims a 'universal jailbreak' capable of bypassing restrictions across multiple large language models, including Claude, GPT, and Gemini. If the technique holds across different a…
A pseudonymous hacker known as Pliny the Liberator claims to have broken every major AI model at once, including GPT-5.6 Sol Ultra, Opus 5, and Fable, with a universal jailbreak announced July 24, tho…
A security researcher known as Pliny the Liberator claims to have discovered a universal jailbreak technique effective on all AI models, including heavily guardrailed flagships like Opus 5, GPT-5.6 So…
Developer Fernando Irarrázaval's AI assistant Fiu survived over 6,000 prompt injection attempts from more than 2,000 attackers without leaking its secrets.env file, though the experiment triggered a G…
Researchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell found that large language models often ignore explicit role tags like <system> or <user> and instead infer roles from text tone, enabling…
The U.S. government ordered Anthropic to disable its Fable 5 and Mythos 5 AI models over national security concerns after a hacker bypassed safety guardrails, marking the first time a Western governme…
On June 12, 2026, the U.S. Department of Commerce ordered Anthropic to suspend access to its advanced AI models Claude Fable 5 and Mythos 5 for all foreign nationals, citing national security export c…