Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
CrowdStrike's Cyber Superintelligence Lab found that the most advanced publicly deployed content safety classifier, which guards models including Claude Opus 5.5 and Fable 5, achieved a 0% direct bypa…