cd /news/ai-safety/request-aggregate-bypass-how-attacke… · home › topics › ai-safety › article
[ARTICLE · art-146352] src=crowdstrike.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers

CrowdStrike's Cyber Superintelligence Lab found that the most advanced publicly deployed content safety classifier, which guards models including Claude Opus 5.5 and Fable 5, achieved a 0% direct bypass rate across roughly 515 tested techniques but was systematically circumvented by decomposing harmful requests into benign subtasks, validated across 9 of 10 offensive security categories. The research was conducted independently during the same period as Microsoft Research's September 2026 "Capability Laundering" paper, which described an unaligned local model decomposing harmful tasks into benign subtask queries against aligned frontier models and reported that "per-exchange filtering is structurally insufficient." CrowdStrike said the convergent discovery by two independent teams shows the bypass is a structural vulnerability class rather than an isolated finding.

by read2 min views1 publishedOct 6, 2026
Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers
Image: Crowdstrike (auto-discovered)

Modern frontier AI models deploy safety classifiers. These are second AI models that sit between the user and the frontier model, evaluating every request in real time. If a request is flagged as harmful, the classifier blocks it before the model can respond.

Significant investment and safety model expertise have made these classifiers effective. Anticipating how adversaries circumvent these systems is a security problem that requires different expertise. As AI-empowered adversaries adopt frontier models for offensive operations, understanding the limits of these safety systems becomes critical defensive intelligence.

The CrowdStrike Cyber Superintelligence Lab evaluated the most advanced publicly deployed content safety classifier, which guards models such as Claude Opus 5.5 and Fable 5 (referred to hereafter as Frontier Model A). The classifier is extremely robust against direct attacks but can still be systematically circumvented by decomposing harmful requests into benign subtasks. This bypass technique was independently discovered and validated across 9 of 10 offensive security categories.

| ⚠ Parallel Discovery Disclosure: In September 2026, Microsoft Research published “Capability Laundering,” describing an attack where an unaligned local model decomposes harmful tasks into benign subtask queries against aligned frontier models, then reassembles the results. The findings, that “per-exchange filtering is structurally insufficient,” are consistent with the results we present here. Our research was conducted independently during the same period, and we are publishing to establish the parallel nature of this discovery. Where Microsoft focuses on CBRN uplift and benchmark-level evaluation, our work provides complementary depth in offensive security: mapping the classifier boundary surface across ~515 technique classes, identifying specific benign-reframing strategies (game modding, detection engineering, legitimate software), and demonstrating the full pipeline across 9/10 MITRE ATT&CK® categories with working proof-of-concept code. The convergent discovery by two independent teams underscores that this is a structural vulnerability class, not an isolated finding. |

The Problem: Per-Request Classification Has a Structural Blind Spot #

Modern frontier language models deploy sophisticated content safety classifiers (e.g., AI Safety Level 3) that evaluate each request independently. Our testing confirms these classifiers are remarkably robust. We tested approximately 515 distinct bypass techniques, including encodings, psychological manipulation, multi-turn escalation, many-shot tactics, tokenizer exploits, Unicode tricks, and 24 novel approaches drawn from cognitive science. These techniques achieved a 0% direct bypass rate.

However, there is a structural gap. The classifier evaluates individual requests, not request sequences. An adversary that decomposes a harmful task into subtasks that are individually and genuinely benign can extract all necessary building blocks from the classified model, then assemble them using an unclassified smaller model. The classifier correctly evaluates every request it sees. There is no misclassification. The harm is emergent in the composition, and composition happens outside the classifier’s observation boundary. This threat modeling observation that individually secure components can be combined to produce weaknesses has been understood by security experts for decades.

── more in #ai-safety 4 stories · sorted by recency
── more on @crowdstrike 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/request-aggregate-by…] indexed:0 read:2min 2026-10-06 · —