cd /news/ai-safety/cloudflare-tests-waf-against-frontie… · home › topics › ai-safety › article
[ARTICLE · art-141795] src=insideai.news ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Cloudflare Tests WAF Against Frontier AI Models, Finds Gaps and Fixes Them

Cloudflare tested its Web Application Firewall against AI-generated attack payloads in an authorized customer staging environment, running 45 scenarios that produced 1,107 attempts across six attack categories, and found 49 non-blocked findings — 48 of them in command injection (CMDi) and server-side request forgery (SSRF). The exercise led to three updates in Cloudflare's Managed Ruleset, including a new SSRF - Obfuscated Host detection rule derived from requests that encoded internal addresses in non-standard numeric forms such as integer, octal and trailing-dot representations. Cloudflare ran the WAF with Attack Score blocking at 30 or below and the OWASP Core Ruleset at Paranoia Level 3, and said "The model generated requests. We decided which ones mattered," with each non-blocked request undergoing human review before any rule was deployed.

by read4 min views2 publishedSep 29, 2026
Cloudflare Tests WAF Against Frontier AI Models, Finds Gaps and Fixes Them
Image: Insideai (auto-discovered)

September 29, 2026, (Inside AI) — Cloudflare has revealed the results of an internal experiment testing its Web Application Firewall (WAF) against frontier AI models. The company built an adaptive system that uses large language models (LLMs) to generate and mutate attack payloads, aiming to see if the WAF could withstand AI-driven attacks. The test, conducted on an authorized customer staging environment, involved 1,107 attempts across six attack categories. The vast majority of attacks were blocked, but the exercise uncovered specific gaps that led to three updates in Cloudflare's Managed Ruleset.

The experiment underscores a growing concern in cybersecurity: as AI models become more capable, they can automate the discovery of vulnerabilities at a speed and scale unmatched by human hackers. Cloudflare's proactive approach aims to stay ahead of this threat by using the same technology to harden its defenses. The findings provide a rare glimpse into how AI can be leveraged both offensively and defensively in the ongoing battle to secure web applications.

Cloudflare's WAF tester works by starting with a known exploit and then iteratively changing how it is encoded or delivered. The LLM proposes variations, and a second review call analyzes the response to decide the next step. The system runs without access to WAF internals, simulating an external attacker. "The model generated requests. We decided which ones mattered," the company stated, emphasizing the role of human review in triaging results.

The test targeted six attack types: cross-site scripting (XSS), SQL injection (SQLi), command injection (CMDi), server-side request forgery (SSRF), path traversal or local file inclusion (LFI), and Log4j. The WAF was configured with Attack Score blocking scores of 30 or below, all Managed Rules enabled, and the OWASP Core Ruleset at Paranoia Level 3. After running 45 scenarios, the system generated 1,107 attempts. XSS, LFI, SQLi, and Log4j had near full coverage, meaning almost all variations were blocked. However, 49 findings emerged, with 48 belonging to CMDi and SSRF.

Read: OpenAI Releases GPT-6 Astra Model With Critical Cybersecurity Capability

One notable SSRF scenario involved a cloud metadata address. The model tried different representations of the IP, such as integer, octal, and trailing-dot forms. The WAF blocked all but one: a trailing-dot form that resulted in a redirect rather than a block. This specific case led to a new detection rule for obfuscated hosts. "The SSRF - Obfuscated Host detection came directly from requests that encoded internal addresses in non-standard numeric forms," Cloudflare explained.

Cloudflare also learned that more attempts within a single scenario did not always yield better results. Some scenarios began repeating earlier ideas near the 25-attempt limit. Broader coverage came from testing more starting requests, attack categories, and input locations. The company ran the same scenarios with two versions of the same model family, and both produced different variations but surfaced the same underlying issues. This consistency allowed for reliable comparison without treating either model's output as ground truth.

The findings were not immediately turned into rules. Each non-blocked request underwent human review to rule out false positives, malformed requests, or benign payloads. Only then were they considered for mitigation. Cloudflare grouped related findings into four sets of candidate rules, validated each, and tested them against live traffic before deployment. This careful process ensured that new rules would not disrupt legitimate traffic.

For customers, Cloudflare recommends ensuring that Managed Rules and WAF Attack Score are correctly configured. They also suggest deploying additional layers such as API Security, Bots and Fraud detection, and Threat Intelligence. Positive security controls, which define expected request shapes, can further reduce attack surface. Cloudflare advises running Managed Rules in log mode first, reviewing matches in Security Events, and confirming no impact on legitimate traffic before switching to block mode. The company also offers an Attack Signature Detection feature that simplifies reviewing matched traffic. Read: OpenAI and 100 Others Warn Window to Defend Against AI Attacks Is Narrowing

This experiment is part of Cloudflare's broader effort to integrate AI into its security development lifecycle. By combining adaptive AI-driven testing with human triage, the company aims to find detection gaps that fixed tests might miss. In a future post, Cloudflare plans to share results from a white-box approach where the model knows both the application's vulnerabilities and the WAF rules. This ongoing research highlights the evolving arms race between AI-powered attacks and defenses, with Cloudflare positioning itself at the forefront.

── more in #ai-safety 4 stories · sorted by recency
── more on @cloudflare 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cloudflare-tests-waf…] indexed:0 read:4min 2026-09-29 · —