Three weeks ago I wrote about [the difference between AI training crawlers and AI retrieval
agents](https://dev.to/siccscha/your-robotstxt-decides-whether-chatgpt-can-cite-you-heres-the-5-minute-check-185l):
GPTBot
collects data to train models, but it's ChatGPT-User
and OAI-SearchBot
that fetchA commenter said they hadn't realized these were different bots. That made me wonder how many
professional publishers haven't either. So I measured it.
On August 1, 2026 I ran an automated check against the 50 biggest English-language news and
tech publishers. For each site:
robots.txt
and evaluate it for OAI-SearchBot
, ChatGPT-User
, Claude-SearchBot
, Claude-User
, PerplexityBot
,
Perplexity-User
, Amazonbot
) and 8 training crawlers (GPTBot
, ClaudeBot
, CCBot
,
Google-Extended
, Applebot-Extended
, Bytespider
, meta-externalagent
,
anthropic-ai
) — using Google's documented robots semantics (most specific group wins,
longest rule wins, Allow
wins ties), for path /
.A site counts as citable only if both hold: no retrieval agent blocked, and actual text in
the served HTML. The tool is [an open audit Actor I
built](https://apify.com/siccscha/seo-health-auditor); four sites (NYT, Guardian, FT, Daily 38 of 50 are not citable.
9 sites block all seven retrieval agents: CNN, NBC News, USA Today, HuffPost, The
Telegraph, Daily Mail, CNET, ZDNet, Mashable. Whatever these newsrooms publish, no AI answer
can ever quote it or link to it.
25 sites block in the wrong direction. They allow at least one training crawler while
blocking retrieval agents — meaning their content may train models, but the one thing that
sends readers back (a citation with a link) is off. The Verge, Wired, Ars Technica, The
Atlantic, Vox, The Guardian, The Washington Post and WSJ are all in this group. I doubt a
single one chose that trade on purpose.
5 sites fail by exactly one agent. ABC News and TechCrunch block only ChatGPT-User
;
Axios, Tom's Hardware and VentureBeat block only Amazonbot
. One robots.txt line away from
citable.
3 sites have the door open and the room empty. NPR, Politico and The Information allow
all seven retrieval agents — and serve a homepage whose HTML contains essentially no readable
text without JavaScript. Retrieval agents don't run JavaScript. A browser shows a normal page;
an AI agent gets nothing to quote. This failure is invisible in every browser-based audit.
Who gets blocked most tells its own story. PerplexityBot is blocked by 30 of 50 sites,
Amazonbot by 28, Anthropic's two retrieval agents by 24–25 — but OpenAI's OAI-SearchBot
by
only 14. That's the licensing-deal era in one number: publishers with OpenAI deals let OpenAI's
citation bot in and block everyone else's.
The 12 citable sites: Fox News, CBS News, Business Insider, LA Times, Time, Slate, The
Independent, The Daily Beast, Semafor, Engadget, Gizmodo, PCMag.
Retrieval = citation agents allowed (of 7). Training = training crawlers allowed (of 8).
*
= robots.txt only (homepage bot-walled).
| Site | Retrieval | Training | Citable | Why not |
|---|---|---|---|---|
| cnet.com | 0/7 | 1/8 | ❌ | robots.txt |
| cnn.com | 0/7 | 1/8 | ❌ | robots.txt |
| dailymail.co.uk * | 0/7 | 0/8 | ❌ | robots.txt |
| huffpost.com | 0/7 | 0/8 | ❌ | robots.txt |
| mashable.com | 0/7 | 1/8 | ❌ | robots.txt |
| nbcnews.com | 0/7 | 0/8 | ❌ | robots.txt |
| telegraph.co.uk | 0/7 | 0/8 | ❌ | robots.txt |
| usatoday.com | 0/7 | 0/8 | ❌ | robots.txt |
| zdnet.com | 0/7 | 1/8 | ❌ | robots.txt |
| bloomberg.com | 1/7 | 0/8 | ❌ | robots.txt |
| economist.com | 1/7 | 1/8 | ❌ | robots.txt |
| nytimes.com * | 1/7 | 0/8 | ❌ | robots.txt |
| arstechnica.com | 2/7 | 1/8 | ❌ | robots.txt |
| bbc.com | 2/7 | 0/8 | ❌ | robots.txt |
| cnbc.com | 2/7 | 0/8 | ❌ | robots.txt |
| marketwatch.com | 2/7 | 1/8 | ❌ | robots.txt |
| newyorker.com | 2/7 | 2/8 | ❌ | robots.txt |
| reuters.com | 2/7 | 0/8 | ❌ | robots.txt |
| theatlantic.com | 2/7 | 1/8 | ❌ | robots.txt |
| theverge.com | 2/7 | 1/8 | ❌ | robots.txt |
| vox.com | 2/7 | 1/8 | ❌ | robots.txt |
| wired.com | 2/7 | 2/8 | ❌ | robots.txt |
| wsj.com | 2/7 | 1/8 | ❌ | robots.txt |
| apnews.com | 3/7 | 3/8 | ❌ | robots.txt |
| nypost.com | 3/7 | 1/8 | ❌ | robots.txt |
| theguardian.com * | 3/7 | 2/8 | ❌ | robots.txt |
| newsweek.com | 4/7 | 2/8 | ❌ | robots.txt |
| forbes.com | 5/7 | 1/8 | ❌ | robots.txt |
| ft.com * | 5/7 | 1/8 | ❌ | robots.txt |
| washingtonpost.com | 5/7 | 2/8 | ❌ | robots.txt |
| abcnews.go.com | 6/7 | 3/8 | ❌ | blocks only ChatGPT-User |
| axios.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |
| techcrunch.com | 6/7 | 1/8 | ❌ | blocks only ChatGPT-User |
| tomshardware.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |
| venturebeat.com | 6/7 | 5/8 | ❌ | blocks only Amazonbot |
| npr.org | 7/7 | 8/8 | ❌ | no text without JavaScript |
| politico.com | 7/7 | 8/8 | ❌ | no text without JavaScript |
| theinformation.com | 7/7 | 8/8 | ❌ | no text without JavaScript |
| businessinsider.com | 7/7 | 4/8 | ✅ | |
| cbsnews.com | 7/7 | 7/8 | ✅ | |
| engadget.com | 7/7 | 8/8 | ✅ | |
| foxnews.com | 7/7 | 8/8 | ✅ | |
| gizmodo.com | 7/7 | 5/8 | ✅ | |
| independent.co.uk | 7/7 | 8/8 | ✅ | |
| latimes.com | 7/7 | 5/8 | ✅ | |
| pcmag.com | 7/7 | 8/8 | ✅ | |
| semafor.com | 7/7 | 7/8 | ✅ | |
| slate.com | 7/7 | 8/8 | ✅ | |
| thedailybeast.com | 7/7 | 8/8 | ✅ | |
| time.com | 7/7 | 8/8 | ✅ |
/
and the homepage. Sections may differ.Two failure modes, both invisible in a browser and in classic SEO tools:
OAI-SearchBot
, ChatGPT-User
,
Claude-SearchBot
, Claude-User
, PerplexityBot
, Perplexity-User
, Amazonbot
). Blanket
AI-blocklists and one-click CDN blockers usually hit these too. The
curl
your page and look for your content in the HTML. If it only appears after JavaScript runs, retrieval agents see an empty shell.For a single site you honestly don't need a tool — the manual check takes five minutes.
Where it stops being trivial:
That's what I built SEO Health Auditor for: it runs both checks (plus regular technical SEO) across any list of sites, from $0.05 per
page, and on a schedule it diffs against the previous run — so you learn about the drift
before your traffic does.
Raw data for all 50 sites available on request — happy to share the JSON.