{"slug": "i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most", "title": "I checked whether ChatGPT can cite the top 50 news sites. 38 are invisible — most by accident.", "summary": "An automated audit of the 50 largest English-language news and tech publishers found that 38 are not citable by AI retrieval agents like ChatGPT, with 9 sites blocking all seven major retrieval bots and 25 blocking retrieval while allowing training crawlers. The audit, run by developer siccscha on August 1, 2026, revealed that many publishers inadvertently block citations due to robots.txt misconfigurations or JavaScript-dependent pages, while PerplexityBot is the most blocked agent and OpenAI's OAI-SearchBot is the least, reflecting licensing deals.", "body_md": "Three weeks ago I wrote about [the difference between AI training crawlers and AI retrieval\nagents](https://dev.to/siccscha/your-robotstxt-decides-whether-chatgpt-can-cite-you-heres-the-5-minute-check-185l):\n\n`GPTBot`\n\ncollects data to train models, but it's `ChatGPT-User`\n\nand `OAI-SearchBot`\n\nthat fetchA commenter said they hadn't realized these were different bots. That made me wonder how many\n\nprofessional publishers haven't either. So I measured it.\n\nOn August 1, 2026 I ran an automated check against the 50 biggest English-language news and\n\ntech publishers. For each site:\n\n`robots.txt`\n\nand evaluate it for `OAI-SearchBot`\n\n, `ChatGPT-User`\n\n, `Claude-SearchBot`\n\n, `Claude-User`\n\n, `PerplexityBot`\n\n,\n`Perplexity-User`\n\n, `Amazonbot`\n\n) and 8 training crawlers (`GPTBot`\n\n, `ClaudeBot`\n\n, `CCBot`\n\n,\n`Google-Extended`\n\n, `Applebot-Extended`\n\n, `Bytespider`\n\n, `meta-externalagent`\n\n,\n`anthropic-ai`\n\n) — using Google's documented robots semantics (most specific group wins,\nlongest rule wins, `Allow`\n\nwins ties), for path `/`\n\n.A site counts as **citable** only if both hold: no retrieval agent blocked, and actual text in\n\nthe served HTML. The tool is [an open audit Actor I\nbuilt](https://apify.com/siccscha/seo-health-auditor); four sites (NYT, Guardian, FT, Daily\n\n**38 of 50 are not citable.**\n\n**9 sites block all seven retrieval agents:** CNN, NBC News, USA Today, HuffPost, The\n\nTelegraph, Daily Mail, CNET, ZDNet, Mashable. Whatever these newsrooms publish, no AI answer\n\ncan ever quote it or link to it.\n\n**25 sites block in the wrong direction.** They allow at least one *training* crawler while\n\nblocking *retrieval* agents — meaning their content may train models, but the one thing that\n\nsends readers back (a citation with a link) is off. The Verge, Wired, Ars Technica, The\n\nAtlantic, Vox, The Guardian, The Washington Post and WSJ are all in this group. I doubt a\n\nsingle one chose that trade on purpose.\n\n**5 sites fail by exactly one agent.** ABC News and TechCrunch block only `ChatGPT-User`\n\n;\n\nAxios, Tom's Hardware and VentureBeat block only `Amazonbot`\n\n. One robots.txt line away from\n\ncitable.\n\n**3 sites have the door open and the room empty.** NPR, Politico and The Information allow\n\nall seven retrieval agents — and serve a homepage whose HTML contains essentially no readable\n\ntext without JavaScript. Retrieval agents don't run JavaScript. A browser shows a normal page;\n\nan AI agent gets nothing to quote. This failure is invisible in every browser-based audit.\n\n**Who gets blocked most tells its own story.** PerplexityBot is blocked by 30 of 50 sites,\n\nAmazonbot by 28, Anthropic's two retrieval agents by 24–25 — but OpenAI's `OAI-SearchBot`\n\nby\n\nonly 14. That's the licensing-deal era in one number: publishers with OpenAI deals let OpenAI's\n\ncitation bot in and block everyone else's.\n\n**The 12 citable sites:** Fox News, CBS News, Business Insider, LA Times, Time, Slate, The\n\nIndependent, The Daily Beast, Semafor, Engadget, Gizmodo, PCMag.\n\nRetrieval = citation agents allowed (of 7). Training = training crawlers allowed (of 8).\n\n`*`\n\n= robots.txt only (homepage bot-walled).\n\n| Site | Retrieval | Training | Citable | Why not |\n|---|---|---|---|---|\n| cnet.com | 0/7 | 1/8 | ❌ | robots.txt |\n| cnn.com | 0/7 | 1/8 | ❌ | robots.txt |\n| dailymail.co.uk * | 0/7 | 0/8 | ❌ | robots.txt |\n| huffpost.com | 0/7 | 0/8 | ❌ | robots.txt |\n| mashable.com | 0/7 | 1/8 | ❌ | robots.txt |\n| nbcnews.com | 0/7 | 0/8 | ❌ | robots.txt |\n| telegraph.co.uk | 0/7 | 0/8 | ❌ | robots.txt |\n| usatoday.com | 0/7 | 0/8 | ❌ | robots.txt |\n| zdnet.com | 0/7 | 1/8 | ❌ | robots.txt |\n| bloomberg.com | 1/7 | 0/8 | ❌ | robots.txt |\n| economist.com | 1/7 | 1/8 | ❌ | robots.txt |\n| nytimes.com * | 1/7 | 0/8 | ❌ | robots.txt |\n| arstechnica.com | 2/7 | 1/8 | ❌ | robots.txt |\n| bbc.com | 2/7 | 0/8 | ❌ | robots.txt |\n| cnbc.com | 2/7 | 0/8 | ❌ | robots.txt |\n| marketwatch.com | 2/7 | 1/8 | ❌ | robots.txt |\n| newyorker.com | 2/7 | 2/8 | ❌ | robots.txt |\n| reuters.com | 2/7 | 0/8 | ❌ | robots.txt |\n| theatlantic.com | 2/7 | 1/8 | ❌ | robots.txt |\n| theverge.com | 2/7 | 1/8 | ❌ | robots.txt |\n| vox.com | 2/7 | 1/8 | ❌ | robots.txt |\n| wired.com | 2/7 | 2/8 | ❌ | robots.txt |\n| wsj.com | 2/7 | 1/8 | ❌ | robots.txt |\n| apnews.com | 3/7 | 3/8 | ❌ | robots.txt |\n| nypost.com | 3/7 | 1/8 | ❌ | robots.txt |\n| theguardian.com * | 3/7 | 2/8 | ❌ | robots.txt |\n| newsweek.com | 4/7 | 2/8 | ❌ | robots.txt |\n| forbes.com | 5/7 | 1/8 | ❌ | robots.txt |\n| ft.com * | 5/7 | 1/8 | ❌ | robots.txt |\n| washingtonpost.com | 5/7 | 2/8 | ❌ | robots.txt |\n| abcnews.go.com | 6/7 | 3/8 | ❌ | blocks only ChatGPT-User |\n| axios.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |\n| techcrunch.com | 6/7 | 1/8 | ❌ | blocks only ChatGPT-User |\n| tomshardware.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |\n| venturebeat.com | 6/7 | 5/8 | ❌ | blocks only Amazonbot |\n| npr.org | 7/7 | 8/8 | ❌ | no text without JavaScript |\n| politico.com | 7/7 | 8/8 | ❌ | no text without JavaScript |\n| theinformation.com | 7/7 | 8/8 | ❌ | no text without JavaScript |\n| businessinsider.com | 7/7 | 4/8 | ✅ | |\n| cbsnews.com | 7/7 | 7/8 | ✅ | |\n| engadget.com | 7/7 | 8/8 | ✅ | |\n| foxnews.com | 7/7 | 8/8 | ✅ | |\n| gizmodo.com | 7/7 | 5/8 | ✅ | |\n| independent.co.uk | 7/7 | 8/8 | ✅ | |\n| latimes.com | 7/7 | 5/8 | ✅ | |\n| pcmag.com | 7/7 | 8/8 | ✅ | |\n| semafor.com | 7/7 | 7/8 | ✅ | |\n| slate.com | 7/7 | 8/8 | ✅ | |\n| thedailybeast.com | 7/7 | 8/8 | ✅ | |\n| time.com | 7/7 | 8/8 | ✅ |\n\n`/`\n\nand the homepage. Sections may differ.Two failure modes, both invisible in a browser and in classic SEO tools:\n\n`OAI-SearchBot`\n\n, `ChatGPT-User`\n\n,\n`Claude-SearchBot`\n\n, `Claude-User`\n\n, `PerplexityBot`\n\n, `Perplexity-User`\n\n, `Amazonbot`\n\n). Blanket\nAI-blocklists and one-click CDN blockers usually hit these too. The\n`curl`\n\nyour page and look for your content in the HTML. If it\nonly appears after JavaScript runs, retrieval agents see an empty shell.For a single site you honestly don't need a tool — the manual check takes five minutes.\n\nWhere it stops being trivial:\n\nThat's what I built [SEO Health Auditor](https://apify.com/siccscha/seo-health-auditor) for:\n\nit runs both checks (plus regular technical SEO) across any list of sites, from $0.05 per\n\npage, and on a schedule it diffs against the previous run — so you learn about the drift\n\nbefore your traffic does.\n\n*Raw data for all 50 sites available on request — happy to share the JSON.*", "url": "https://wpnews.pro/news/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most", "canonical_source": "https://dev.to/siccscha/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most-by-accident-55kh", "published_at": "2026-08-03 20:09:32+00:00", "updated_at": "2026-08-03 20:43:36.696356+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-policy", "ai-tools", "ai-infrastructure"], "entities": ["OpenAI", "Perplexity", "Anthropic", "CNN", "NBC News", "The Verge", "Wired", "Ars Technica"], "alternates": {"html": "https://wpnews.pro/news/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most", "markdown": "https://wpnews.pro/news/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most.md", "text": "https://wpnews.pro/news/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most.txt", "jsonld": "https://wpnews.pro/news/i-checked-whether-chatgpt-can-cite-the-top-50-news-sites-38-are-invisible-most.jsonld"}}