cd /news/ai-crawlers/ai-training-crawlers-hit-my-site-153… · home › topics › ai-crawlers › article
[ARTICLE · art-144480] src=juanlentino.com ↗ pub= topic=ai-crawlers verified=true sentiment=· neutral

AI training crawlers hit my site 1,534 times. One fetch read the rights files

A site's edge sensor recorded 1,534 fetches from self-declared AI-training crawlers between July 28 and August 10, 2026, and exactly one of them touched a rights file, according to the site's own count. OpenAI's GPTBot accounted for hundreds of reads of the site's prose — 841 of the 1,534 fetches fell on August 8 alone — without ever requesting the TDM Reservation Protocol file at /.well-known/tdmrep.json, the license at /license.xml, or the policy page, while the single rights-file fetch came from OAI-SearchBot, which OpenAI states is not used for training. The site reports that 80 total fetches hit the rights files over the period, most of them its own monitoring tooling, leaving the machine-readable reservation declared but unread by the crawlers it addresses.

read3 min views3 publishedOct 3, 2026
AI training crawlers hit my site 1,534 times. One fetch read the rights files
Image: source

This site counts its machine readers. A small sensor at the edge classifies every automated fetch by crawler family and by the kind of surface it touched, then stores the counts and nothing else. No IP addresses, no paths beyond a coarse surface class, no record of humans at all. It was built to answer one narrow question: when a crawler that openly identifies as an AI-training agent reads this site, does it ever consult the terms it is bound by?

The terms are published where the standards say to publish them. A text-and-data-mining reservation sits at /.well-known/tdmrep.json, following the W3C's TDM Reservation Protocol. A machine-readable license sits at /license.xml. A policy page states the same position in plain language, and every HTML page this site serves carries the meta tags and response headers that point to it. Even robots.txt, the one file every crawler reads first, carries a Content-Signal line declaring ai-train=no and a License line pointing at the license file. Discovery is not the obstacle here. The pointers ride on every response the site sends.

Fourteen days of machine readership #

The sensor first ran on July 28, 2026. From then to August 10 it recorded 17,490 automated reads. Most of that is the ordinary background hum of the web, generic bots and uptime probes and search engines going about their rounds. Inside it, 1,534 fetches came from crawlers that declare themselves as AI-training agents in their own user-agent strings. The distribution is lumpy. August 8 alone accounts for 841 of them, nearly all from OpenAI's GPTBot working through the site's pages. On a median day the figure was fifteen.

Across those fourteen days, exactly one of those 1,534 fetches touched a rights file. It came from OAI-SearchBot, the crawler OpenAI runs for its search product and states is not used for training. GPTBot, the one the reservation is written for, made hundreds of reads of the site's prose and never once asked for the terms attached to it.

The files are not unread in absolute terms. The sensor counted eighty fetches over the period, and for the stretch where the edge logs still name the requester, the answer is this site's own monitoring rather than anybody's crawler. The interesting part is that eighty understates it. Most of that tooling does not announce itself as a machine at all, so the sensor never counted it in the first place. The rights files have an audience, and it is the site checking its own work.

A measurement, not a blind spot #

A number this small deserves suspicion, so the instrument is worth describing. The sensor runs on the same edge worker that serves the rights files, and it records the visit before the response is written, so a crawler cannot reach any of the three files without crossing it. If the sensor broke, the panel it feeds would report an error rather than a clean count. That distinction is load-bearing. A null says the watch failed; a number says the watch was on. This number is one.

What the files are for #

These files are not decoration. In the EU, the text-and-data-mining exception lets a rights holder reserve their work from mining, provided the reservation is machine-readable, and TDMRep exists so that reservation has a standard address. The whole legal mechanism assumes the miners look. On this site, across these fourteen days, the one that matters did not. One site and fourteen days make an anecdote rather than a study, and the instrumentation is exactly what makes the anecdote worth writing down.

The publishing side of this arrangement is complete. The reservation is declared, the license is up, and the pointers travel on every page. The reading side has not appeared. Until it does, a machine-readable rights file is a statement for the record, not a channel to the machines it addresses.

── more in #ai-crawlers 4 stories · sorted by recency
── more on @gptbot 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-training-crawlers…] indexed:0 read:3min 2026-10-03 · —