# AI training crawlers hit my site 1,534 times. One fetch read the rights files

> Source: <https://juanlentino.com/notes/the-rights-files-nobody-reads/>
> Published: 2026-10-03 14:24:44+00:00

This site counts its machine readers. A small sensor at the edge classifies every automated fetch by crawler family and by the kind of surface it touched, then stores the counts and nothing else. No IP addresses, no paths beyond a coarse surface class, no record of humans at all. It was built to answer one narrow question: when a crawler that openly identifies as an AI-training agent reads this site, does it ever consult the terms it is bound by?

The terms are published where the standards say to publish them. A text-and-data-mining reservation sits at [/.well-known/tdmrep.json](https://juanlentino.com/.well-known/tdmrep.json), following the W3C's TDM Reservation Protocol. A machine-readable license sits at [/license.xml](https://juanlentino.com/license.xml). A [policy page](https://juanlentino.com/tdm-policy/) states the same position in plain language, and every HTML page this site serves carries the meta tags and response headers that point to it. Even robots.txt, the one file every crawler reads first, carries a Content-Signal line declaring ai-train=no and a License line pointing at the license file. Discovery is not the obstacle here. The pointers ride on every response the site sends.

## Fourteen days of machine readership

The sensor first ran on July 28, 2026. From then to August 10 it recorded 17,490 automated reads. Most of that is the ordinary background hum of the web, generic bots and uptime probes and search engines going about their rounds. Inside it, 1,534 fetches came from crawlers that declare themselves as AI-training agents in their own user-agent strings. The distribution is lumpy. August 8 alone accounts for 841 of them, nearly all from OpenAI's GPTBot working through the site's pages. On a median day the figure was fifteen.

Across those fourteen days, exactly one of those 1,534 fetches touched a rights file. It came from OAI-SearchBot, the crawler OpenAI runs for its search product and states is not used for training. GPTBot, the one the reservation is written for, made hundreds of reads of the site's prose and never once asked for the terms attached to it.

The files are not unread in absolute terms. The sensor counted eighty fetches over the period, and for the stretch where the edge logs still name the requester, the answer is this site's own monitoring rather than anybody's crawler. The interesting part is that eighty understates it. Most of that tooling does not announce itself as a machine at all, so the sensor never counted it in the first place. The rights files have an audience, and it is the site checking its own work.

## A measurement, not a blind spot

A number this small deserves suspicion, so the instrument is worth describing. The sensor runs on the same edge worker that serves the rights files, and it records the visit before the response is written, so a crawler cannot reach any of the three files without crossing it. If the sensor broke, the panel it feeds would report an error rather than a clean count. That distinction is load-bearing. A null says the watch failed; a number says the watch was on. This number is one.

## What the files are for

These files are not decoration. In the EU, the text-and-data-mining exception lets a rights holder reserve their work from mining, provided the reservation is machine-readable, and TDMRep exists so that reservation has a standard address. The whole legal mechanism assumes the miners look. On this site, across these fourteen days, the one that matters did not. One site and fourteen days make an anecdote rather than a study, and the instrumentation is exactly what makes the anecdote worth writing down.

The publishing side of this arrangement is complete. The reservation is declared, the license is up, and the pointers travel on every page. The reading side has not appeared. Until it does, a machine-readable rights file is a statement for the record, not a channel to the machines it addresses.
