Can't remember where I stumbled upon a SEO-type guide that was suggesting agent-optimization via a robot.txt-like file llms.txt which lives in the web root and helps agents use your site. Perfect, I thought, my sightread.org has a zillion options, all bookmarkable. I imagined a world where someone asks ChatGPT for a sight reading exercise or maybe even a teachable gradual sequence of exercises, and said ChatGPT creates the URLs and send the person to the correct exercise. Wonderful! So I dully created my sightread.org/llms.txt
I guess something was still bothering me, mainly how do I even know it's useful and do the agent's crawlers are finding it ok. Asked Gemini back then, 6-7 months ago. It told me to link to it from my robots.txt like
Sitemap: https://sightread.org/llms.txt
Next time Gemini (3-pro) told me this is invalid and wrong and not even a sitemap. Agh!
I also monkey-see-monkey-do'd a link to it from my HTML head, like so:
<link rel="llms" href="https://www.phpied.com/llms.txt" title="LLM Guide">
7 months later #
Recently I asked Claude about opinions and it told me that no one cares. Agent crawlers ignore it. My coding agent might find it useful in the absence of agents.md, but no SEO/Agent-O at all. Ditch it, it said. Write a public facing HTML instead. And I did.
Couple days ago #
Lo and behold, Lighthouse now has an "agenting browsing" audit. I ran it, it complained about my missing llms.txt. I restored it, LH complained it has no links. My markdown had URLs but not formatted as links. Fine. Fixed. Bye. Now the HTML meta is like described by the docs:
<link rel="describedby" href="https://www.phpied.com/llms.txt" type="text/markdown">
I am ready! Come to me my agents!
Today #
Still something was bugging me. Does llms.txt matter or not at all? One company that knows crawling and has an agent and a browser and developer tools in the browser says YES! It's WRONG if it's missing. Another agentic company says meh! Who cares?
What to do but look at server logs to see who's crawling my llms.txt
First I compared number of requests to robots.txt to the number of requests to llms.txt. In terms of percentage it's 1 - 1.2% in the last 5 days. That's it? Not impressed. For every 100 requests to robots.txt I get 1 request for llms.txt.
Next - who is requesting it? Here are the top user-agents (or should we say agent-agents?)
| User-agent | % of total |
|---|---|
| Chrome 148/Linux | 66.28% |
| Chrome 126/Mac | 16.09% |
| SEOJuice-SearchBot | 5.13% |
| Chrome 144/Windows | 1.15% |
| Chrome 124/Mac | 0.81% |
| Chrome 136/Android (Chrome-Lighthouse) | 0.69% |
| Chrome 136/Mac (Chrome-Lighthouse) | 0.65% |
| ... | ... |
| Claude-User (Anthropic) | 0.10% |
| ... | ... |
There's some Claude, no GPT, no Gemini... Meta + Googlebot total of 0.08%
One UA is responsible for 66%. Running latest Chrome on Linux, not identifying itself. Just one poorly vibecoded crawler going rogue on my poor server? Maybe.
Conclusions? #
Is it me? My syntax was wrong? I dunno. I'm inclined to to say llms.txt doesn't matter, whatever Lighthouse says. Until I see googlebot/gemini in the logs, I don't trust it.
LLMs-schmellms.
Comments? Find me on
[BlueSky](https://bsky.app/profile/stoyan.org),
[Mastodon](https://indieweb.social/@stoyan),
[LinkedIn](https://www.linkedin.com/in/stoyanstefanov/),
[Threads](https://www.threads.net/@stoyanstefanov),
[Twitter](https://x.com/stoyanstefanov)