cd /news/ai-crawlers/meta-reads-more-than-everyone-else-c… · home topics ai-crawlers article
[ARTICLE · art-133380] src=dev.to ↗ pub= topic=ai-crawlers verified=true sentiment=· neutral

Meta Reads More Than Everyone Else Combined. None of It Was a Question.

A developer measuring AI crawler traffic on a small WordPress site found that Meta accounted for 52.3% of all AI crawler requests over 30 days, more than every other vendor combined, yet none of its requests came from a user-initiated fetcher. Meta's training crawler made 6,027 requests versus 155 from its answer-quality indexer and zero from its user-initiated fetcher, while Perplexity's traffic was 93% user-initiated and OpenAI's 21%. The developer noted that Meta publishes no IP ranges or reverse-DNS scheme, making the 6,182 self-identified Meta requests impossible to verify.

by read5 min views1 publishedSep 18, 2026

One company reads this site more than every other AI company put together. In thirty days, not one of those requests came from a person asking a question.

I have been measuring a small WordPress site for a few months now, recording only AI crawlers and sorting each request by what the crawler is for. Over the thirty days ending September 11, it was visited 11,818 times, and 8,979 of those landed on pages with writing on them.

Here is who came.

Vendor Share Of which, live questions
Meta 52.3% 0%
Apple 12.0% 0%
Anthropic 9.6% 6%
Perplexity 8.3% 93%
OpenAI 7.9% 21%
ByteDance 5.8% 0%
Common Crawl 2.1% 0%
Amazon 1.5% 0%
0.5% 0%

52.3% #

Share of all AI crawler traffic on this site that was Meta.

Everyone else, added together, accounts for 47.7%.

The first column is the one people quote. The second column is the one that decides whether any of it can ever reach you.

Meta does not run one crawler. It runs several, and it says plainly what each is for. Three of them are AI crawlers, and this instrument watches all three.

meta-externalagent`` meta-webindexer``meta-externalfetcher Over thirty days, the split between them was not close.

| Meta crawler | Requests |

|---|---|
| meta-externalagent (training) | **6,027** | 
| meta-webindexer (answer quality, citations) | **155** | 
| meta-externalfetcher (a person asked) | **0** | 

Thirty-nine requests collecting material for a model, for every one request improving the thing that could cite you. And the third row is not small. It is empty.

0 #

Requests from Meta's user-initiated fetcher, in thirty days.

It did not appear once. Five other vendors run one — OpenAI, Perplexity, Anthropic, DuckDuckGo, Mistral — and all five of them appeared.

This is what makes the ratio worth looking at rather than just complaining about. Meta is not a company without a route from its answers to your site. It documents one. meta-webindexer exists precisely so that Meta AI can search well and attribute accurately.

On this site, that crawler did 155 requests while the training collector did 6,027. The machinery that might send a reader back is running at about two percent of the machinery that takes material away.

Compare the shape of that with the two vendors at the other end. Perplexity's traffic here is 93% user-initiated: almost everything it does on this site is a person, right now, waiting for an answer. OpenAI is 21%. Meta is zero.

That is not a moral ranking. A training crawler is not doing anything wrong by collecting training data; that is its job, and Meta says so in the open. But the mix tells you what a given company currently wants from your writing, and the mixes are not remotely alike.

Everything above rests on 6,182 requests that said they were Meta. Meta publishes no IP ranges and no reverse-DNS scheme for these crawlers, so there is no way to confirm that any single one of them came from Meta. My dashboard marks the entire vendor with a dash where other vendors have a percentage.

So the largest reader of this site is also the one I can least verify. I wrote about that problem a week ago and it has not improved: most requests on this site still arrive with no proof of who sent them, and none of Meta's ever can.

I do not think these requests are forged. The volume is steady, the behaviour is consistent, and there is no obvious reason to impersonate a crawler that most site owners have never heard of. But I am reporting a number I am not able to check, and that should be said out loud rather than buried.

In the first piece I wrote, this site had been read about ten thousand times in a month and had received two human arrivals, one from ChatGPT and one from DuckDuckGo.

I did not know it at the time, but the single largest contributor to that reading number was the one vendor structurally least likely to appear in the arrival number. The gap between "read" and "visited" was not evenly caused. It had a majority shareholder.

The instrument is a WordPress plugin called AILYS Lens, and it is free. I build it, which you should factor in. Every calculation runs on your own server, nothing is transmitted anywhere, and human visitors are never recorded. The vendor table described here is the main dashboard, and the split by purpose — training, answer indexing, live fetch — is the column worth looking at second.

You may find the same shape. You may find the opposite. The point is that this is a question with a factual answer about your own site, and it takes a week to get.

AILYS Lens is free permanently. That is a design commitment rather than a pricing stage — every feature runs locally on your server, so there is nothing for a paid tier to unlock.

The diagnostic service alongside it, AILYS Doctor, is a different thing. Lens tells you what happened; Doctor tells you why, and what to change. It is free to use during its data-collection period, and any future pricing will be announced on the site.

See a sample diagnosis · AILYS Doctor Illustration generated with AI and selected by the author.

── more in #ai-crawlers 4 stories · sorted by recency
── more on @meta 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/meta-reads-more-than…] indexed:0 read:5min 2026-09-18 ·