The bots already won the front door TollBit's State of the Bots report reveals that unauthorized AI crawlers are increasingly difficult to block, with bad bots masquerading as legitimate ones like Google's, rotating IP addresses, and using residential device networks, according to People Inc. Chief Innovation Officer Jonathan Roberts. The report suggests the media industry's focus should shift from blocking scrapers to controlling how content is retrieved and surfaced, as retrieval, not training, is the real threat to media business models. The bots https://www.fastcompany.com/section/chatbots are winning. That’s my clearest takeaway after reading the first section of the latest State of the Bots report https://tollbit.com/state-of-the-bots/q1-q2-2026/ from TollBit, which builds payment rails between publishers https://www.fastcompany.com/91504520/publishers-are-finally-getting-serious-about-ai-scraping and AI crawlers https://www.fastcompany.com/91589354/ai-crawlers-hammered-a-volunteer-run-lgbt-history-archive-2 . In it, People Inc.’s Chief Innovation Officer, Jonathan Roberts https://mediacopilot.substack.com/p/dotdash-merediths-bold-bet-on-aiand , outlines the company’s approach to AI bots scraping content https://www.fastcompany.com/91539092/ai-scraping-become-own-media-business from its many media properties: aggressively block unauthorized bots while allowing access to legitimate crawlers https://www.fastcompany.com/91572651/dont-block-the-bots-build-the-gate , which typically means some kind of licensing agreement. However, Roberts concedes that the unauthorized crawlers have become harder and harder to identify and block. The report shows evidence that “bad” bots sometimes try to masquerade as legit bots like Google’s, they often rotate IP address if their first scrape is blocked, and some industrial-scale scraping companies have resorted to using huge networks of devices in people’s homes to make it look like their traffic https://www.fastcompany.com/91579817/dead-internet-theory-is-real-web-agents is coming from real people. Even People Inc. hasn’t been entirely successful at blocking it all. So if a large media company with lots of resources and deep expertise on AI https://www.fastcompany.com/section/artificial-intelligence search crawlers can’t keep all the bad bots out, what hope is there for the rest of us? That’s why the fight over bot access can’t be where the future of publishing in the AI era is defined. If it is, the media has already lost. Instead, the focus needs to shift from what content is scraped to how that content is used. While it would be unwise for publishers to simply allow all crawlers unfettered access to their content, they should start thinking harder about how that content surfaces for the end user, and what can be done there to preserve value, encourage and enforce good behavior, and ultimately build their business. Let’s be clear: This is about retrieval, not training data. The media-AI fight has largely moved on from training, not because it was somehow “OK” for AI companies to use crawled information to train their models, but because the use case that truly threatens media business models is information retrieval—people using AI as a discovery surface. That requires accurate and up-to-date information, which is not what training is about. Training large language models is something very few companies actually do, mostly because it’s expensive. Training runs can cost in the hundreds of millions or billions, and it’s mostly about leveraging vast data sets to teach models to predict better, not for informational queries. This was clear from the early days of AI: you’d ask a chatbot about, say, the Enlightenment, and its answers were directionally good but often confused and wrong when you got into details like names and dates. For accurate information, the AI needs to retrieve it in real time. That’s a different kind of bot, with different stakes, since it’s about serving information to a specific user, not tossing it into a pile of data. None of this is to say publishers shouldn’t block training bots. They absolutely should, but they should also drop any expectation they’ll get paid for training data. Few publishers have the scale to make their corpus valuable enough to AI companies, and licensing deals have largely moved their focus away from training to retrieval, something Rob Kelly of the Media and the Machine Substack has observed. Kelly’s tally of 94 publicly announced deals https://mediaandthemachine.substack.com/p/ai-content-licensing-fewer-deals found that only about four in 10 now include training rights, and that the market is shifting from “buy content to build better models” toward “license content to deliver better answers.” A recent story in Digiday https://digiday.com/marketing/cmos-are-struggling-to-link-ai-visibility-with-sales/ zeroed in on the value of appearing in AI answers. Although the story focused on the struggles that brands are having to connect the dots between AI presence and good business outcomes, it echoes what publishers have felt for a long time: There is some value in appearing as the authoritative source https://www.fastcompany.com/91463716/ai-isnt-stealing-your-traffic-stealing-your-authority when answering a question, but it isn’t, in and of itself, monetizable. That may be beginning to change https://www.fastcompany.com/91584067/next-ad-market-built-machines , but for the most part when an AI uses a publisher’s content in an answer, that’s the end of the journey. Or more precisely, the journey never begins. Study after study finds that the vast majority of AI users never click through to sources from AI answers: Pew Research https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/ clocked clicks on links inside a Google AI summary at about 1% of visits, compared with 15% on a results page with no AI answer on it at all. Even though the content supplier sees no financial upside, there is undoubtedly value to the user in getting the information. That’s what all the current approaches at a business model seek to quantify, whether it’s pay-per-crawl https://developers.cloudflare.com/ai-crawl-control/features/pay-per-crawl/what-is-pay-per-crawl/ , pay-per-use https://techcrunch.com/2026/07/01/cloudflares-new-policy-pushes-ai-companies-to-pay-for-publishers-content/ , licensing https://www.niemanlab.org/2026/05/the-emerging-ai-content-licensing-market-puts-news-publishers-in-a-double-bind-a-new-report-warns/ , or serving ads to bots https://mediacopilot.ai/time-ads-ai-agents-markdown/ . None of those approaches has taken off to become an industry standard, and a big part of why is the difficulty in enforcing them. At the end of the day, there are just too many paths for content to find its way into the ecosystem, whether it’s stealthy bots crawling information they shouldn’t, AI companies purchasing data from gray-market scraper companies, or honest retrieval of republished and repackaged content. But what if the focus was less on keeping content from being scraped, and more on whether or not that content appears in the answer? Let’s imagine how policing answers would work in an ideal world. A person inputs a query, and the answer engine goes and finds the content to use. We know that attribution works well in retrieval systems. All AI engines today give citations, and ProRata’s entire business model https://prorata.ai/content/fairattribution/ relies on accurately showing which sources contributed to an answer and how much. So what if the engine, after checking sources, then performed another check to ensure it had legitimate access to all those sources, whether through licensing or any of the other business models on the table? If the content doesn’t pass the check, it can’t be used. In the case of company-level licensing, it’s an easy check. In the case of pay-per-use/crawl models, the service might build the average spend of that into its fees. Or the user might allocate a budget for it. A good analogy is how evidence is adjudicated in courts. When evidence is obtained improperly, it can’t be used in court, even when it’s true. Process matters. But with online content, the presumption runs the other way. Crawlers tend to default to a stance along the lines of, “It was out there, it was reachable, so it must be free.” As bots evolve beyond anyone’s ability to keep up with the ever-expanding game of whack-a-mole, the answer layer becomes the obvious leverage point. Most of the scrapers out there Common Crawl, Parallel, Diffbot and the like don’t operate major AI engines with market share. The places people actually get information are a short list of very large companies. You can’t chase every scraper, but you can write rules for the handful of surfaces where the content is served to a human. The natural pricing model that follows: each retrieval is its own customer. Just because ChatGPT retrieved information from your site twice a day for two different users, those are essentially proxies for those individual users, not for ChatGPT itself. In other words, ChatGPT can’t just pay for a single subscription to retrieve content for everybody. This is Sam Altman’s micropayments idea https://www.niemanlab.org/2026/05/sam-altman-backs-micropayment-model-for-ai-agents-to-compensate-publishers/ arriving through the back door, and it’s where Cloudflare is already moving https://www.fastcompany.com/91572651/dont-block-the-bots-build-the-gate . None of this happens voluntarily. The companies that would need to run the check are the same ones that benefit from skipping it, and publishers have virtually no leverage to make them do otherwise. The only party that does, really, is the government, which is why some form of regulation feels inevitable here. New York’s Stealth Crawler Prohibition Act https://www.dataguidance.com/news/new-york-stealth-crawler-prohibition-act-passes , which would require bots to identify themselves, is a reasonable first step. It’s also a low bar, and the fact that it took a law to clear it says a lot about where we are. Roberts points out that the past year has produced around 30 “Napsters of content,” but no Spotify. He’s right, and the reason is that nobody has built the part where using someone’s work improperly actually costs something. That check doesn’t belong at the crawler, where publishers keep losing. It belongs at the answer, where a handful of companies decide what billions of people see. Publishers have spent three years defending the door. It’s time to start making noise about the other end of the pipe.