cd /news/artificial-intelligence/how-film-industry-data-website-the-n… · home topics artificial-intelligence article
[ARTICLE · art-73979] src=hackaday.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

How Film Industry Data Website The-Numbers.com got Mauled by Bots

The Numbers, a film industry data website founded in 1997, was forced offline for months after a surge of automated bot traffic, primarily from LLM-related crawlers and prediction markets like Polymarket, overwhelmed its server architecture and drove up hosting costs. Founder and CEO Bruce Nash told Stephen Follows that despite mitigations such as LLM-targeted text, the site collapsed under the load, requiring a major architectural rework to restore features. The incident highlights the growing challenge of AI-driven scraping, which services like Cloudflare now offer blocking features to counter.

read2 min views1 publishedJul 26, 2026
How Film Industry Data Website The-Numbers.com got Mauled by Bots
Image: Hackaday (auto-discovered)

A lot has been made about the increase of automated traffic on the Internet, with the past years LLM-related crawlers having quite literally swarmed the picture here. Not only does this drive up traffic, it also increases load on web servers, whose owners find themselves faced with increased hosting costs. This recently led to The-Numbers.com going offline for a while as automated traffic was quite literally destroying their bottom line.

This saga is covered by [Stephen Follows], who had a chance to talk with the founder and CEO of the site, [Bruce Nash], after the site went basically offline for a few months. Since the website both licenses data for commercial purposes as well as offering the free access on its website, there were accusations of this being a ‘rug pull’.

The site was started in 1997, as a static HTML site on Geocities where [Bruce] provided box office analyses for investment purposes. Since that beginning traffic was generally polite, with human visitors and usually well-behaved search engine crawlers. Then around 2024 the first wave of scraper bots arrived, followed by a larger wave around December of 2025.

Despite implementing a few mitigations, such as LLM-targeted text, the increased traffic and the resulting load on a site architecture that was never designed for this ultimately led to a collapse. One of the major sources of traffic turned out to be from so-called ‘prediction markets’, like Polymarket, whose bots absolutely hammered the site.

Fortunately for [Bruce] and his team they do not rely on the free website for income, but they have had to massively rework the site’s architecture to bring back a semblance of the original features. As noted in the article, the amount of crawling traffic by these LLMs and ‘agentic AI’ tools is logarithmically more than that for search engines, which makes this a major challenge.

Issues like these is why services such as Cloudflare are offering blocking features for such automated traffic. After all, unless such traffic is of use to you, you may as well treat it like a DDoS attack and cut it off at the root.

Thanks to [Ben] for the tip.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @the numbers 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-film-industry-da…] indexed:0 read:2min 2026-07-26 ·