{"slug": "ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history", "title": "AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive", "summary": "AI crawlers from Meta and Alibaba nearly took down the LGBT History Project, a volunteer-run archive founded by Jonathan Harborne, after one Meta crawler made roughly 26,000 requests in a day and pulled about 12GB of data, forcing Harborne to double his server size and pay higher costs. Harborne, who runs the site from London, said the crawlers ignored robots.txt and that he had to learn Cloudflare and bot-blocking rules to protect the site, which has over 50 million views since 2011 and is archived by the British Library. Cloudflare estimates that 30% of all web activity is now bot traffic, and Harborne warned that AI companies risk damaging the human-created information ecosystem they rely on.", "body_md": "Jonathan Harborne founded his LGBT history encyclopedia 15 years ago to document some of the struggles his peers had encountered in their lives. He worried that even as changing social values made life better for LGBTQ+ people, the struggles past generations had endured for recognition and respect could be wiped from memory, along with the pain caused by the AIDS pandemic.\n\nThe site, [the LGBT History Project](https://lgbthistoryuk.org/index.php?title=Main_Page), recently passed 50 million views since its launch in 2011 and has been archived by the British Library for posterity. But an onslaught of [AI](https://www.fastcompany.com/section/artificial-intelligence) bots seeking to scrape its content nearly took it offline, bringing the site to a crawl while also making it more expensive to run.\n\nHarborne only realized what was happening when he began modernizing the site and its hosting, migrating it from a platform that had been in place since the project’s founding. He moved it onto a new AWS server, expecting it to become faster and more secure. Instead, the opposite happened.\n\nHarborne, who lives in London, connected the server to Claude Code and asked it to help diagnose the problem. Its analysis of his server logs pointed to an enormous volume of automated traffic. On one day, Harborne says, a crawler belonging to Meta made roughly 26,000 requests, trawling not only published articles but years of editing histories, login pages, and other parts of the site. He says it [pulled around 12GB of data](https://medium.com/towards-artificial-intelligence/metas-ai-keeps-scraping-my-community-website-and-i-get-charged-25-a-month-for-the-privilege-500e55f24755).\n\nThat matters because Harborne pays for the project himself and was forced to double the size of the server just to accommodate the scraping. “I was paying the cost for people like Meta to train their AIs,” he says. “there’s me that’s paying for all this stuff out of my back pocket”.\n\nMeta wasn’t the only one. Harborne says crawlers linked to Alibaba and other companies have hammered the site over the past three weeks, forcing him to spend nights learning how to configure Cloudflare, write blocking rules, and introduce checks to distinguish humans from bots. “This is quite advanced stuff that I’m having to learn,” he says.\n\nHe alleges that Meta’s crawlers did not respect the site’s robots.txt file, which is meant to tell crawlers whether a site owner wants them accessing particular parts of a site. (Meta, Alibaba, and Tencent, another company Harborne mentioned had hit his server, didn’t respond to *Fast Company*‘s request for comment.)\n\nHarborne’s experience is far from unusual. Independent website operators increasingly find themselves squeezed between rampant bot traffic and an internet infrastructure dominated by a handful of giant companies. [Repeated reports](https://next.ink/wp-content/uploads/2025/03/TollBit-State-of-the-Bots-Q4-2024.pdf) have documented the rise in scraping by AI companies. Cloudflare [estimates that 30% of all web activity](https://blog.cloudflare.com/from-googlebot-to-gptbot-whos-crawling-your-site-in-2025/) is now bot traffic.\n\n“I used to run a web server off a computer that sat in my front room,” says Catherine Flick, a professor of ethics and games technology at Staffordshire University. “You just can’t really get away with doing that anymore, unfortunately, and it’s to the detriment of the web and the internet because the internet really was supposed to be an open space.”\n\nDespite the experience, Harborne plans to continue running the site for as long as he can fund it and has tightened the controls governing how bots can access the LGBT History Project. But he sees the problem as much broader than his own website. Harborne believes AI companies risk damaging the very ecosystem of human-created information that makes their products useful.\n\n“It’s like a virus that’s killing its own host,” he says. “If they’re bringing down the sites that they’re actually creaming the data off, that’s destroying it for everybody.”\n\nPart of the problem, Harborne says, is that small website owners have little recourse when something like this happens. Blocking crawlers becomes their responsibility, while getting the attention of the companies operating them can be almost impossible. “How do you complain to somebody as big as Meta or Apple or Alibaba?” he says. “There’s just no policing of the internet. There’s nobody there.”\n\nThat leaves volunteer operators fighting a largely technical battle themselves, even when they lack the expertise, time, or money of the companies whose automated systems they are trying to keep out.\n\n“This is what happens when profit is allowed to be the prime value, instead of human rights,” adds Carissa Veliz, an AI ethicist at the University of Oxford. “The internet is concerningly close to being a Wild West dominated by bullies, corporate and otherwise.”", "url": "https://wpnews.pro/news/ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history", "canonical_source": "https://www.fastcompany.com/91589354/ai-crawlers-hammered-a-volunteer-run-lgbt-history-archive-2", "published_at": "2026-08-13 10:11:00+00:00", "updated_at": "2026-08-13 10:36:15.581071+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-ethics", "ai-policy"], "entities": ["Meta", "Alibaba", "LGBT History Project", "Jonathan Harborne", "British Library", "AWS", "Claude Code", "Cloudflare"], "alternates": {"html": "https://wpnews.pro/news/ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history", "markdown": "https://wpnews.pro/news/ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history.md", "text": "https://wpnews.pro/news/ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history.txt", "jsonld": "https://wpnews.pro/news/ai-crawlers-from-meta-and-alibaba-almost-destroyed-a-volunteer-run-lgbt-history.jsonld"}}