{"slug": "how-to-stop-bad-bots-and-ai-scrapers", "title": "How to Stop Bad Bots and AI Scrapers", "summary": "A developer has published a tutorial on using Aegis, an open-source, self-hosted web application firewall and reverse proxy, to block automated scrapers, AI training crawlers, and credential-stuffing bots. The guide walks through enabling bot protection, setting per-crawler policies for GPTBot, CCBot, Anthropic, and Bytespider, and deploying client-side cryptographic browser challenges that drop headless scripts at the edge. It notes that automated traffic accounts for over 40% of all internet traffic.", "body_md": "--\n\nAutomated scrapers, AI training crawlers, and credential stuffing bots account for over 40% of all internet traffic.\n\nAllowing bots to scrape your web application without restrictions inflates hosting bills, degrades performance for real users, and drains database connection pools.\n\nIn this tutorial, we are going to use [Aegis](https://github.com/divinelabio/aegis)—an open-source, self-hosted Web Application Firewall (WAF) and reverse proxy—to configure automated bot defense, manage AI scrapers, and deploy client-side cryptographic browser challenges.\n\n## \n  \n  \n  Step 1: Access Bot Defense Settings\n\n1. Open your Aegis Admin Console at `http://server-ip:8081` .\n2. In the left navigation menu, navigate to **Traffic Control > Bot Defense** .\n3. Toggle on **Enable Bot Protection** .\n\nAegis evaluates incoming client fingerprints, request headers, and traffic behavioral patterns in real time.\n\n## \n  \n  \n  Step 2: Configure AI Crawlers and Scraper Policies\n\nManage how automated crawlers access your site:\n\n1. \n**Search Engine Crawlers:** Toggle on**Allow Verified Search Engines** (Googlebot, Bingbot, DuckDuckGo). Aegis verifies reverse-DNS signatures to ensure fake bots cannot spoof search engine user-agents.\n2. \n**AI Scrapers & Training Bots:** Select your policy for known AI crawlers (GPTBot, CCBot, Anthropic, Bytespider):  - \n**Block:** Immediately drop requests with`HTTP 403 Forbidden` .\n  - \n**Challenge:** Require the bot to pass a client-side challenge.\n3. Click **Apply Policy** .\n\n## \n  \n  \n  Step 3: Deploy Client-Side Browser Challenges (Smart Challenge)\n\nFor requests that exhibit suspicious behavior or exceed baseline request velocities:\n\n1. Under **Mitigation Strategy** , select**Browser Challenge (Smart Challenge)** .\n2. When a suspected bot requests a protected page, Aegis returns a lightweight HTML payload that executes a fast cryptographic challenge in the background.\n3. Genuine human visitors running modern web browsers solve the challenge automatically in milliseconds without seeing a CAPTCHA.\n4. Headless scripts, scraping tools, and automated botnets fail the execution and are dropped at the edge.\n\n## \n  \n  \n  Step 4: Verify Bot Mitigation\n\nTest the endpoint with an automated CLI client:\n\nResponse:\n\nThe request is rejected at the ingress proxy and never reaches your origin web server.\n\n## \n  \n  \n  Resources\n\nThe Community Edition is free to self-host:", "url": "https://wpnews.pro/news/how-to-stop-bad-bots-and-ai-scrapers", "canonical_source": "https://dev.to/divinelab/how-to-stop-bad-bots-and-ai-scrapers-3561", "published_at": "2026-10-09 08:17:54+00:00", "updated_at": "2026-10-09 08:21:29.131324+00:00", "lang": "en", "topics": ["ai-crawlers", "ai-tools", "developer-tools"], "entities": ["Aegis", "GPTBot", "CCBot", "Anthropic", "Bytespider", "Googlebot", "Bingbot", "DuckDuckGo"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-stop-bad-bots-and-ai-scrapers", "markdown": "https://wpnews.pro/news/how-to-stop-bad-bots-and-ai-scrapers.md", "text": "https://wpnews.pro/news/how-to-stop-bad-bots-and-ai-scrapers.txt", "jsonld": "https://wpnews.pro/news/how-to-stop-bad-bots-and-ai-scrapers.jsonld"}}