# How to Stop Bad Bots and AI Scrapers

> Source: <https://dev.to/divinelab/how-to-stop-bad-bots-and-ai-scrapers-3561>
> Published: 2026-10-09 08:17:54+00:00

--

Automated scrapers, AI training crawlers, and credential stuffing bots account for over 40% of all internet traffic.

Allowing bots to scrape your web application without restrictions inflates hosting bills, degrades performance for real users, and drains database connection pools.

In this tutorial, we are going to use [Aegis](https://github.com/divinelabio/aegis)—an open-source, self-hosted Web Application Firewall (WAF) and reverse proxy—to configure automated bot defense, manage AI scrapers, and deploy client-side cryptographic browser challenges.

## 
  
  
  Step 1: Access Bot Defense Settings

1. Open your Aegis Admin Console at `http://server-ip:8081` .
2. In the left navigation menu, navigate to **Traffic Control > Bot Defense** .
3. Toggle on **Enable Bot Protection** .

Aegis evaluates incoming client fingerprints, request headers, and traffic behavioral patterns in real time.

## 
  
  
  Step 2: Configure AI Crawlers and Scraper Policies

Manage how automated crawlers access your site:

1. 
**Search Engine Crawlers:** Toggle on**Allow Verified Search Engines** (Googlebot, Bingbot, DuckDuckGo). Aegis verifies reverse-DNS signatures to ensure fake bots cannot spoof search engine user-agents.
2. 
**AI Scrapers & Training Bots:** Select your policy for known AI crawlers (GPTBot, CCBot, Anthropic, Bytespider):  - 
**Block:** Immediately drop requests with`HTTP 403 Forbidden` .
  - 
**Challenge:** Require the bot to pass a client-side challenge.
3. Click **Apply Policy** .

## 
  
  
  Step 3: Deploy Client-Side Browser Challenges (Smart Challenge)

For requests that exhibit suspicious behavior or exceed baseline request velocities:

1. Under **Mitigation Strategy** , select**Browser Challenge (Smart Challenge)** .
2. When a suspected bot requests a protected page, Aegis returns a lightweight HTML payload that executes a fast cryptographic challenge in the background.
3. Genuine human visitors running modern web browsers solve the challenge automatically in milliseconds without seeing a CAPTCHA.
4. Headless scripts, scraping tools, and automated botnets fail the execution and are dropped at the edge.

## 
  
  
  Step 4: Verify Bot Mitigation

Test the endpoint with an automated CLI client:

Response:

The request is rejected at the ingress proxy and never reaches your origin web server.

## 
  
  
  Resources

The Community Edition is free to self-host:
