cd /news/ai-crawlers/see-which-bots-and-ai-crawlers-are-v… · home › topics › ai-crawlers › article
[ARTICLE · art-141972] src=dev.to ↗ pub= topic=ai-crawlers verified=true sentiment=· neutral

See which bots and AI crawlers are visiting your Angular site

WebDecoy published a walkthrough showing how to place its detection middleware in front of an Angular server-side rendering handler so operators can observe bot and AI crawler requests in monitor mode. The example, built on Angular 22.2 with WebDecoy's Express and Node packages at 0.18.0, runs the middleware before Angular's rendering handler and the public catalog API, using a tripwire rule for paths like /.env and a per-process rate limit of 60 requests per minute. The company notes that crawlers hitting static files or CDN cache hits never reach the Node middleware, so collection must happen at that layer instead.

by read5 min views1 publishedSep 29, 2026

If a crawler downloads a page without running JavaScript, an Angular service or HTTP interceptor will never see it. The request still reaches the server that delivers your site.

This walkthrough puts WebDecoy in front of Angular's server-side rendering handler. You can observe requests in monitor mode, inspect the results, and keep serving the application while you decide what needs a response.

We build WebDecoy. This tutorial was prepared with AI assistance and uses a runnable example from our public SDK repository. It covers WebDecoy request detection, not FCaptcha.

The example uses Angular SSR with an Express server:

Crawler or browser → Express → WebDecoy → Angular SSR

The same middleware also sees requests to the application's public catalog API. It runs before both the API route and Angular's rendering handler.

A separate API server cannot observe crawlers that only request HTML from another host. If your Angular app is deployed as static files, put collection at that hosting or CDN layer. Likewise, a CDN cache hit that never reaches Node is outside this middleware's coverage.

Use Node.js 26.5, which was used for these checks, or another version supported by your Angular release.

git clone https://github.com/WebDecoy/node.git
cd node
git checkout 1a6b0700d45f4437311e359d3ab1ea91927fc0bf
cd examples/angular-monitor/app
npm ci --workspaces=false
npm run build
npm test
npm run serve

The example installs the published SDK packages independently of the repository's workspaces. Open http://127.0.0.1:4300 and select Load public catalog.

Use the built Node server for this check. ng serve is not the production server whose middleware order we are testing.

The complete source includes the Angular component, server, lockfile, and integration tests. The example uses Angular 22.2 and WebDecoy's Express and Node packages at 0.18.0.

In src/server.ts, WebDecoy runs before the call to angularApp.handle(req):

import { webdecoy } from '@webdecoy/express';
import { tripwire, rateLimit } from '@webdecoy/node';

app.use(webdecoy({
  apiKey: process.env['WEBDECOY_API_KEY'] || undefined,
  mode: 'monitor',
  honeytoken: false,
  rules: [
    tripwire({ paths: ['/.env', '/wp-config.php'] }),
    rateLimit({ max: 60, window: 60 }),
  ],
}));

This is the middleware block to integrate into your existing Express server, not a complete replacement for server.ts. Keep your existing routes and rendering handler.

mode: 'monitor' records a refusal without enforcing it. The tripwire paths let us test a concrete rule without needing a real crawler. The sample does not expose a real environment file or WordPress configuration file at either path.

The rate limit is intentionally low for the demonstration. It is per process, not a shared limit across a fleet. Adjust or remove it for your workload, or use a shared store if you need one limit across replicas.

honeytoken: false turns off automatic HTML trap injection in this example. That keeps the initial setup focused on request observation without adding DOM changes during hydration.

The server serves real static assets before the middleware and excludes its health endpoint. Other requests pass through detection. The Angular server routes use:

import { RenderMode, ServerRoute } from '@angular/ssr';

export const serverRoutes: ServerRoute[] = [
  { path: '**', renderMode: RenderMode.Server },
];

Angular distinguishes server rendering, prerendering, and client rendering. Its hybrid-rendering documentation explains those choices. Check how your host serves each route before assuming every page request reaches this Node process.

With the server running, send these requests from another terminal:

curl -i http://127.0.0.1:4300/api/products
curl -i http://127.0.0.1:4300/.env

The catalog should return HTTP 200 with two demo items. The tripwire request should produce a terminal record with conclusion: "DENY", tripwire: true, and wouldBlock: true.

Its HTTP response stays 404 because the application defines that response and monitor mode lets the request continue. A refusal in the decision log is not the same thing as a blocked request.

The example logs limited structured decisions. It does not log visitor IPs, full headers, request bodies, or credentials.

For cloud reporting, set your own WEBDECOY_API_KEY in the server environment and restart the server. Create the key in your WebDecoy account.

Keep the key out of Angular configuration, components, and HTTP interceptors. Those can ship to the browser.

Then send the reserved test request:

curl -i -A 'WebDecoy-Test/1.0' http://127.0.0.1:4300/api/products

Look for the labeled Test record in Detections. This checks the reporting pipeline; it does not represent a real AI crawler visit.

Without a key, the example explicitly logs dashboardConfigured: false and reportingError: true for this test. A configured key alone is not proof of delivery either. Check the dashboard record and any reporting errors.

The production build and five integration tests passed locally. Those checks cover server-rendered HTML, catalog availability, the tripwire decision, the reserved test trigger, and rate-limit observation while responses continue.

No API key was used in those tests, so they do not verify cloud delivery or detection accuracy against real AI crawlers.

Start with the requested paths, supporting signals, timing, and crawler classification. Keep a client's claimed user-agent identity separate from any verified identity. A familiar crawler name can be copied by another client.

Before deploying behind a proxy, configure trusted proxies for your actual network path. The local example trusts no forwarding headers and binds only to loopback. Incorrect proxy settings can attribute many visitors to the same address.

Finally, detections are not a complete access log. Requests may be allowed locally without being reported, and caching can affect reporting frequency. Keep access logs for total request counts. Monitor mode gives you a way to investigate before changing how the site responds.

── more in #ai-crawlers 4 stories · sorted by recency
── more on @webdecoy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/see-which-bots-and-a…] indexed:0 read:5min 2026-09-29 · —