cd /news/ai-agents/we-configured-residential-proxies-on… · home › topics › ai-agents › article
[ARTICLE · art-141269] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

We configured residential proxies on our MCP server. They did nothing. Six releases later.

A developer documented six CrawlForge MCP server releases shipped between September 16 and 27, 2026, after discovering that routing traffic through residential proxies did nothing to bypass Cloudflare bot detection. The releases replaced proxy rotation with a multi-rung escalation ladder — an honest CrawlForge User-Agent fetch, a Chrome-TLS client called impit, and a Camoufox stealth browser with Turnstile clicking and clearance-cookie replay — that charges 2 credits for a plain fetch and 2 + 5 when a stealth rung retrieves the page.

by read10 min views1 publishedSep 28, 2026

"What browser MCP to bypass Cloudflare?" is a live question on r/ClaudeAI, and the honest answer for most of 2026 was: none of them reliably, ours included.

In mid-September a bot-detection benchmark we ran against CrawlForge MCP v6.6.2 recommended routing through residential proxies. We configured it. It did nothing. Finding out why turned into six releases in eleven days, and this is the record of what they changed for anyone who needs to scrape a Cloudflare-protected site from an MCP server without pretending to be something they are not.

The plain fetch still runs first and still costs 2 credits. What changed is everything that happens after it is refused.

Version Date The one-line version
v6.7.0 2026-09-16 proxyRotation routes traffic for real; Camoufox stops announcing itself as Chrome
v6.8.0 2026-09-22 The stealth benchmark becomes npm run bench:stealth ; fingerprint leaks closed;"auto" engine defaults to Camoufox
v6.9.0 2026-09-22 agent retries a walled page in the stealth browser, 8 credits + 5 per retry that gets the page
v6.10.0 2026-09-25 Clearance jar: cf_clearance ,__cf_bm anddatadome cookies replayed to the next stealth context
v6.11.0 2026-09-26 A Chrome TLS try ( impit ) before any browser launches; Turnstile checkbox click on Chromium; escalation audit rows
v6.12.0 2026-09-27 App-error fallbacks caught as soft blocks; CRAWLFORGE_IMPIT=off per deployment

No tool was added or renamed, and no price changed except the agent ceiling, which now depends on what got through. Every line above is drawn from the server's changelog.

Cloudflare scores a request before the page is served, and a Node HTTP client fails that score in three places at once: the IP is a known datacenter range, the TLS handshake does not look like a browser's, and the interstitial needs JavaScript the client will never run. The result is a 403, or a 200 carrying nothing but a challenge shell.

Since v5.6.11, scrape has reported that wall as success: false with blocked.vendor naming Cloudflare, DataDome, PerimeterX, Akamai, Amazon or Vercel, whatever the HTTP status, and charged nothing for it. Since v5.9.0, escalate: true has let the same call render the page once in the stealth browser instead of returning the block. The six releases below are what that escalation stage turned into.

With escalate: true and the default engine of "auto", one scrape call now climbs a ladder, and stops at the first rung that returns the page:

CrawlForge/<version> User-Agent. 2 credits. If the host walled a request in the last 24 hours, the MCP server skips this rung rather than repeat a doomed fetch.impit, new in v6.11.0. It presents Chrome's TLS fingerprint and still carries the honest CrawlForge User-Agent. No browser starts. From a residential IP where Cloudflare blocks the plain fetch, indeed.com came back in about a second.challenges.cloudflare.com frame on the page, gets one click on the Turnstile checkbox (v6.11.0). The price is 2 + 5 whichever of the last two rungs got the page, and a call the plain fetch served still pays 2. A page that meets the wall on every rung comes back as a block carrying escalated: true and is charged nothing.

// npm install crawlforge-sdk
import { CrawlForge } from 'crawlforge-sdk';

const client = new CrawlForge({ apiKey: process.env.CRAWLFORGE_API_KEY });

// One call. The plain fetch runs first; only a wall triggers the rest.
const result = await client.scrape({
  url: 'https://www.indeed.com/cmp/Burger-King/reviews',
  formats: ['markdown', 'metadata'],
  escalate: true
});

const data = result.data as {
  escalated?: boolean;
  stealth?: { engine: string };
  content: { markdown: string };
};

// "impit" means the TLS rung got it and no browser ever launched.
console.log(data.escalated, data.stealth?.engine, data.content.markdown.length);

Two guards sit inside the impit rung. Every redirect hop is SSRF-checked with its own DNS resolution, because impit resolves names itself. And a page with under 200 characters of visible text goes on to the browser rather than being returned: quora.com answers a Chrome handshake with its client app's "Something went wrong" fallback under a perfectly normal title, and v6.12.0 now treats that fallback as a soft block on every path.

One deliberate difference: the hosted REST API's scrape escalation stays browser-only. Measured from the hosted instance's datacenter IP, the impit step cleared none of the benchmark's walls, so the hosted service runs with CRAWLFORGE_IMPIT=off while an npm install keeps it on.

Before v6.10.0, every stealth render started cold. If Cloudflare's interstitial let the browser through, that clearance died with the browser context, and the next call to the same host solved it again.

The clearance jar keeps exactly three cookies from a render that got past a wall: cf_clearance, __cf_bm and datadome. They are replayed to the next stealth context with the same engine, User-Agent and proxy, which is the identity the clearance was issued to. That covers the scrape escalation stage, stealth_mode, the agent's stealth retry, browser_session and scrape_with_actions.

Measured on stackoverflow.com from a residential IP: the first stealth call went through Cloudflare's interstitial in 4.1 seconds, with the orchestrate and fo challenge requests visible in the trace. The second went straight to 200 with neither, in 1.6 seconds. A second process reused a clearance from disk and loaded indeed.com with zero challenge-platform requests.

The boundaries matter as much as the feature:

~/.crawlforge/stealth-clearance.json`` CRAWLFORGE_CLEARANCE_JAR=off turns it off. A persistent Chromium profile pool was considered for the same phase and deliberately not built. A shared profile keeps logins and site storage, and on a hosted instance that would leak between customers.

Before v6.9.0, the agent tool's act stage ran only the plain fetch. A challenged seed page was dropped and the answer was assembled from search snippets. In our own review that produced the sentence "3.3 stars, based on 3.3 reviews" for a Burger King page on Indeed, which is what happens when a rating leaks into the count field of a snippet.

Now a page that hits a challenge, a 403 or 429, an empty shell or a timeout is retried through the same escalation stage scrape uses. URLs the caller named are retried first. A discovered URL is retried only when it has no relevant search snippet. Refusals, 404s and 5xx errors are never retried, a run makes at most two retries, and none starts with under 20 seconds of wall clock left.

With the Chromium engine, the same Indeed test read the seed page itself and answered 3.3 stars from 58,942 reviews, with the evidence marked via: "stealth".

Run Credits
No retry needed, or every retry met the wall 8
One retry that got its page 13
Two retries that got their pages (the ceiling) 18

The result reports stealth_retries (attempts) and stealth_retries_charged (billed). A retry that meets the wall again, or throws, is free.

CrawlForge supplies no proxies. What v6.7.0 fixed is that the ones you supply are actually used. stealthConfig.proxyRotation had been a no-op in three separate ways: the proxy went onto Chromium's --proxy-server flag, which cannot carry the user:pass every residential proxy is issued with, so authenticating proxies answered 407; Camoufox's launch path returned before the argument list was built, so the Firefox engine ran unproxied whatever was asked; and the list was read once at launch, so rotationInterval could never elapse.

All three are fixed. Proxies are ordinary URLs (http, https, socks4, socks5, credentials percent-encoded), applied per browser context so both engines authenticate, and a malformed entry is an error rather than a silently unproxied request. Camoufox now runs with its own geoip, block_webrtc and humanize features on, so behind a proxy it derives locale, timezone and location from the exit IP.

v6.8.0 added the server-level form, CRAWLFORGE_STEALTH_PROXIES: a comma-separated list used by every stealth path when a call passes none. A proxy on the call always wins.

The same release fixed a Camoufox identity problem that predated all of this. About two thirds of Camoufox contexts had presented a Chrome User-Agent on a Gecko engine and sent sec-ch-ua client hints Firefox has never implemented. A detector could act on that from the request headers alone, before a line of script ran. Camoufox now presents its own Firefox identity, with nothing Chromium-shaped injected over it.

Every number in this post comes from npm run bench:stealth, which v6.8.0 turned from a hand-run table into a command. It drives the stealth browser and the plain fetch against twelve bot walls and five detector pages, and heads its matrix with the host OS, exit-IP class, engine and browser versions, because a Cloudflare result from a residential connection and one from a datacenter are different measurements.

Measured against it, the fingerprint leaks were closed: navigator.webdriver reads false instead of being deleted, nothing is an own property of navigator, userAgentData drops its HeadlessChrome brand, the Chrome major comes from the installed binary, and a Web Worker answers what the document answers. Detector self-probes went to 19 pass, 0 fail, 1 skip on both engines. npm run bench:stealth:ci runs ten of those probes with no third-party site and fails only on a regression, with a negative control that forces navigator.webdriver to true and must be caught.

Two findings worth carrying with you:

CRAWLFORGE_STEALTH_ENGINE exists to pin a deployment, and why the npm default still prefers Camoufox. What is not closed is stated rather than glossed. Chromium still leaks HeadlessChrome from a SharedWorker, a target Playwright attaches nothing to. And the Turnstile click was verified against Cloudflare's forced-interactive test sitekey, which proves the mechanism only; whether a real site accepts the click still depends on the IP and fingerprint it scores.

The searcher's word is "bypass". Ours is narrower, and these releases held the line on it:

cf_clearance cookie through CrawlForge product token the plain fetch uses. A respect_robots: false override is written to the compliance audit log.stealth_escalation row to logs/compliance-audit.log, after the compliance gate and before the browser navigates, carrying the URL, the tool, the resolved engine and a truncated hash of the API key. A refused request writes none. If you run the server for several clients, the pieces above compose into one deployment:

The environment for a self-hosted, multi-client install

CRAWLFORGE_STEALTH_PROXIES=http://user:pass@proxy-a:8000,socks5://user:pass@proxy-b:1080

CRAWLFORGE_STEALTH_ENGINE=chromium


Then measure before you promise anything: npm run bench:stealth from the box that will do the work, so the matrix carries your exit-IP class and not ours. The audit log gives you a per-key record of every stealth escalation, and the clearance jar's three-cookie rule means one client's login never sits in a context another client's job reuses.

For the hosted API, the honest description is this: stealth traffic exits from a datacenter IP, the escalation is browser-only, and which walls that clears is a property of that IP. Where a client's targets sit behind IP-reputation checks, the self-hosted MCP server with your own proxy list is the setup that measures well.

npm install -g crawlforge-mcp-server@latest
crawlforge --version   # 6.12.0

An MCP client that launches the server with npx picks up v6.12.0 on its next restart. impit is an optional dependency with prebuilt binaries; when it is absent, escalation behaves exactly as it did in v6.10.0. Camoufox is optional too, needs Node 22 for its transitive dependencies, and "auto" falls back to Chromium with a warning that says so. No schema or output shape changed on any tool.

The package is crawlforge-mcp-server on npm, and the full release history is on the changelog. If you have a page that has been blocking you, the free plan's 1,000 one-time credits are enough to see which rung of the ladder clears it from your IP.

This article was written with the help of AI and reviewed by the author.

── more in #ai-agents 4 stories · sorted by recency
── more on @crawlforge 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/we-configured-reside…] indexed:0 read:10min 2026-09-28 · —