cd /news/developer-tools/stop-reaching-for-a-headless-browser… · home topics developer-tools article
[ARTICLE · art-92845] src=dev.to ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Stop Reaching for a Headless Browser to Scrape Documentation Sites

A developer outlines lightweight alternatives to headless browsers for scraping documentation sites, noting that static site generators like Next.js and Docusaurus leave content in predictable places such as __NEXT_DATA__ tags or pre-rendered HTML. The post recommends checking the initial HTML response, using curl to verify content presence, and sparse-cloning markdown from public repos for cleaner input, while cautioning about license checks.

read3 min views1 publishedAug 11, 2026

Every few days someone in a scraping forum asks a version of the same question: "I'm collecting documentation text for an AI tool, but the pages render with JavaScript. What's the lightest way to get the content?"

The answers are always the same — open DevTools, check the Network tab, find the XHR. That advice is correct, and for documentation sites specifically it's usually unnecessary work.

Documentation sites are not arbitrary web apps. They're overwhelmingly built by a handful of static site generators, and those generators leave the content sitting in predictable places. Three routes cover most of what you'll hit, and none of them need a browser.

Next.js-based docs (which includes a large share of company developer portals) embed the full page payload in a __NEXT_DATA__

script tag. It's in the initial HTML response — no JavaScript execution needed.

import json, re, httpx

html = httpx.get(url, follow_redirects=True).text
m = re.search(r'<script id="__NEXT_DATA__" type="application/json">(.*?)</script>', html, re.S)
if m:
    data = json.loads(m.group(1))
    print(json.dumps(data["props"]["pageProps"], indent=2)[:2000])

Docusaurus (the other big one) doesn't use __NEXT_DATA__

, but it pre-renders the full article text into the static HTML. A plain httpx.get

plus a main

or article

selector gets you everything. The "JavaScript rendering" you see in the browser is hydration for navigation and search — the prose was already there in the initial response.

Thirty-second check: curl -s <url> | grep -c "some sentence you can see on the page"

. If that returns 1 or more, the content is in the HTML and you're done. This one check resolves the majority of "it's JavaScript-rendered" cases, because people judge from DevTools' Elements panel — which shows the hydrated DOM, not what the server actually sent.

Most open-source project documentation is markdown in the same repository as the code, usually under docs/

. Scraping the rendered HTML means fetching pages one at a time, parsing them, and stripping navigation chrome you didn't want. Cloning gets you the clean source in one shot:

git clone --depth 1 --filter=blob:none --sparse https://github.com/org/project
cd project && git sparse-checkout set docs

You get the original markdown — headings intact, code fences intact, no nav sidebars, no cookie banners, no per-page rate limiting. For a RAG pipeline this is strictly better input than parsed HTML, and it's one request instead of several hundred.

Worth saying plainly: check the license before you ingest. Documentation is frequently licensed separately from the code, and "the repo is public" is not the same as "you may redistribute this."

A growing number of documentation hosts expose plain-text views:

/llms.txt

/sitemap.xml

Signal Route
curl output contains visible page text
Parse the static HTML — done
__NEXT_DATA__ in the HTML
Extract and walk the JSON
Public repo with a docs/ directory
Sparse-clone the markdown
/llms.txt returns 200
Start there
ReadTheDocs / GitBook host Look for the download build
None of the above
Now open DevTools

Some cases are genuinely dynamic, and it's worth knowing them so you don't over-apply the above:

For a few hundred pages of prose, though, these are the exception. The default assumption should be that the text is already reachable, and the browser is the fallback.

The reason this matters isn't purity — it's that headless browsers change the shape of your project. You go from a script anyone can run to a pipeline with a browser binary, a memory ceiling, per-page startup cost, and a new class of flaky failures that only reproduce sometimes. On a few hundred documentation pages, Route 1 or Route 2 typically finishes before a browser-based run has finished launching.

Check whether the content is already sitting there in plain text. Most of the time, it is.

We publish code examples and testing notes for developers who scrape and automate at RoamProxy. More runnable examples: github.com/roamproxy/proxy-examples.

── more in #developer-tools 4 stories · sorted by recency
── more on @next.js 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/stop-reaching-for-a-…] indexed:0 read:3min 2026-08-11 ·