cd /news/ai-agents/how-to-scrape-javascript-heavy-pages… · home › topics › ai-agents › article
[ARTICLE · art-148929] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

How to Scrape JavaScript-Heavy Pages for AI Agents

A developer published a TypeScript guide for scraping JavaScript-heavy pages on behalf of AI agents, using a two-tier approach that fetches static HTML first and only falls back to headless Chromium rendering when the result looks thin. The method relies on AgentSearch's Web Extract and Web Render tools on the Pocket Agentic Portal, each priced at $0.005 in USDC via x402 on Base or MPP on Tempo with no signup or API key. The guide detects client-rendered pages heuristically by flagging markdown under 400 characters or text containing no-JavaScript notices.

by read8 min views1 publishedOct 10, 2026

Point a plain HTTP fetcher at a React or Vue single-page app and you often get back a shell: a title, an empty <div id="root">, and a line asking you to enable JavaScript. Hand that to an LLM and it will either say the page is empty or, worse, answer from memory.

The fix is not to run a headless browser on every URL. It is to fetch cheaply first, notice when the result is thin, and render only those pages. This guide builds that fallback in TypeScript on the AgentSearch tools on the Pocket Agentic Portal: Web Extract for static HTML and Web Render for pages that need a browser. Each call is $0.005 in USDC, paid over x402 on Base or MPP on Tempo, with no signup and no API key.

A server-rendered page ships its text in the HTML. A client-rendered page ships JavaScript that builds the text after load. Extract reads the HTML it receives and converts it to markdown, so on a client-rendered page it faithfully converts the empty template.

Render runs the page in headless Chromium, then returns what the browser ended up with: markdown, text, html, title, links, and an optional screenshot. It is a real browser, so it is the tool for JS-heavy pages. It is also the same $0.005 per call as Extract, which is why it works as a fallback: one thin Extract plus one Render is $0.01 for that page, and pages that extract cleanly never pay for the second call.

This is the same wrapper as in Error Handling for AI Agent Web Tools, trimmed. payFetch is the x402 client from Pay-per-call APIs for AI agents, with a spend cap installed.

// js-pages.ts
const PORTAL = "https://agent.pocket.network/v1";
declare const payFetch: typeof fetch; // from the x402 tutorial

type PortalFailure = { kind: "portal"; status: number; code: string; retryable: boolean };
type Delivered = { kind: "data"; data: Record<string, unknown> };

async function callService(serviceId: string, path: string, body: unknown): Promise<PortalFailure | Delivered> {
  const res = await payFetch(`${PORTAL}/${serviceId}${path}`, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify(body),
  });
  const json = await res.json();
  if (!res.ok) {
    const err = json.error ?? {};
    return {
      kind: "portal",
      status: res.status,
      code: String(err.code ?? `HTTP_${res.status}`),
      retryable: err.retryable === true,
    };
  }
  return { kind: "data", data: json.data };
}

Extract does not tell you "this page needed JavaScript." You infer it from what came back. Two cheap signals cover most cases: the markdown is very short, or it contains the page's own no-JavaScript notice.

type PageData = {
  url?: string;
  title?: string | null;
  markdown?: string;
  error?: { code: string; retryable?: boolean } | null;
};

const THIN = 400; // local heuristic, tune for your sources
const JS_HINTS = ["enable javascript", "requires javascript", "javascript is disabled"];

function looksClientRendered(page: PageData): boolean {
  if (page.error) return false; // an error is a different problem; see Step 4
  const md = (page.markdown ?? "").trim();
  const lower = md.toLowerCase();
  return md.length < THIN || JS_HINTS.some((h) => lower.includes(h));
}

THIN is your number, not a service field. A short page with no error is not proof of a JavaScript shell (some pages are just short), but it is a reasonable trigger for one Render call. Do not retry Extract on a thin result: it will fetch the same HTML again.

Request only the formats you will use. Markdown is what an LLM wants; add links if the agent needs to crawl onward.

type ModelPage =
  | { status: "ok"; url: string; title: string | null; markdown: string; note?: string }
  | { status: "skip"; url: string; code: string }
  | { status: "retry_later"; url: string; code: string };

export async function readJsPage(url: string): Promise<ModelPage> {
  const extract = await callService("agentsearch-web-extract-v1", "/v1/extract", {
    url,
    formats: ["markdown"],
    max_chars: 8000,
  });
  if (extract.kind === "portal") return portalResult(url, extract);

  const first = extract.data as PageData;
  if (!looksClientRendered(first)) return toModel(first, url);

  const render = await callService("agentsearch-web-render-v1", "/v1/render", {
    url,
    formats: ["markdown"],
    max_chars: 8000,
  });
  if (render.kind === "portal") return portalResult(url, render);

  const page = toModel(render.data as PageData, url);
  if (page.status === "ok" && page.markdown.length < THIN && !page.note) {
    return { status: "skip", url: page.url, code: "EMPTY_PAGE" };
  }
  return page;
}

If Render also comes back nearly empty, the page probably has nothing public to read (or needs a login, which Render does not do). Skip it rather than paying again.

Render accepts a few options for slow apps: wait_until (domcontentloaded by default, or load or networkidle), wait_for_selector (a CSS selector, waited on for at most 1.5 s) and wait_ms (0 to 1500), plus block to skip image, media, font or stylesheet requests (default ["media", "font"]). Waiting longer eats into the deadline below, so start with the defaults.

Three outcomes need their own handling.

robots.txt blocks the page. Render honors robots.txt (RFC 9309, user-agent token AgentSearchRender; its browser identifies itself with a Chrome user-agent string ending in AgentSearchRender/0.1 (+https://agentsearchhq.com/agents.md)). A disallowed URL is refused before any fetch. Called directly, Render answers HTTP 422:

{"error": {"code": "ROBOTS_DISALLOWED", "message": "robots.txt disallows this URL for AgentSearchRender.", "retryable": false, "request_id": "…"}}

Through the Pocket portal this surfaces as HTTP 400 with code UPSTREAM_REJECTED, and the call is not charged on x402. For a plain disallow the answer will not change, so treat it as "use a different source" and respect the site's choice. (If the site's robots.txt itself answered with a server error, Render also refuses, but marks that one retryable: true.)

The render hits its deadline. Render's hard deadline is 4.2 s, counted from when the request arrives. When it runs out, the response is still HTTP 200 with error.code set to TARGET_DEADLINE, which is always retryable: true, and markdown may already hold the text that rendered in time. Keep that partial page and tell the model it may be incomplete.

Other website errors. These arrive inside data.error on an HTTP 200, and the call counts as delivered. Branch on code and retryable, never on message.

function portalResult(url: string, f: PortalFailure): ModelPage {
  // robots.txt refusal: 400 UPSTREAM_REJECTED via the portal (not charged on x402), 422 ROBOTS_DISALLOWED direct
  if (f.code === "UPSTREAM_REJECTED" || f.code === "ROBOTS_DISALLOWED") return { status: "skip", url, code: f.code };
  return f.retryable ? { status: "retry_later", url, code: f.code } : { status: "skip", url, code: f.code };
}

function toModel(page: PageData, requested: string): ModelPage {
  const url = page.url || requested; // final URL after redirects: cite this
  const md = (page.markdown ?? "").trim();
  const err = page.error ?? null;
  if (err) {
    if (err.code === "TARGET_DEADLINE" && md) {
      return { status: "ok", url, title: page.title ?? null, markdown: md, note: "partial: render hit its 4.2s deadline" };
    }
    return err.retryable ? { status: "retry_later", url, code: err.code } : { status: "skip", url, code: err.code };
  }
  return { status: "ok", url, title: page.title ?? null, markdown: md };
}

retry_later means "try this URL on another turn," not "loop now." Each delivered retry is a new $0.005 call. The full code list, including Extract's TARGET_TIMEOUT and Render's TARGET_BOT_WALL (a bot check or challenge page, which Render reports rather than bypasses; skip that source), is in the error-handling guide.

The fallback already does most of the work: static pages cost one call, JS-heavy pages cost two. A few more habits keep a loop predictable:

url are one page.markdown alone is enough for most agent reads; skip html and the screenshot unless you need them. Pass the resulting markdown to the model as a tool result, never as part of the system prompt. Rendered or not, it is still someone else's text.

If you would rather not write a client, Pocket's MCP server pays over x402 on Base with no backend of your own:

{
  "mcpServers": {
    "pocket-network": {
      "command": "npx",
      "args": ["-y", "@pocket-network/agentic-portal-mcp"],
      "env": {
        "POCKET_PRIVATE_KEY": "0x…",
        "POCKET_MAX_TOTAL_ATOMIC": "1000000",
        "POCKET_MAX_PER_CALL_ATOMIC": "5000"
      }
    }
  }
}

Amounts are USDC atomic units (6 decimals): POCKET_MAX_TOTAL_ATOMIC of 1000000 caps total signing at $1.00 while the server runs (it resets on restart and is required to pay), and POCKET_MAX_PER_CALL_ATOMIC of 5000 caps any one call at $0.005. The limits are checked before anything is signed. Use a wallet that holds only what you are willing to spend. Tell the model the same rule in its instructions: call Extract first, and call Render only when the markdown is thin or asks for JavaScript.

POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract (POST https://agent.pocket.network/v1/agentsearch-web-render-v1/v1/render (POST https://agent.pocket.network/v1/agentsearch-web-search-v1/v1/search ( For choosing between the three tools in general, see Search, Extract, Render: three web tools for AI agents.

Fetch it with a static extractor first. If the markdown is very short or asks you to enable JavaScript, send the same URL to a headless browser such as Web Render, which runs the page in Chromium and returns markdown.

Look at the extracted text: a near-empty body, an empty app container, or a "please enable JavaScript" notice are the usual signs. Use a length threshold you tune for your sources.

Web Render returns markdown, text, HTML, the title, links and an optional screenshot, so an LLM can read the page directly.

Not on x402. Render refuses the URL before fetching it (HTTP 422 ROBOTS_DISALLOWED directly); through the Pocket portal that surfaces as HTTP 400 UPSTREAM_REJECTED and the call is not charged.

Render stops at 4.2 s and returns HTTP 200 with error.code TARGET_DEADLINE, which is retryable and may include the markdown that rendered in time.

Search, Extract and Render are each $0.005 per call in USDC, over x402 on Base or MPP on Tempo. No signup and no API key.

Add looksClientRendered and the Render fallback to any loop that already calls Extract. Start with Web Extract, fall back to Web Render, and use Web Search when the agent does not have a URL yet. More on the AgentSearch blog.

Originally published on the AgentSearch blog.

── more in #ai-agents 4 stories · sorted by recency
── more on @agentsearch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-scrape-javasc…] indexed:0 read:8min 2026-10-10 · —