# How to Scrape JavaScript-Heavy Pages for AI Agents

> Source: <https://dev.to/agentsearchhq/how-to-scrape-javascript-heavy-pages-for-ai-agents-ifj>
> Published: 2026-10-10 22:10:10+00:00

Point a plain HTTP fetcher at a React or Vue single-page app and you often get back a shell: a title, an empty `<div id="root">`, and a line asking you to enable JavaScript. Hand that to an LLM and it will either say the page is empty or, worse, answer from memory.

The fix is not to run a headless browser on every URL. It is to fetch cheaply first, notice when the result is thin, and render only those pages. This guide builds that fallback in TypeScript on the AgentSearch tools on the Pocket Agentic Portal: [Web Extract](https://agentsearchhq.com/extract?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents) for static HTML and [Web Render](https://agentsearchhq.com/render?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents) for pages that need a browser. Each call is $0.005 in USDC, paid over x402 on Base or MPP on Tempo, with no signup and no API key.

A server-rendered page ships its text in the HTML. A client-rendered page ships JavaScript that builds the text after load. Extract reads the HTML it receives and converts it to markdown, so on a client-rendered page it faithfully converts the empty template.

Render runs the page in headless Chromium, then returns what the browser ended up with: `markdown`, `text`, `html`, `title`, `links`, and an optional screenshot. It is a real browser, so it is the tool for JS-heavy pages. It is also the same $0.005 per call as Extract, which is why it works as a fallback: one thin Extract plus one Render is $0.01 for that page, and pages that extract cleanly never pay for the second call.

This is the same wrapper as in [Error Handling for AI Agent Web Tools](https://agentsearchhq.com/blog/error-handling-ai-agent-web-tools?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents), trimmed. `payFetch` is the x402 client from [Pay-per-call APIs for AI agents](https://agentsearchhq.com/blog/x402-pay-per-call-api-for-ai-agents?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents), with a spend cap installed.

``` js
// js-pages.ts
const PORTAL = "https://agent.pocket.network/v1";
declare const payFetch: typeof fetch; // from the x402 tutorial

type PortalFailure = { kind: "portal"; status: number; code: string; retryable: boolean };
type Delivered = { kind: "data"; data: Record<string, unknown> };

async function callService(serviceId: string, path: string, body: unknown): Promise<PortalFailure | Delivered> {
  const res = await payFetch(`${PORTAL}/${serviceId}${path}`, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify(body),
  });
  const json = await res.json();
  if (!res.ok) {
    const err = json.error ?? {};
    return {
      kind: "portal",
      status: res.status,
      code: String(err.code ?? `HTTP_${res.status}`),
      retryable: err.retryable === true,
    };
  }
  return { kind: "data", data: json.data };
}
```

Extract does not tell you "this page needed JavaScript." You infer it from what came back. Two cheap signals cover most cases: the markdown is very short, or it contains the page's own no-JavaScript notice.

```
type PageData = {
  url?: string;
  title?: string | null;
  markdown?: string;
  error?: { code: string; retryable?: boolean } | null;
};

const THIN = 400; // local heuristic, tune for your sources
const JS_HINTS = ["enable javascript", "requires javascript", "javascript is disabled"];

function looksClientRendered(page: PageData): boolean {
  if (page.error) return false; // an error is a different problem; see Step 4
  const md = (page.markdown ?? "").trim();
  const lower = md.toLowerCase();
  return md.length < THIN || JS_HINTS.some((h) => lower.includes(h));
}
```

`THIN` is your number, not a service field. A short page with no error is not proof of a JavaScript shell (some pages are just short), but it is a reasonable trigger for one Render call. Do not retry Extract on a thin result: it will fetch the same HTML again.

Request only the formats you will use. Markdown is what an LLM wants; add `links` if the agent needs to crawl onward.

```
type ModelPage =
  | { status: "ok"; url: string; title: string | null; markdown: string; note?: string }
  | { status: "skip"; url: string; code: string }
  | { status: "retry_later"; url: string; code: string };

export async function readJsPage(url: string): Promise<ModelPage> {
  const extract = await callService("agentsearch-web-extract-v1", "/v1/extract", {
    url,
    formats: ["markdown"],
    max_chars: 8000,
  });
  if (extract.kind === "portal") return portalResult(url, extract);

  const first = extract.data as PageData;
  if (!looksClientRendered(first)) return toModel(first, url);

  const render = await callService("agentsearch-web-render-v1", "/v1/render", {
    url,
    formats: ["markdown"],
    max_chars: 8000,
  });
  if (render.kind === "portal") return portalResult(url, render);

  const page = toModel(render.data as PageData, url);
  if (page.status === "ok" && page.markdown.length < THIN && !page.note) {
    return { status: "skip", url: page.url, code: "EMPTY_PAGE" };
  }
  return page;
}
```

If Render also comes back nearly empty, the page probably has nothing public to read (or needs a login, which Render does not do). Skip it rather than paying again.

Render accepts a few options for slow apps: `wait_until` (`domcontentloaded` by default, or `load` or `networkidle`), `wait_for_selector` (a CSS selector, waited on for at most 1.5 s) and `wait_ms` (0 to 1500), plus `block` to skip loading `image`, `media`, `font` or `stylesheet` requests (default `["media", "font"]`). Waiting longer eats into the deadline below, so start with the defaults.

Three outcomes need their own handling.

**robots.txt blocks the page.** Render honors robots.txt (RFC 9309, user-agent token `AgentSearchRender`; its browser identifies itself with a Chrome user-agent string ending in `AgentSearchRender/0.1 (+https://agentsearchhq.com/agents.md)`). A disallowed URL is refused before any fetch. Called directly, Render answers HTTP 422:

```
{"error": {"code": "ROBOTS_DISALLOWED", "message": "robots.txt disallows this URL for AgentSearchRender.", "retryable": false, "request_id": "â€¦"}}
```

Through the Pocket portal this surfaces as HTTP 400 with code `UPSTREAM_REJECTED`, and the call is not charged on x402. For a plain disallow the answer will not change, so treat it as "use a different source" and respect the site's choice. (If the site's robots.txt itself answered with a server error, Render also refuses, but marks that one `retryable: true`.)

**The render hits its deadline.** Render's hard deadline is 4.2 s, counted from when the request arrives. When it runs out, the response is still HTTP 200 with `error.code` set to `TARGET_DEADLINE`, which is always `retryable: true`, and `markdown` may already hold the text that rendered in time. Keep that partial page and tell the model it may be incomplete.

**Other website errors.** These arrive inside `data.error` on an HTTP 200, and the call counts as delivered. Branch on `code` and `retryable`, never on `message`.

```
function portalResult(url: string, f: PortalFailure): ModelPage {
  // robots.txt refusal: 400 UPSTREAM_REJECTED via the portal (not charged on x402), 422 ROBOTS_DISALLOWED direct
  if (f.code === "UPSTREAM_REJECTED" || f.code === "ROBOTS_DISALLOWED") return { status: "skip", url, code: f.code };
  return f.retryable ? { status: "retry_later", url, code: f.code } : { status: "skip", url, code: f.code };
}

function toModel(page: PageData, requested: string): ModelPage {
  const url = page.url || requested; // final URL after redirects: cite this
  const md = (page.markdown ?? "").trim();
  const err = page.error ?? null;
  if (err) {
    if (err.code === "TARGET_DEADLINE" && md) {
      return { status: "ok", url, title: page.title ?? null, markdown: md, note: "partial: render hit its 4.2s deadline" };
    }
    return err.retryable ? { status: "retry_later", url, code: err.code } : { status: "skip", url, code: err.code };
  }
  return { status: "ok", url, title: page.title ?? null, markdown: md };
}
```

`retry_later` means "try this URL on another turn," not "loop now." Each delivered retry is a new $0.005 call. The full code list, including Extract's `TARGET_TIMEOUT` and Render's `TARGET_BOT_WALL` (a bot check or challenge page, which Render reports rather than bypasses; skip that source), is in the [error-handling guide](https://agentsearchhq.com/blog/error-handling-ai-agent-web-tools?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents).

The fallback already does most of the work: static pages cost one call, JS-heavy pages cost two. A few more habits keep a loop predictable:

`url` are one page.`markdown` alone is enough for most agent reads; skip `html` and the screenshot unless you need them.
Pass the resulting markdown to the model as a tool result, never as part of the system prompt. Rendered or not, it is still someone else's text.

If you would rather not write a client, Pocket's MCP server pays over x402 on Base with no backend of your own:

```
{
  "mcpServers": {
    "pocket-network": {
      "command": "npx",
      "args": ["-y", "@pocket-network/agentic-portal-mcp"],
      "env": {
        "POCKET_PRIVATE_KEY": "0xâ€¦",
        "POCKET_MAX_TOTAL_ATOMIC": "1000000",
        "POCKET_MAX_PER_CALL_ATOMIC": "5000"
      }
    }
  }
}
```

Amounts are USDC atomic units (6 decimals): `POCKET_MAX_TOTAL_ATOMIC` of `1000000` caps total signing at $1.00 while the server runs (it resets on restart and is required to pay), and `POCKET_MAX_PER_CALL_ATOMIC` of `5000` caps any one call at $0.005. The limits are checked before anything is signed. Use a wallet that holds only what you are willing to spend. Tell the model the same rule in its instructions: call Extract first, and call Render only when the markdown is thin or asks for JavaScript.

`POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract` (`POST https://agent.pocket.network/v1/agentsearch-web-render-v1/v1/render` (`POST https://agent.pocket.network/v1/agentsearch-web-search-v1/v1/search` (
For choosing between the three tools in general, see [Search, Extract, Render: three web tools for AI agents](https://agentsearchhq.com/blog/search-extract-render-web-tools-for-ai-agents?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents).

Fetch it with a static extractor first. If the markdown is very short or asks you to enable JavaScript, send the same URL to a headless browser such as Web Render, which runs the page in Chromium and returns markdown.

Look at the extracted text: a near-empty body, an empty app container, or a "please enable JavaScript" notice are the usual signs. Use a length threshold you tune for your sources.

Web Render returns markdown, text, HTML, the title, links and an optional screenshot, so an LLM can read the page directly.

Not on x402. Render refuses the URL before fetching it (HTTP 422 `ROBOTS_DISALLOWED` directly); through the Pocket portal that surfaces as HTTP 400 `UPSTREAM_REJECTED` and the call is not charged.

Render stops at 4.2 s and returns HTTP 200 with `error.code` `TARGET_DEADLINE`, which is retryable and may include the markdown that rendered in time.

Search, Extract and Render are each $0.005 per call in USDC, over x402 on Base or MPP on Tempo. No signup and no API key.

Add `looksClientRendered` and the Render fallback to any loop that already calls Extract. Start with [Web Extract](https://agentsearchhq.com/extract?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents), fall back to [Web Render](https://agentsearchhq.com/render?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents), and use [Web Search](https://agentsearchhq.com/search?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents) when the agent does not have a URL yet. More on the [AgentSearch blog](https://agentsearchhq.com/blog/?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents).

*Originally published on the [AgentSearch blog](https://agentsearchhq.com/blog/scrape-javascript-heavy-pages-for-ai-agents?utm_source=devto&utm_medium=social&utm_campaign=scrape-javascript-heavy-pages-for-ai-agents).*
