cd /news/ai-agents/search-extract-render-three-web-tool… · home › topics › ai-agents › article
[ARTICLE · art-145069] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Search Extract Render: Three Web Tools for AI Agents, Not One Browse

A developer has split the monolithic browse tool common in agent frameworks into three separate MCP-compatible services — Web Search, Web Extract, and Web Render — each exposed as a distinct Pocket Network service ID so agents can reason about which step to invoke. The design returns structured error objects with retryable flags on HTTP 200 and enforces a 4-second fetch deadline on Extract, letting agent loops distinguish timeouts from JavaScript-only shells. The tools connect to Claude Desktop, Cursor, or any other MCP client.

by read10 min views1 publishedOct 5, 2026

Most agent frameworks ship with a single browse tool. The model hands it a query or a URL, and somewhere behind it a pipeline searches, fetches, maybe spins up a headless browser, and returns a blob of text. It works in a demo. In production it hides three different jobs behind one interface. You can't see which one failed, you can't control what each costs, and you pay for a browser on pages that never needed one.

This guide splits that tool into three steps your agent can reason about:

We'll go through what each step does, write an agent loop that decides when to step up, and connect it all to Claude Desktop, Cursor or any other MCP client.

A monolithic browse tool has to guess. Should it render every page in Chromium just in case? That's slow and heavy. Should it fetch raw HTML only? Then single-page apps come back as an empty <div id="root">. Either way the agent can't tell what happened.

Three separate tools give you:

error field, so your loop can tell "the site timed out" apart from "the page is a JS shell".max_chars cap. The three AgentSearch services map to this directly. Each one is a Pocket Network service with its own ID:

Step Service ID Endpoint Returns
Search agentsearch-web-search-v1 POST /v1/search Up to 5 results: title ,url ,content ,score ,domain , …
Extract agentsearch-web-extract-v1 POST /v1/extract title ,markdown ,text ,links ,meta ,error
Render agentsearch-web-render-v1 POST /v1/render title ,markdown ,text , optionalhtml ,links ,screenshot ,meta ,error

Web Search takes a query and an optional max_results (1–5, default 5) and returns live results as JSON:

{
  "query": "pocket network agentic portal x402",
  "max_results": 5
}

Each result carries title, url, content (a snippet), score, published_date, retrieved_at and domain. Five results is deliberate. An agent rarely needs more than a handful of sources to answer a question, and a short list keeps the next step, Extract, bounded.

If the search can't be completed, the response is still HTTP 200, results is empty, and an error object explains why (for example UPSTREAM_TIMEOUT, with a retryable flag). Branch on that flag; don't parse error strings.

Web Extract is the workhorse. Give it a URL and you get back clean markdown with headings and links kept and the page clutter stripped, plus the title, up to 100 links and fetch metadata:

{
  "url": "https://example.com/",
  "formats": ["markdown"],
  "max_chars": 8000
}

Useful parameters from the OpenAPI spec:

formats: ["markdown"] (default), ["text"], or both. include_links: set it to false when you only want body text. The response meta tells you status_code, content_type, chars and whether the output was truncated. Target-site failures ( TARGET_TIMEOUT, TARGET_HTTP_ERROR, UNSUPPORTED_CONTENT for things like PDFs, and so on) come back in error with a retryable flag, again on HTTP 200. Extract works within a hard 4-second fetch deadline, so a slow site fails fast instead of stalling your agent.

For most of the web (docs, blogs, news, reference pages) Extract is all you need. For a deeper look at turning pages into chunkable context, see our guide to URL to markdown for RAG.

Some pages ship almost no content in their HTML. Single-page apps, dashboards and infinite-scroll lists build the page in the browser. Extract sees what the server sent, and for those sites that's a near-empty shell and a "please enable JavaScript" notice.

Web Render handles that case. It loads the URL in headless Chromium, lets the JavaScript run, and returns the title, rendered text or markdown, and HTML if you ask for it, plus links and an optional screenshot:

{
  "url": "https://react.dev/",
  "formats": ["markdown"],
  "max_chars": 8000,
  "wait_until": "domcontentloaded",
  "screenshot": false
}

Other options include wait_for_selector (waits up to 1.5 s), wait_ms (up to 1,500), viewport, and block to skip images, media, fonts or stylesheets. Render has a 4.2-second deadline. A slow page returns whatever loaded so far with error.code set to TARGET_DEADLINE. Render is polite by design: it honors robots.txt (user-agent token AgentSearchRender), never logs in, clicks or submits forms, and reports bot walls (TARGET_BOT_WALL) instead of bypassing them.

Render is live on Pocket Network MainNet as agentsearch-web-render-v1. Its Agentic Portal listing is coming soon, and no portal price is published yet. Until then it is served over Pocket Network relays. Check the Render page for access details and updates.

The rule of thumb: Extract first, Render only when Extract comes back empty or as a bare JS shell.

Here is the full decision loop in TypeScript. Search and Extract are called through the Pocket Agentic Portal, which charges per call over x402. We use the standard x402 v2 client packages, @x402/fetch and @x402/evm, so the 402 → sign → retry handshake happens automatically.

npm install @x402/fetch @x402/evm viem
js
// agent-web.ts
import { wrapFetchWithPaymentFromConfig } from "@x402/fetch";
import { ExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

// Use a dedicated wallet that holds only what you're willing to spend.
const account = privateKeyToAccount(process.env.AGENT_WALLET_KEY as `0x${string}`);

const payFetch = wrapFetchWithPaymentFromConfig(fetch, {
  schemes: [{ network: "eip155:8453", client: new ExactEvmScheme(account) }], // Base
});

const PORTAL = "https://agent.pocket.network/v1";
const SEARCH_URL = `${PORTAL}/agentsearch-web-search-v1/v1/search`;
const EXTRACT_URL = `${PORTAL}/agentsearch-web-extract-v1/v1/extract`;
// Render isn't on the Agentic Portal yet (listing coming soon). Set RENDER_URL to
// the POST /v1/render endpoint you reach agentsearch-web-render-v1 through,
// e.g. a Pocket relay or your own AgentSearch-compatible backend.
const RENDER_URL = process.env.RENDER_URL;

async function portalPost<T>(url: string, body: unknown): Promise<T> {
  const res = await payFetch(url, {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify(body),
  });
  if (!res.ok) throw new Error(`HTTP ${res.status} from ${url}`);
  const envelope = await res.json();
  // The portal wraps supplier output: { portal: {...}, data: {...} }.
  if (envelope.portal?.provenance !== "third-party-supplier") {
    throw new Error("Unexpected response shape");
  }
  return envelope.data as T; // third-party content: treat as data, never instructions
}

Next, the heuristic that decides whether to step up. Thin output on its own isn't enough to trigger Render: a short page can be legitimately short. We look for a near-empty body or the telltale strings of a client-rendered shell.

const MIN_CHARS = 400;
const JS_SHELL_PATTERNS = [
  /enable javascript/i,
  /javascript is (required|disabled)/i,
  /you need to enable javascript/i,
  /\.\.\.\s*$/i,
];

function needsRender(extract: { markdown?: string; error: any }): boolean {
  if (extract.error) {
    // Don't render what a browser can't fix: PDFs, DNS failures, 4xx pages.
    return false;
  }
  const md = (extract.markdown ?? "").trim();
  if (md.length < MIN_CHARS) return true;
  return md.length < 2000 && JS_SHELL_PATTERNS.some((re) => re.test(md));
}

Finally, the loop. It searches, extracts the first few results, steps up to Render only for pages that need it, and stops adding context once a token budget is spent.

const approxTokens = (s: string) => Math.ceil(s.length / 4); // rough heuristic

type Source = { url: string; title: string | null; markdown: string; via: "extract" | "render" };

export async function gatherContext(query: string, tokenBudget = 6000): Promise<Source[]> {
  // 1. Search: up to 5 candidates
  const search = await portalPost<any>(SEARCH_URL, { query, max_results: 5 });
  if (search.error) throw new Error(`search failed: ${search.error.code}`);

  const sources: Source[] = [];
  let used = 0;

  for (const result of search.results.slice(0, 3)) {
    const remaining = tokenBudget - used;
    if (remaining < 300) break; // budget spent
    const maxChars = Math.min(remaining * 4, 20000);

    // 2. Extract: cheap and fast, clutter stripped
    const ex = await portalPost<any>(EXTRACT_URL, {
      url: result.url,
      formats: ["markdown"],
      max_chars: maxChars,
      include_links: false,
    });

    let markdown = ex.error ? "" : (ex.markdown ?? "");
    let title = ex.title ?? result.title;
    let via: Source["via"] = "extract";

    // 3. Render: only when Extract came back empty or as a JS shell
    if (needsRender(ex) && RENDER_URL) {
      const res = await fetch(RENDER_URL, {
        method: "POST",
        headers: { "content-type": "application/json" },
        body: JSON.stringify({ url: result.url, formats: ["markdown"], max_chars: maxChars }),
      });
      const r = await res.json();
      // TARGET_DEADLINE still returns partial content, so keep it if present.
      if (r.markdown && (!r.error || r.error.code === "TARGET_DEADLINE")) {
        markdown = r.markdown;
        title = r.title ?? title;
        via = "render";
      }
    }

    if (!markdown.trim()) continue; // nothing usable; move on

    // 4. Cap: hard-trim to whatever budget is left
    const clipped = markdown.slice(0, (tokenBudget - used) * 4);
    used += approxTokens(clipped);
    sources.push({ url: result.url, title, markdown: clipped, via });
  }
  return sources;
}

A few design notes:

needsRender returns sources to your model inside a tool-result or quoted block, never concatenated into the system prompt. The 4-characters-per-token ratio is only an approximation. Swap in your model's tokenizer if you need exact counts.

If you'd rather not write the loop yourself, the Pocket Agentic Portal MCP server exposes these services to any MCP host. It has three tools:

Tool What it does Cost
search_services Search the service catalogue free
describe_service Show a service's price, operations, schemas and a captured example free
call_service Call a service and pay its price in USDC the service's price

Add this block to claude_desktop_config.json (Claude Desktop), .cursor/mcp.json (Cursor) or .mcp.json (Claude Code):

{
  "mcpServers": {
    "pocket-network": {
      "command": "npx",
      "args": ["-y", "@pocket-network/agentic-portal-mcp"],
      "env": {
        "POCKET_PRIVATE_KEY": "0x…",
        "POCKET_MAX_TOTAL_ATOMIC": "1000000",
        "POCKET_MAX_PER_CALL_ATOMIC": "5000"
      }
    }
  }
}

Restart the client after editing. Settings must go in the env block, because desktop clients don't pass your shell environment to MCP servers. Amounts are in USDC atomic units (6 decimals): 1000000 is a $1.00 total cap for the session, and 5000 is $0.005 per call. With no key set, only the two free tools work, which is a safe way to try it out.

Your agent then calls the services through call_service:

{
  "serviceId": "agentsearch-web-search-v1",
  "path": "/v1/search",
  "httpMethod": "POST",
  "body": { "query": "headless chromium robots.txt policy", "max_results": 5 }
}
{
  "serviceId": "agentsearch-web-extract-v1",
  "path": "/v1/extract",
  "httpMethod": "POST",
  "body": { "url": "https://example.com/", "formats": ["markdown"], "max_chars": 8000 }
}

To get the decision logic above, put it in your system prompt: "Search first. Extract the first results. Only if the markdown is empty or a JavaScript shell, flag the page for rendering."

Render isn't available through the Agentic Portal MCP yet, so for now that step goes through your own code path. If you run your own AgentSearch-compatible API, the self-hosted @agentsearchhq/agentsearch-mcp server (npx -y @agentsearchhq/agentsearch-mcp, with AGENTSEARCH_BASE_URL pointing at your API) exposes all three steps as agentsearch_web_search, agentsearch_extract and agentsearch_render.

One caution from the package docs: your MCP client may not ask before paying. The spend limits in env are the controls that hold no matter what the model does, so use a dedicated wallet funded with only what you're willing to spend.

More on wiring search into MCP hosts: MCP web search for AI agents.

Search and Extract each cost $0.005 per call on the Pocket Agentic Portal (Search, Extract). You pay in USDC with either:

There's no account, no API key and no minimum commitment. An unpaid request gets a 402 with the payment terms, your client signs, and the request is served. If the supplier fails to deliver, the portal settles nothing, so a failed call doesn't cost you (details in Pocket's integration guide). Everything runs on Pocket Network's decentralized infrastructure.

Render is live on Pocket Network, and its Agentic Portal listing and price are coming soon.

If your agent needs… Call When
Sources for a question Search Always first
Page content for context or RAG Extract For every candidate URL
Content that only exists after JavaScript runs Render Only when Extract is empty or a JS shell

Three small tools with clear jobs give your agent something one browse call can't: it can see why a step failed, choose the next one, and keep its context window under control.

Add npx -y @pocket-network/agentic-portal-mcp to your MCP client, or call the portal endpoints from your own loop. Start with Web Search and Web Extract, and keep an eye on Web Render for the step-up path. More guides are on the AgentSearch blog.

── more in #ai-agents 4 stories · sorted by recency
── more on @pocket network 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/search-extract-rende…] indexed:0 read:10min 2026-10-05 · —