Turn Any URL into Clean Markdown for RAG (API Guide) A developer published an API guide for converting arbitrary URLs into clean, RAG-ready markdown using AgentSearch Web Extract (agentsearch-web-extract-v1), a paid endpoint that returns headings, paragraphs, lists and links plus citation metadata. Each call costs $0.005 in USDC via x402 or MPP on the Pocket Agentic Portal, requires no account or API key, and exposes options for formats, max_chars, include_links and a hard 4-second fetch deadline. The guide covers the POST /v1/extract request, a sample response, chunking on markdown headings, and error handling. Raw HTML makes poor context for an LLM. A typical page is mostly navigation, scripts, cookie banners and footer links. Putting all of that into an embedding model or a prompt wastes tokens and lowers retrieval quality. What you want is the content : headings, paragraphs, lists and links, as markdown, plus enough metadata to cite the source and to tell when it goes stale. This guide uses AgentSearch Web Extract https://agentsearchhq.com/extract?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag agentsearch-web-extract-v1 to turn a URL into RAG-ready markdown. It covers the request, the response, chunking and error handling. Each call costs $0.005, paid in USDC over x402 or MPP on the Pocket Agentic Portal, with no account and no API key. , tell you where sections begin, which gives you natural chunk boundaries and section titles to store as metadata. text url keeps the citation trail, so an agent can follow a link or quote its source. The endpoint is POST /v1/extract , which takes a JSON body. Only url is required. curl -X POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract \ -H 'content-type: application/json' \ -d '{"url":"https://example.com"}' → 402 Payment Required with the terms. An x402 or MPP client signs and retries. These are the options from the extract OpenAPI spec https://agentsearchhq.com/specs/extract-openapi.json?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag : | Field | Type | Default | Notes | |---|---|---|---| | url | string | required | An absolute http s URL | | formats | "markdown" , "text" or both | "markdown" | Which renderings to return | | max chars | integer, 1–100000 | 50000 | Caps the returned characters. meta.truncated tells you whether the cap was hit | | include links | boolean | true | Up to 100 absolute links from the page | | timeout ms | integer | 4000 | Fetch deadline. Values above 4000 are capped a hard 4 s deadline | This is a real response captured on the portal service page 2026-09-25 UTC : { "portal": { "provenance": "third-party-supplier", "serviceId": "agentsearch-web-extract-v1", "schemaCheck": "unchecked" }, "data": { "request id": "05f5a4ebe2684c79867aa9bf3515c33d", "url": "https://example.com/", "title": "Example Domain", "markdown": " Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n", "text": "", "links": { "href": "https://iana.org/domains/example", "text": "Learn more" } , "meta": { "status code": 200, "content type": "text/html", "fetched at": "2026-09-25T15:53:07.843243+00:00", "chars": 167, "truncated": false }, "error": null } } The fields that matter for RAG: url is the title is a ready-made citation label. meta.fetched at lets you re-extract stale documents on a schedule. meta.truncated tells you whether you hit max chars . If you did, raise the cap and fetch again. error is null on success. Standard x402 v2 clients work with the portal unchanged. Here's @x402/fetch with @x402/evm , paying on Base: js import { wrapFetchWithPaymentFromConfig } from "@x402/fetch"; import { ExactEvmScheme } from "@x402/evm"; import { privateKeyToAccount } from "viem/accounts"; const EXTRACT = "https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract"; // Use a dedicated wallet that holds only what you're willing to spend. const account = privateKeyToAccount process.env.AGENT WALLET KEY as 0x${string} ; const pay = wrapFetchWithPaymentFromConfig fetch, { schemes: { network: "eip155:8453", client: new ExactEvmScheme account } , } ; export async function extract url: string, maxChars = 50000 { const res = await pay EXTRACT, { method: "POST", headers: { "content-type": "application/json" }, body: JSON.stringify { url, formats: "markdown" , max chars: maxChars } , } ; if res.ok { // Portal errors aren't wrapped in the envelope: { error: { code, message, retryable } } const body = await res.json .catch = {} ; throw new Error portal ${res.status}: ${body?.error?.code ?? "unknown"} ; } const envelope = await res.json ; if envelope.portal?.provenance == "third-party-supplier" { throw new Error "Unexpected response shape; refusing to use it." ; } const doc = envelope.data; if doc.error { // The target site failed, but the HTTP status was still 200. throw Object.assign new Error doc.error.code , { retryable: doc.error.retryable } ; } return doc as { url: string; title: string | null; markdown: string; links: { href: string; text: string } ; meta: { fetched at: string; chars: number; truncated: boolean }; }; } If delivery fails, the portal settles nothing, so you aren't charged for it. Payment details are in Pocket's integration guide https://agent.pocket.network/agent-integration . In Claude Desktop, Cursor or Claude Code, add npx -y @pocket-network/agentic-portal-mcp and ask the agent to call service on agentsearch-web-extract-v1 . The full setup, including spending limits, is in How to give your AI agent web search with MCP https://agentsearchhq.com/blog/mcp-web-search-for-ai-agents?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag . Split on markdown headings first, then on size. Every chunk keeps the path of headings above it, which improves both retrieval and citations. python import re def chunk markdown doc: dict, max chars: int = 1500, overlap: int = 150 : """doc is the data object returned by agentsearch-web-extract-v1.""" chunks, path, buf = , , def flush : text = "\n".join buf .strip if not text: return for i in range 0, len text , max chars - overlap : chunks.append { "text": text i:i + max chars , "source url": doc "url" , "title": doc.get "title" , "section": " ".join path , "fetched at": doc "meta" "fetched at" , } for line in doc "markdown" .splitlines : m = re.match r"^ {1,6} \s+ . ", line if m: flush ; buf.clear level = len m.group 1 path : = path :level - 1 + m.group 2 .strip buf.append line flush return chunks Embed chunk "text" and store the rest as metadata. When your agent answers, cite title and source url . Extract separates target-site failures from request problems : | Where | HTTP | Codes | What to do | |---|---|---|---| | data.error target site | 200 | TARGET HTTP ERROR , TARGET TIMEOUT , TARGET DNS ERROR , TARGET CONNECT ERROR , TARGET FETCH FAILED , TARGET TOO MANY REDIRECTS , UNSUPPORTED CONTENT | Retry only when retryable is true timeouts, connection failures, target 5xx/429 | | Request | 400 | INVALID REQUEST , SSRF BLOCKED private, loopback or metadata addresses | Fix the input. Don't retry | | Request | 413 / 429 | REQUEST TOO LARGE , CAPACITY LIMIT , APPLICATION RATE LIMITED | Back off and retry | UNSUPPORTED CONTENT means the URL points to something other than HTML or text, such as a PDF or an image. Send those to a document parser instead. Any page can contain text written to manipulate your agent. Before content reaches a model: portal.schemaCheck . Only passed means a schema check ran. Extract fetches the HTML the server sends, which is fast and works for docs, blogs, news and most marketing pages. Single-page apps and dashboards often render their content only after JavaScript runs. For those, AgentSearch has Web Render https://agentsearchhq.com/render?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag agentsearch-web-render-v1 , live on Pocket MainNet. It loads the page in headless Chromium and returns the rendered markdown, text, HTML, links and an optional screenshot. It honors robots.txt and doesn't bypass bot walls. Render isn't on the Agentic Portal yet. Its request and response contract is in agents.md https://agentsearchhq.com/agents.md?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag and the render spec https://agentsearchhq.com/specs/render-openapi.json?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag . A common pattern is to search with AgentSearch Web Search https://agentsearchhq.com/search?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag up to 5 results per call , pick the most relevant results, then extract them in full. See the search spec https://agentsearchhq.com/specs/search-openapi.json?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag , and compare per-call costs across vendors in Cheapest web search API for AI agents in 2026 https://agentsearchhq.com/blog/cheapest-web-search-api-for-ai-agents-2026?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag . How long can an extracted page be? Up to 100,000 characters per call through max chars . The default is 50,000. Does it follow redirects? Yes. url in the response is the final URL, and requested url appears on errors. Can it extract PDFs? No. Non-HTML content returns UNSUPPORTED CONTENT . Who serves the requests? Decentralized Pocket Network suppliers. Machine-readable details are at agentsearchhq.com/llms.txt https://agentsearchhq.com/llms.txt?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag . Service pages: Web Search API https://agentsearchhq.com/search?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag · Web Extract API https://agentsearchhq.com/extract?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag · Web Render API https://agentsearchhq.com/render?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag Originally published at agentsearchhq.com https://agentsearchhq.com/blog/url-to-markdown-for-rag?utm source=devto&utm medium=social&utm campaign=url-to-markdown-for-rag .