{"slug": "turn-any-url-into-clean-markdown-for-rag-api-guide", "title": "Turn Any URL into Clean Markdown for RAG (API Guide)", "summary": "A developer published an API guide for converting arbitrary URLs into clean, RAG-ready markdown using AgentSearch Web Extract (agentsearch-web-extract-v1), a paid endpoint that returns headings, paragraphs, lists and links plus citation metadata. Each call costs $0.005 in USDC via x402 or MPP on the Pocket Agentic Portal, requires no account or API key, and exposes options for formats, max_chars, include_links and a hard 4-second fetch deadline. The guide covers the POST /v1/extract request, a sample response, chunking on markdown headings, and error handling.", "body_md": "Raw HTML makes poor context for an LLM. A typical page is mostly navigation, scripts, cookie banners and footer links. Putting all of that into an embedding model or a prompt wastes tokens and lowers retrieval quality. What you want is the **content**: headings, paragraphs, lists and links, as markdown, plus enough metadata to cite the source and to tell when it goes stale.\n\nThis guide uses [AgentSearch Web Extract](https://agentsearchhq.com/extract?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) (`agentsearch-web-extract-v1`) to turn a URL into RAG-ready markdown. It covers the request, the response, chunking and error handling. Each call costs $0.005, paid in USDC over x402 or MPP on the Pocket Agentic Portal, with no account and no API key.\n\n`#`, `##`) tell you where sections begin, which gives you natural chunk boundaries and section titles to store as metadata.`[text](url)` keeps the citation trail, so an agent can follow a link or quote its source.\nThe endpoint is `POST /v1/extract`, which takes a JSON body. Only `url` is required.\n\n```\ncurl -X POST https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract \\\n  -H 'content-type: application/json' \\\n  -d '{\"url\":\"https://example.com\"}'\n# → 402 Payment Required with the terms. An x402 or MPP client signs and retries.\n```\n\nThese are the options from the [extract OpenAPI spec](https://agentsearchhq.com/specs/extract-openapi.json?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag):\n\n| Field | Type | Default | Notes | \n|---|---|---|---|\n| `url` | string | (required) | An absolute http(s) URL | \n| `formats` | `[\"markdown\"]` ,`[\"text\"]` or both | `[\"markdown\"]` | Which renderings to return | \n| `max_chars` | integer, 1–100000 | 50000 | Caps the returned characters. `meta.truncated` tells you whether the cap was hit | \n| `include_links` | boolean | true | Up to 100 absolute links from the page | \n| `timeout_ms` | integer | 4000 | Fetch deadline. Values above 4000 are capped (a hard 4 s deadline) | \n\nThis is a real response captured on the portal service page (2026-09-25 UTC):\n\n```\n{\n  \"portal\": {\n    \"provenance\": \"third-party-supplier\",\n    \"serviceId\": \"agentsearch-web-extract-v1\",\n    \"schemaCheck\": \"unchecked\"\n  },\n  \"data\": {\n    \"request_id\": \"05f5a4ebe2684c79867aa9bf3515c33d\",\n    \"url\": \"https://example.com/\",\n    \"title\": \"Example Domain\",\n    \"markdown\": \"# Example Domain\\n\\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\\n\",\n    \"text\": \"\",\n    \"links\": [{ \"href\": \"https://iana.org/domains/example\", \"text\": \"Learn more\" }],\n    \"meta\": {\n      \"status_code\": 200,\n      \"content_type\": \"text/html\",\n      \"fetched_at\": \"2026-09-25T15:53:07.843243+00:00\",\n      \"chars\": 167,\n      \"truncated\": false\n    },\n    \"error\": null\n  }\n}\n```\n\nThe fields that matter for RAG:\n\n`url` is the `title` is a ready-made citation label.`meta.fetched_at` lets you re-extract stale documents on a schedule.`meta.truncated` tells you whether you hit `max_chars`. If you did, raise the cap and fetch again.` error` is `null` on success.\nStandard x402 v2 clients work with the portal unchanged. Here's `@x402/fetch` with `@x402/evm`, paying on Base:\n\n``` js\nimport { wrapFetchWithPaymentFromConfig } from \"@x402/fetch\";\nimport { ExactEvmScheme } from \"@x402/evm\";\nimport { privateKeyToAccount } from \"viem/accounts\";\n\nconst EXTRACT =\n  \"https://agent.pocket.network/v1/agentsearch-web-extract-v1/v1/extract\";\n\n// Use a dedicated wallet that holds only what you're willing to spend.\nconst account = privateKeyToAccount(process.env.AGENT_WALLET_KEY as `0x${string}`);\nconst pay = wrapFetchWithPaymentFromConfig(fetch, {\n  schemes: [{ network: \"eip155:8453\", client: new ExactEvmScheme(account) }],\n});\n\nexport async function extract(url: string, maxChars = 50000) {\n  const res = await pay(EXTRACT, {\n    method: \"POST\",\n    headers: { \"content-type\": \"application/json\" },\n    body: JSON.stringify({ url, formats: [\"markdown\"], max_chars: maxChars }),\n  });\n\n  if (!res.ok) {\n    // Portal errors aren't wrapped in the envelope: { error: { code, message, retryable } }\n    const body = await res.json().catch(() => ({}));\n    throw new Error(`portal ${res.status}: ${body?.error?.code ?? \"unknown\"}`);\n  }\n\n  const envelope = await res.json();\n  if (envelope.portal?.provenance !== \"third-party-supplier\") {\n    throw new Error(\"Unexpected response shape; refusing to use it.\");\n  }\n  const doc = envelope.data;\n  if (doc.error) {\n    // The target site failed, but the HTTP status was still 200.\n    throw Object.assign(new Error(doc.error.code), { retryable: doc.error.retryable });\n  }\n  return doc as {\n    url: string; title: string | null; markdown: string;\n    links: { href: string; text: string }[];\n    meta: { fetched_at: string; chars: number; truncated: boolean };\n  };\n}\n```\n\nIf delivery fails, the portal settles nothing, so you aren't charged for it. Payment details are in Pocket's [integration guide](https://agent.pocket.network/agent-integration).\n\nIn Claude Desktop, Cursor or Claude Code, add `npx -y @pocket-network/agentic-portal-mcp` and ask the agent to `call_service` on `agentsearch-web-extract-v1`. The full setup, including spending limits, is in [How to give your AI agent web search with MCP](https://agentsearchhq.com/blog/mcp-web-search-for-ai-agents?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag).\n\nSplit on markdown headings first, then on size. Every chunk keeps the path of headings above it, which improves both retrieval and citations.\n\n``` python\nimport re\n\ndef chunk_markdown(doc: dict, max_chars: int = 1500, overlap: int = 150):\n    \"\"\"doc is the `data` object returned by agentsearch-web-extract-v1.\"\"\"\n    chunks, path, buf = [], [], []\n\n    def flush():\n        text = \"\\n\".join(buf).strip()\n        if not text:\n            return\n        for i in range(0, len(text), max_chars - overlap):\n            chunks.append({\n                \"text\": text[i:i + max_chars],\n                \"source_url\": doc[\"url\"],\n                \"title\": doc.get(\"title\"),\n                \"section\": \" > \".join(path),\n                \"fetched_at\": doc[\"meta\"][\"fetched_at\"],\n            })\n\n    for line in doc[\"markdown\"].splitlines():\n        m = re.match(r\"^(#{1,6})\\s+(.*)\", line)\n        if m:\n            flush(); buf.clear()\n            level = len(m.group(1))\n            path[:] = path[:level - 1] + [m.group(2).strip()]\n        buf.append(line)\n    flush()\n    return chunks\n```\n\nEmbed `chunk[\"text\"]` and store the rest as metadata. When your agent answers, cite `title` and `source_url`.\n\nExtract separates **target-site failures** from **request problems**:\n\n| Where | HTTP | Codes | What to do | \n|---|---|---|---|\n| `data.error` (target site) | 200 | `TARGET_HTTP_ERROR` ,`TARGET_TIMEOUT` ,`TARGET_DNS_ERROR` ,`TARGET_CONNECT_ERROR` ,`TARGET_FETCH_FAILED` ,`TARGET_TOO_MANY_REDIRECTS` ,`UNSUPPORTED_CONTENT` | Retry only when `retryable` is true (timeouts, connection failures, target 5xx/429) | \n| Request | 400 | `INVALID_REQUEST` ,`SSRF_BLOCKED` (private, loopback or metadata addresses) | Fix the input. Don't retry | \n| Request | 413 / 429 | `REQUEST_TOO_LARGE` ,`CAPACITY_LIMIT` ,`APPLICATION_RATE_LIMITED` | Back off and retry | \n\n`UNSUPPORTED_CONTENT` means the URL points to something other than HTML or text, such as a PDF or an image. Send those to a document parser instead.\n\nAny page can contain text written to manipulate your agent. Before content reaches a model:\n\n`portal.schemaCheck`. Only `passed` means a schema check ran.\nExtract fetches the HTML the server sends, which is fast and works for docs, blogs, news and most marketing pages. Single-page apps and dashboards often render their content only after JavaScript runs. For those, AgentSearch has [**Web Render**](https://agentsearchhq.com/render?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) (`agentsearch-web-render-v1`), live on Pocket MainNet. It loads the page in headless Chromium and returns the rendered markdown, text, HTML, links and an optional screenshot. It honors robots.txt and doesn't bypass bot walls. Render isn't on the Agentic Portal yet. Its request and response contract is in [agents.md](https://agentsearchhq.com/agents.md?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) and the [render spec](https://agentsearchhq.com/specs/render-openapi.json?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag).\n\nA common pattern is to search with [AgentSearch Web Search](https://agentsearchhq.com/search?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) (up to 5 results per call), pick the most relevant results, then extract them in full. See the [search spec](https://agentsearchhq.com/specs/search-openapi.json?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag), and compare per-call costs across vendors in [Cheapest web search API for AI agents in 2026](https://agentsearchhq.com/blog/cheapest-web-search-api-for-ai-agents-2026?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag).\n\n**How long can an extracted page be?**\n\nUp to 100,000 characters per call through `max_chars`. The default is 50,000.\n\n**Does it follow redirects?**\n\nYes. `url` in the response is the final URL, and `requested_url` appears on errors.\n\n**Can it extract PDFs?**\n\nNo. Non-HTML content returns `UNSUPPORTED_CONTENT`.\n\n**Who serves the requests?**\n\nDecentralized Pocket Network suppliers. Machine-readable details are at [agentsearchhq.com/llms.txt](https://agentsearchhq.com/llms.txt?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag).\n\nService pages: [Web Search API](https://agentsearchhq.com/search?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) · [Web Extract API](https://agentsearchhq.com/extract?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag) · [Web Render API](https://agentsearchhq.com/render?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag)\n\nOriginally published at [agentsearchhq.com](https://agentsearchhq.com/blog/url-to-markdown-for-rag?utm_source=devto&utm_medium=social&utm_campaign=url-to-markdown-for-rag).", "url": "https://wpnews.pro/news/turn-any-url-into-clean-markdown-for-rag-api-guide", "canonical_source": "https://dev.to/agentsearchhq/turn-any-url-into-clean-markdown-for-rag-api-guide-40hp", "published_at": "2026-10-02 23:58:53+00:00", "updated_at": "2026-10-03 00:07:43.684001+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "agent-protocols", "developer-tools"], "entities": ["AgentSearch Web Extract", "Pocket Agentic Portal", "x402", "MPP", "@x402/fetch", "Base"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/turn-any-url-into-clean-markdown-for-rag-api-guide", "markdown": "https://wpnews.pro/news/turn-any-url-into-clean-markdown-for-rag-api-guide.md", "text": "https://wpnews.pro/news/turn-any-url-into-clean-markdown-for-rag-api-guide.txt", "jsonld": "https://wpnews.pro/news/turn-any-url-into-clean-markdown-for-rag-api-guide.jsonld"}}