{"slug": "from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch", "title": "From a URL to Markdown: a first collection and MCP setup with DeepSFetch", "summary": "A developer building DeepSFetch, a hosted web-content collection service, published a developer kit with client libraries, CLI examples, and English/Portuguese documentation, and walked through a first scrape that returns Markdown via the /v2/scrape API using an idempotency key, async job polling, and a one-credit cost cap. The writeup also covers connecting the hosted MCP server at mcp.deepsfetch.com/mcp over Streamable HTTP with a Bearer workspace key, and warns that a completed job status alone does not prove a page was fully captured.", "body_md": "I'm building DeepSFetch, a hosted service for collecting web content and connecting it to applications and AI tools. We have published a developer kit with client libraries, CLI examples, and documentation in English and Portuguese.\n\nThis tutorial walks through one bounded request, retrieves the completed Markdown, and explains how to connect the hosted MCP server. The developer kit is public; the hosted collection backend is private.\n\nCreate an account at [DeepSFetch](https://deepsfetch.com/login?tab=signup&utm_source=devto&utm_medium=tutorial&utm_campaign=developer_kit_launch), check the selected workspace and its available credits, then open **API Keys** and create a key. API and MCP calls require this key, including calls from a free account.\n\nKeep it in your local environment. Never paste it into a GitHub issue, public MCP configuration, or a comment.\n\n```\nexport DEEPSFETCH_API_KEY=\"YOUR_WORKSPACE_KEY\"\n```\n\nWindows PowerShell:\n\n```\n$env:DEEPSFETCH_API_KEY = \"YOUR_WORKSPACE_KEY\"\n```\n\nThe current collection engine is selected explicitly with `collectionEngine: \"auto\"`. A valid collection costs one DeepSFetch credit on this path, including its browser attempt when used. The legacy collection path has different pricing and response shapes.\n\nSave the following as `first-collection.mjs`. It uses Node.js 20+ and built-in APIs, so you do not need to install a package to try it.\n\n``` js\nimport { randomUUID } from 'node:crypto';\n\nconst key = process.env.DEEPSFETCH_API_KEY;\nif (!key) throw new Error('Set DEEPSFETCH_API_KEY first');\n\nconst base = 'https://api.deepsfetch.com';\nconst headers = {\n  Authorization: `Bearer ${key}`,\n  'Content-Type': 'application/json',\n};\n\nasync function read(response) {\n  const data = await response.json();\n  if (!response.ok) {\n    throw new Error(`HTTP ${response.status}: ${data.code ?? 'request failed'}`);\n  }\n  return data;\n}\n\nconst requestId = randomUUID();\nconsole.error('Keep this request ID for submission recovery:', requestId);\nlet job = await read(await fetch(`${base}/v2/scrape`, {\n  method: 'POST',\n  headers: { ...headers, 'Idempotency-Key': requestId },\n  body: JSON.stringify({\n    url: 'https://quotes.toscrape.com/',\n    formats: ['markdown'],\n    collectionEngine: 'auto',\n    engineStrategy: 'auto',\n    async: true,\n    costControl: {\n      maxCredits: 1,\n      allowFallback: false,\n      cacheMode: 'fresh',\n    },\n  }),\n  signal: AbortSignal.timeout(30000),\n}));\n\nconsole.error('Save this job ID:', job.id);\nconst deadline = Date.now() + 180000;\nwhile (['queued', 'running'].includes(job.status)) {\n  if (Date.now() >= deadline) {\n    throw new Error(`Local timeout. Retrieve the same job later: ${job.id}`);\n  }\n  await new Promise(resolve => setTimeout(resolve, 2000));\n  job = await read(await fetch(`${base}/v2/jobs/${encodeURIComponent(job.id)}`, {\n    headers,\n    signal: AbortSignal.timeout(30000),\n  }));\n}\n\nif (job.status !== 'completed' || typeof job.result?.markdown !== 'string') {\n  throw new Error(`Inspect job ${job.id}: ${job.status}`);\n}\nconsole.error('Billing:', job.result.billing);\nconsole.log(job.result.markdown);\n```\n\nRun it:\n\n```\nnode first-collection.mjs\n```\n\nThe first response can mean **accepted**, rather than completed. Polling retrieves that same job; it does not submit another scrape. If your terminal closes, keep the printed job ID and retrieve it with `GET /v2/jobs/{id}` using the same workspace key.\n\nInspect the content as well as the status. For this target, check that the Markdown contains actual quotes and authors. A completed job alone does not prove that a page was fully captured.\n\nThe cap is expressed in DeepSFetch credits, not dollars or another provider's credit units. A premium offer is separate and requires explicit approval; this example does not approve premium collection.\n\nThe hosted endpoint is `https://mcp.deepsfetch.com/mcp`, using Streamable HTTP and a Bearer workspace key. For VS Code versions supporting remote HTTP servers, put the following in `.vscode/mcp.json`:\n\n```\n{\n  \"inputs\": [\n    {\n      \"type\": \"promptString\",\n      \"id\": \"deepsfetch-api-key\",\n      \"description\": \"DeepSFetch workspace API key\",\n      \"password\": true\n    }\n  ],\n  \"servers\": {\n    \"deepsfetch\": {\n      \"type\": \"http\",\n      \"url\": \"https://mcp.deepsfetch.com/mcp\",\n      \"headers\": {\n        \"Authorization\": \"Bearer ${input:deepsfetch-api-key}\"\n      }\n    }\n  }\n}\n```\n\nList the tools before issuing a request: the live input schemas matter. Start by asking the agent to collect one public page and inspect its job status.\n\n**The current MCP 2.0.1 scrape schema uses the legacy collection path.** It does not expose the automatic engine selector used in the HTTP example above. A basic MCP scrape can use `plannedLane: \"http_basic\"` with a three-credit cap. It may return artifact references rather than inline Markdown. Do not assume that the HTTP and MCP examples have identical costs or outputs.\n\nThe [MCP guide](https://deepsfetch.com/docs/mcp?lang=en) describes each tool and connection options. The server is also listed in the [official MCP registry](https://registry.modelcontextprotocol.io/v0.1/servers/io.github.ermultimidia%2Fdeepsfetch/versions/2.0.1).\n\nThe automatic HTTP engine has a 512 KiB HTML limit. Its supported formats include HTML, Markdown, links, and screenshots; structured field extraction is a separate operation. Custom headers, cookies, action sequences, and several browser tuning options are not supported on this engine path.\n\nWe are testing page quality and latency against other collectors. This post does not claim universal coverage or parity with Firecrawl. JavaScript detection and content cleanup are areas we are actively checking.\n\nThe [scrape reference](https://deepsfetch.com/docs/scrape?lang=en) and [troubleshooting guide](https://deepsfetch.com/docs/troubleshooting?lang=en) document path-specific options, failures, and recovery.\n\nThe [public developer kit](https://github.com/ermultimidia/deepsfetch-developer-kit) includes JavaScript and Python clients, CLI examples, and a service-by-service manual.\n\nI'm looking for developers who will try one real workflow and share:\n\nYou can reply here or join the [first-user feedback discussion](https://github.com/ermultimidia/deepsfetch-developer-kit/discussions/1). Share a public reproducible target when possible; keep credentials and private data out of feedback.\n\nDisclosure: this article was generated with an AI coding assistant from the implementation, documentation, and live checks performed for the project.", "url": "https://wpnews.pro/news/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch", "canonical_source": "https://dev.to/rafael_souza_2f0438f4d82e/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch-4jdn", "published_at": "2026-10-11 19:28:16+00:00", "updated_at": "2026-10-11 19:32:19.375721+00:00", "lang": "en", "topics": ["ai-tools", "agent-protocols", "developer-tools", "ai-agents"], "entities": ["DeepSFetch", "MCP", "Node.js", "VS Code", "quotes.toscrape.com"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch", "markdown": "https://wpnews.pro/news/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch.md", "text": "https://wpnews.pro/news/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch.txt", "jsonld": "https://wpnews.pro/news/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch.jsonld"}}