cd /news/ai-tools/from-a-url-to-markdown-a-first-colle… · home › topics › ai-tools › article
[ARTICLE · art-149276] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

From a URL to Markdown: a first collection and MCP setup with DeepSFetch

A developer building DeepSFetch, a hosted web-content collection service, published a developer kit with client libraries, CLI examples, and English/Portuguese documentation, and walked through a first scrape that returns Markdown via the /v2/scrape API using an idempotency key, async job polling, and a one-credit cost cap. The writeup also covers connecting the hosted MCP server at mcp.deepsfetch.com/mcp over Streamable HTTP with a Bearer workspace key, and warns that a completed job status alone does not prove a page was fully captured.

by read4 min views1 publishedOct 11, 2026

I'm building DeepSFetch, a hosted service for collecting web content and connecting it to applications and AI tools. We have published a developer kit with client libraries, CLI examples, and documentation in English and Portuguese.

This tutorial walks through one bounded request, retrieves the completed Markdown, and explains how to connect the hosted MCP server. The developer kit is public; the hosted collection backend is private.

Create an account at DeepSFetch, check the selected workspace and its available credits, then open API Keys and create a key. API and MCP calls require this key, including calls from a free account.

Keep it in your local environment. Never paste it into a GitHub issue, public MCP configuration, or a comment.

export DEEPSFETCH_API_KEY="YOUR_WORKSPACE_KEY"

Windows PowerShell:

$env:DEEPSFETCH_API_KEY = "YOUR_WORKSPACE_KEY"

The current collection engine is selected explicitly with collectionEngine: "auto". A valid collection costs one DeepSFetch credit on this path, including its browser attempt when used. The legacy collection path has different pricing and response shapes.

Save the following as first-collection.mjs. It uses Node.js 20+ and built-in APIs, so you do not need to install a package to try it.

import { randomUUID } from 'node:crypto';

const key = process.env.DEEPSFETCH_API_KEY;
if (!key) throw new Error('Set DEEPSFETCH_API_KEY first');

const base = 'https://api.deepsfetch.com';
const headers = {
  Authorization: `Bearer ${key}`,
  'Content-Type': 'application/json',
};

async function read(response) {
  const data = await response.json();
  if (!response.ok) {
    throw new Error(`HTTP ${response.status}: ${data.code ?? 'request failed'}`);
  }
  return data;
}

const requestId = randomUUID();
console.error('Keep this request ID for submission recovery:', requestId);
let job = await read(await fetch(`${base}/v2/scrape`, {
  method: 'POST',
  headers: { ...headers, 'Idempotency-Key': requestId },
  body: JSON.stringify({
    url: 'https://quotes.toscrape.com/',
    formats: ['markdown'],
    collectionEngine: 'auto',
    engineStrategy: 'auto',
    async: true,
    costControl: {
      maxCredits: 1,
      allowFallback: false,
      cacheMode: 'fresh',
    },
  }),
  signal: AbortSignal.timeout(30000),
}));

console.error('Save this job ID:', job.id);
const deadline = Date.now() + 180000;
while (['queued', 'running'].includes(job.status)) {
  if (Date.now() >= deadline) {
    throw new Error(`Local timeout. Retrieve the same job later: ${job.id}`);
  }
  await new Promise(resolve => setTimeout(resolve, 2000));
  job = await read(await fetch(`${base}/v2/jobs/${encodeURIComponent(job.id)}`, {
    headers,
    signal: AbortSignal.timeout(30000),
  }));
}

if (job.status !== 'completed' || typeof job.result?.markdown !== 'string') {
  throw new Error(`Inspect job ${job.id}: ${job.status}`);
}
console.error('Billing:', job.result.billing);
console.log(job.result.markdown);

Run it:

node first-collection.mjs

The first response can mean accepted, rather than completed. Polling retrieves that same job; it does not submit another scrape. If your terminal closes, keep the printed job ID and retrieve it with GET /v2/jobs/{id} using the same workspace key.

Inspect the content as well as the status. For this target, check that the Markdown contains actual quotes and authors. A completed job alone does not prove that a page was fully captured.

The cap is expressed in DeepSFetch credits, not dollars or another provider's credit units. A premium offer is separate and requires explicit approval; this example does not approve premium collection.

The hosted endpoint is https://mcp.deepsfetch.com/mcp, using Streamable HTTP and a Bearer workspace key. For VS Code versions supporting remote HTTP servers, put the following in .vscode/mcp.json:

{
  "inputs": [
    {
      "type": "promptString",
      "id": "deepsfetch-api-key",
      "description": "DeepSFetch workspace API key",
      "password": true
    }
  ],
  "servers": {
    "deepsfetch": {
      "type": "http",
      "url": "https://mcp.deepsfetch.com/mcp",
      "headers": {
        "Authorization": "Bearer ${input:deepsfetch-api-key}"
      }
    }
  }
}

List the tools before issuing a request: the live input schemas matter. Start by asking the agent to collect one public page and inspect its job status.

The current MCP 2.0.1 scrape schema uses the legacy collection path. It does not expose the automatic engine selector used in the HTTP example above. A basic MCP scrape can use plannedLane: "http_basic" with a three-credit cap. It may return artifact references rather than inline Markdown. Do not assume that the HTTP and MCP examples have identical costs or outputs.

The MCP guide describes each tool and connection options. The server is also listed in the official MCP registry.

The automatic HTTP engine has a 512 KiB HTML limit. Its supported formats include HTML, Markdown, links, and screenshots; structured field extraction is a separate operation. Custom headers, cookies, action sequences, and several browser tuning options are not supported on this engine path.

We are testing page quality and latency against other collectors. This post does not claim universal coverage or parity with Firecrawl. JavaScript detection and content cleanup are areas we are actively checking.

The scrape reference and troubleshooting guide document path-specific options, failures, and recovery.

The public developer kit includes JavaScript and Python clients, CLI examples, and a service-by-service manual.

I'm looking for developers who will try one real workflow and share:

You can reply here or join the first-user feedback discussion. Share a public reproducible target when possible; keep credentials and private data out of feedback.

Disclosure: this article was generated with an AI coding assistant from the implementation, documentation, and live checks performed for the project.

── more in #ai-tools 4 stories · sorted by recency
── more on @deepsfetch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/from-a-url-to-markdo…] indexed:0 read:4min 2026-10-11 · —