# From a URL to Markdown: a first collection and MCP setup with DeepSFetch

> Source: <https://dev.to/rafael_souza_2f0438f4d82e/from-a-url-to-markdown-a-first-collection-and-mcp-setup-with-deepsfetch-4jdn>
> Published: 2026-10-11 19:28:16+00:00

I'm building DeepSFetch, a hosted service for collecting web content and connecting it to applications and AI tools. We have published a developer kit with client libraries, CLI examples, and documentation in English and Portuguese.

This tutorial walks through one bounded request, retrieves the completed Markdown, and explains how to connect the hosted MCP server. The developer kit is public; the hosted collection backend is private.

Create an account at [DeepSFetch](https://deepsfetch.com/login?tab=signup&utm_source=devto&utm_medium=tutorial&utm_campaign=developer_kit_launch), check the selected workspace and its available credits, then open **API Keys** and create a key. API and MCP calls require this key, including calls from a free account.

Keep it in your local environment. Never paste it into a GitHub issue, public MCP configuration, or a comment.

```
export DEEPSFETCH_API_KEY="YOUR_WORKSPACE_KEY"
```

Windows PowerShell:

```
$env:DEEPSFETCH_API_KEY = "YOUR_WORKSPACE_KEY"
```

The current collection engine is selected explicitly with `collectionEngine: "auto"`. A valid collection costs one DeepSFetch credit on this path, including its browser attempt when used. The legacy collection path has different pricing and response shapes.

Save the following as `first-collection.mjs`. It uses Node.js 20+ and built-in APIs, so you do not need to install a package to try it.

``` js
import { randomUUID } from 'node:crypto';

const key = process.env.DEEPSFETCH_API_KEY;
if (!key) throw new Error('Set DEEPSFETCH_API_KEY first');

const base = 'https://api.deepsfetch.com';
const headers = {
  Authorization: `Bearer ${key}`,
  'Content-Type': 'application/json',
};

async function read(response) {
  const data = await response.json();
  if (!response.ok) {
    throw new Error(`HTTP ${response.status}: ${data.code ?? 'request failed'}`);
  }
  return data;
}

const requestId = randomUUID();
console.error('Keep this request ID for submission recovery:', requestId);
let job = await read(await fetch(`${base}/v2/scrape`, {
  method: 'POST',
  headers: { ...headers, 'Idempotency-Key': requestId },
  body: JSON.stringify({
    url: 'https://quotes.toscrape.com/',
    formats: ['markdown'],
    collectionEngine: 'auto',
    engineStrategy: 'auto',
    async: true,
    costControl: {
      maxCredits: 1,
      allowFallback: false,
      cacheMode: 'fresh',
    },
  }),
  signal: AbortSignal.timeout(30000),
}));

console.error('Save this job ID:', job.id);
const deadline = Date.now() + 180000;
while (['queued', 'running'].includes(job.status)) {
  if (Date.now() >= deadline) {
    throw new Error(`Local timeout. Retrieve the same job later: ${job.id}`);
  }
  await new Promise(resolve => setTimeout(resolve, 2000));
  job = await read(await fetch(`${base}/v2/jobs/${encodeURIComponent(job.id)}`, {
    headers,
    signal: AbortSignal.timeout(30000),
  }));
}

if (job.status !== 'completed' || typeof job.result?.markdown !== 'string') {
  throw new Error(`Inspect job ${job.id}: ${job.status}`);
}
console.error('Billing:', job.result.billing);
console.log(job.result.markdown);
```

Run it:

```
node first-collection.mjs
```

The first response can mean **accepted**, rather than completed. Polling retrieves that same job; it does not submit another scrape. If your terminal closes, keep the printed job ID and retrieve it with `GET /v2/jobs/{id}` using the same workspace key.

Inspect the content as well as the status. For this target, check that the Markdown contains actual quotes and authors. A completed job alone does not prove that a page was fully captured.

The cap is expressed in DeepSFetch credits, not dollars or another provider's credit units. A premium offer is separate and requires explicit approval; this example does not approve premium collection.

The hosted endpoint is `https://mcp.deepsfetch.com/mcp`, using Streamable HTTP and a Bearer workspace key. For VS Code versions supporting remote HTTP servers, put the following in `.vscode/mcp.json`:

```
{
  "inputs": [
    {
      "type": "promptString",
      "id": "deepsfetch-api-key",
      "description": "DeepSFetch workspace API key",
      "password": true
    }
  ],
  "servers": {
    "deepsfetch": {
      "type": "http",
      "url": "https://mcp.deepsfetch.com/mcp",
      "headers": {
        "Authorization": "Bearer ${input:deepsfetch-api-key}"
      }
    }
  }
}
```

List the tools before issuing a request: the live input schemas matter. Start by asking the agent to collect one public page and inspect its job status.

**The current MCP 2.0.1 scrape schema uses the legacy collection path.** It does not expose the automatic engine selector used in the HTTP example above. A basic MCP scrape can use `plannedLane: "http_basic"` with a three-credit cap. It may return artifact references rather than inline Markdown. Do not assume that the HTTP and MCP examples have identical costs or outputs.

The [MCP guide](https://deepsfetch.com/docs/mcp?lang=en) describes each tool and connection options. The server is also listed in the [official MCP registry](https://registry.modelcontextprotocol.io/v0.1/servers/io.github.ermultimidia%2Fdeepsfetch/versions/2.0.1).

The automatic HTTP engine has a 512 KiB HTML limit. Its supported formats include HTML, Markdown, links, and screenshots; structured field extraction is a separate operation. Custom headers, cookies, action sequences, and several browser tuning options are not supported on this engine path.

We are testing page quality and latency against other collectors. This post does not claim universal coverage or parity with Firecrawl. JavaScript detection and content cleanup are areas we are actively checking.

The [scrape reference](https://deepsfetch.com/docs/scrape?lang=en) and [troubleshooting guide](https://deepsfetch.com/docs/troubleshooting?lang=en) document path-specific options, failures, and recovery.

The [public developer kit](https://github.com/ermultimidia/deepsfetch-developer-kit) includes JavaScript and Python clients, CLI examples, and a service-by-service manual.

I'm looking for developers who will try one real workflow and share:

You can reply here or join the [first-user feedback discussion](https://github.com/ermultimidia/deepsfetch-developer-kit/discussions/1). Share a public reproducible target when possible; keep credentials and private data out of feedback.

Disclosure: this article was generated with an AI coding assistant from the implementation, documentation, and live checks performed for the project.
