cd /news/ai-tools/my-url-to-markdown-api-served-a-nav-… · home › topics › ai-tools › article
[ARTICLE · art-143145] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

My URL-to-Markdown API served a nav menu as an 'article' — detecting JS-rendered SPA shell pages

A developer hardened a URL-to-Markdown extraction API after auditing request logs and finding that roughly one in six URLs returned under 60 words of valid-looking Markdown — navigation menus and footers from JavaScript-rendered SPA shell pages — silently poisoning downstream RAG pipelines. The fix adds a low_content flag with word count when extraction falls below a threshold or shows a high boilerplate ratio, a fallback ladder to og:description and meta description labeled as summaries, and a CSS selector escape hatch, while deliberately skipping headless Chromium due to its memory and latency cost. Useful-content yield on a 500-URL sample rose from about 83% to 96% after the flag shipped.

by read2 min views1 publishedOct 1, 2026

My URL-to-Markdown API began as a scrappy readability pass: fetch a URL, strip nav/ads/sidebars, hand back clean Markdown. It worked beautifully on every blog, docs site, and news article I tested. Then I audited my own request logs and found the failure mode nobody complains about — because it doesn't error.

About 1 in 6 URLs came back under 60 words. Not errors — 200 OK, valid Markdown, technically correct output. The "content" was Skip to content / Home / About / Sign in plus a footer. I had been silently embedding navigation menus into people's RAG pipelines. Valid output, zero information, and worse than an exception because nothing downstream noticed.

The cause: SPA shells. The server returns real HTML with a mostly-empty root div; the article is assembled client-side. My extraction pass had no actual content to find, so it carved boilerplate into something shaped like an article.

What I changed, in order of impact:

Extraction floor. If word count falls below a threshold or the boilerplate ratio is too high, return an explicit low_content flag with the word count instead of confident garbage. Biggest single win — it made failure visible.

Fallback ladder. On a shell page, try og:description and meta description, clearly labeled as a summary. An honest 40-word summary beats 40 words of menu pretending to be prose.

Escape hatch. A CSS selector option, so callers who know a site can target its article container directly. Print views and RSS endpoints often carry the full text where the main page doesn't.

What I deliberately didn't do: add headless Chromium. I benchmarked it — hundreds of MB of RAM per concurrent render and p95 latency from ~800ms to 6s+. For my traffic, loud cheap failure beat expensive half-rendering. I'll revisit if shell pages ever become the majority.

After the flag shipped, my retry logic routes shell pages elsewhere instead of embedding empty strings, and useful-content yield on a 500-URL sample went from roughly 83% to 96%. Honest limit: pure client-rendered pages with no SSR and no meta tags will never extract well. The point isn't extracting more — it's never lying about what came back.

I ended up packaging the hardened version as my URL-to-Markdown API: https://x402.freeq.one/tools/markdown.html

── more in #ai-tools 4 stories · sorted by recency
── more on @x402.freeq.one 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/my-url-to-markdown-a…] indexed:0 read:2min 2026-10-01 · —