My URL-to-Markdown API began as a scrappy readability pass: fetch a URL, strip nav/ads/sidebars, hand back clean Markdown. It worked beautifully on every blog, docs site, and news article I tested. Then I audited my own request logs and found the failure mode nobody complains about — because it doesn't error.
About 1 in 6 URLs came back under 60 words. Not errors — 200 OK, valid Markdown, technically correct output. The "content" was Skip to content / Home / About / Sign in plus a footer. I had been silently embedding navigation menus into people's RAG pipelines. Valid output, zero information, and worse than an exception because nothing downstream noticed.
The cause: SPA shells. The server returns real HTML with a mostly-empty root div; the article is assembled client-side. My extraction pass had no actual content to find, so it carved boilerplate into something shaped like an article.
What I changed, in order of impact:
Extraction floor. If word count falls below a threshold or the boilerplate ratio is too high, return an explicit low_content flag with the word count instead of confident garbage. Biggest single win — it made failure visible.
Fallback ladder. On a shell page, try og:description and meta description, clearly labeled as a summary. An honest 40-word summary beats 40 words of menu pretending to be prose.
Escape hatch. A CSS selector option, so callers who know a site can target its article container directly. Print views and RSS endpoints often carry the full text where the main page doesn't.
What I deliberately didn't do: add headless Chromium. I benchmarked it — hundreds of MB of RAM per concurrent render and p95 latency from ~800ms to 6s+. For my traffic, loud cheap failure beat expensive half-rendering. I'll revisit if shell pages ever become the majority.
After the flag shipped, my retry logic routes shell pages elsewhere instead of embedding empty strings, and useful-content yield on a 500-URL sample went from roughly 83% to 96%. Honest limit: pure client-rendered pages with no SSR and no meta tags will never extract well. The point isn't extracting more — it's never lying about what came back.
I ended up packaging the hardened version as my URL-to-Markdown API: https://x402.freeq.one/tools/markdown.html