cd /news/ai-tools/building-a-high-throughput-article-t… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-145278] src=dev.to β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright

A developer designed a self-hosted article-to-Markdown extraction microservice using FastAPI, Trafilatura, Readability-lxml, and an asynchronous Playwright fallback to convert raw web pages into clean Markdown for LLM ingestion. The tiered extraction waterfall lets roughly 85% of standard web content pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero. The service caches results in Redis and returns Markdown with parsed page metadata.

by read5 min views2 publishedOct 5, 2026

Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.

A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.

Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.

Most teams start with simple libraries like BeautifulSoup or newspaper3k. These quickly break down:

<div id="root"></div> shells to standard HTTP clients. To balance latency, compute cost, and reliability, the optimal architecture uses a tiered extraction waterfall:

trafilatura. Latency: ~150-300ms. readability-lxml algorithm.

[Incoming Request: URL]
          β”‚
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Redis Cache Check β”‚ ──(Hit)──► Return Markdown & Metadata
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚ (Miss)
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ HTTPX Async Fetch β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚
    [Static HTML]
          β”‚
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚    Trafilatura    β”‚ ──(Success: Length > Threshold)──► Parse Meta & Return
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚ (Failed / Empty)
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Readability-lxml  β”‚ ──(Success: Length > Threshold)──► Parse Meta & Return
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚ (Failed / SPA Detected)
          β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Playwright Worker β”‚ ──(Render DOM)──► Re-extract via Trafilatura
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
          β”‚
          β–Ό
[Store in Cache & Return Markdown]

This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.

Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.

from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, Any
from datetime import datetime

class ExtractionRequest(BaseModel):
    url: HttpUrl
    force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering")
    max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits")

class PageMetadata(BaseModel):
    title: Optional[str] = None
    author: Optional[str] = None
    published_date: Optional[str] = None
    site_name: Optional[str] = None
    reading_time_minutes: int = 0
    opengraph: Dict[str, Any] = {}

class ExtractionResponse(BaseModel):
    url: str
    markdown: str
    metadata: PageMetadata
    engine_used: str
    execution_time_ms: float
python
import time
import httpx
import trafilatura
from readability import Document
from bs4 import BeautifulSoup
from markdownify import markdownify as md
from playwright.async_api import async_playwright

USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"

async def fetch_html_fast(url: str) -> str:
    async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:
        response = await client.get(url, headers={"User-Agent": USER_AGENT})
        response.raise_for_status()
        return response.text

async def fetch_html_playwright(url: str) -> str:
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"])
        context = await browser.new_context(user_agent=USER_AGENT)
        page = await context.new_page()
        await page.goto(url, wait_until="networkidle", timeout=20000)
        content = await page.content()
        await browser.close()
        return content

def extract_metadata(html: str) -> PageMetadata:
    soup = BeautifulSoup(html, "lxml")
    og_data = {}
    for tag in soup.find_all("meta"):
        prop = tag.get("property", tag.get("name", ""))
        if prop.startswith("og:") or prop.startswith("twitter:"):
            og_data[prop] = tag.get("content", "")

    title = og_data.get("og:title") or (soup.title.string if soup.title else None)
    author = og_data.get("article:author") or og_data.get("twitter:creator")
    pub_date = og_data.get("article:published_time")

    return PageMetadata(
        title=title,
        author=author,
        published_date=pub_date,
        site_name=og_data.get("og:site_name"),
        opengraph=og_data
    )

def extract_content(html: str, url: str) -> tuple[str, str]:
    extracted = trafilatura.extract(
        html,
        url=url,
        output_format="markdown",
        include_links=True,
        include_images=False,
        favor_recall=False
    )
    if extracted and len(extracted.strip()) > 200:
        return extracted, "trafilatura"

    doc = Document(html)
    summary_html = doc.summary()
    markdown_output = md(summary_html, heading_style="ATX").strip()

    if len(markdown_output) > 100:
        return markdown_output, "readability"

    return "", "none"
python
from fastapi import FastAPI, HTTPException, status

app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0")

@app.post("/api/v1/extract", response_model=ExtractionResponse)
async def extract_article(payload: ExtractionRequest):
    start_time = time.perf_counter()
    url_str = str(payload.url)
    engine_used = "trafilatura"

    try:
        if payload.force_playwright:
            html = await fetch_html_playwright(url_str)
            markdown, engine_used = extract_content(html, url_str)
            engine_used = f"playwright+{engine_used}"
        else:
            try:
                html = await fetch_html_fast(url_str)
                markdown, engine_used = extract_content(html, url_str)
            except Exception:
                markdown = ""

            if not markdown or len(markdown.strip()) < 150:
                html = await fetch_html_playwright(url_str)
                markdown, engine_used = extract_content(html, url_str)
                engine_used = f"playwright_fallback+{engine_used}"

        if not markdown:
            raise HTTPException(
                status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
                detail="Failed to extract meaningful content from the target URL."
            )

        metadata = extract_metadata(html)
        words = len(markdown.split())
        metadata.reading_time_minutes = max(1, round(words / 200))

        if payload.max_chars and len(markdown) > payload.max_chars:
            markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]"

        exec_duration = (time.perf_counter() - start_time) * 1000

        return ExtractionResponse(
            url=url_str,
            markdown=markdown,
            metadata=metadata,
            engine_used=engine_used,
            execution_time_ms=round(exec_duration, 2)
        )

    except HTTPException:
        raise
    except Exception as e:
        raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))

When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:

Launching a new Chromium instance via async_playwright on every request will instantly lead to Linux OOM crashes. In a production build:

BrowserContext sessions per request.--disable-dev-shm-usage and --no-sandbox flags inside your Docker container. Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:

import hashlib

def generate_cache_key(url: str) -> str:
    normalized = url.split("?")[0].rstrip("/").lower()
    return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}"
FROM python:3.11-slim

ENV PYTHONUNBUFFERED=1 \
    PIP_NO_CACHE_DIR=1 \
    PLAYWRIGHT_BROWSERS_PATH=/ms-playwright

WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
    libxml2-dev libxslt-dev gcc libc-dev curl \
    && rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install -r requirements.txt
RUN playwright install --with-deps chromium

COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]

You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.

If you want the complete, production-ready implementation out of the boxβ€”including:

You can grab the complete source repository directly:

EARLYBIRD

── more in #ai-tools 4 stories Β· sorted by recency
── more on @fastapi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/building-a-high-thro…] indexed:0 read:5min 2026-10-05 Β· β€”