{"slug": "building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and", "title": "Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright", "summary": "A developer designed a self-hosted article-to-Markdown extraction microservice using FastAPI, Trafilatura, Readability-lxml, and an asynchronous Playwright fallback to convert raw web pages into clean Markdown for LLM ingestion. The tiered extraction waterfall lets roughly 85% of standard web content pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero. The service caches results in Redis and returns Markdown with parsed page metadata.", "body_md": "Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.\n\nA standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.\n\nHere is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.\n\nMost teams start with simple libraries like `BeautifulSoup` or `newspaper3k`. These quickly break down:\n\n`<div id=\"root\"></div>` shells to standard HTTP clients.\nTo balance latency, compute cost, and reliability, the optimal architecture uses a **tiered extraction waterfall**:\n\n`trafilatura`. Latency: ~150-300ms.` readability-lxml` algorithm.\n\n```\n[Incoming Request: URL]\n          │\n          ▼\n┌───────────────────┐\n│ Redis Cache Check │ ──(Hit)──► Return Markdown & Metadata\n└───────────────────┘\n          │ (Miss)\n          ▼\n┌───────────────────┐\n│ HTTPX Async Fetch │\n└───────────────────┘\n          │\n    [Static HTML]\n          │\n          ▼\n┌───────────────────┐\n│    Trafilatura    │ ──(Success: Length > Threshold)──► Parse Meta & Return\n└───────────────────┘\n          │ (Failed / Empty)\n          ▼\n┌───────────────────┐\n│ Readability-lxml  │ ──(Success: Length > Threshold)──► Parse Meta & Return\n└───────────────────┘\n          │ (Failed / SPA Detected)\n          ▼\n┌───────────────────┐\n│ Playwright Worker │ ──(Render DOM)──► Re-extract via Trafilatura\n└───────────────────┘\n          │\n          ▼\n[Store in Cache & Return Markdown]\n```\n\nThis setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.\n\nBelow is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.\n\n``` python\nfrom pydantic import BaseModel, HttpUrl, Field\nfrom typing import Optional, Dict, Any\nfrom datetime import datetime\n\nclass ExtractionRequest(BaseModel):\n    url: HttpUrl\n    force_playwright: bool = Field(default=False, description=\"Bypass fast path and force browser rendering\")\n    max_chars: Optional[int] = Field(default=None, description=\"Truncate body content for token limits\")\n\nclass PageMetadata(BaseModel):\n    title: Optional[str] = None\n    author: Optional[str] = None\n    published_date: Optional[str] = None\n    site_name: Optional[str] = None\n    reading_time_minutes: int = 0\n    opengraph: Dict[str, Any] = {}\n\nclass ExtractionResponse(BaseModel):\n    url: str\n    markdown: str\n    metadata: PageMetadata\n    engine_used: str\n    execution_time_ms: float\npython\nimport time\nimport httpx\nimport trafilatura\nfrom readability import Document\nfrom bs4 import BeautifulSoup\nfrom markdownify import markdownify as md\nfrom playwright.async_api import async_playwright\n\nUSER_AGENT = \"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36\"\n\nasync def fetch_html_fast(url: str) -> str:\n    async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:\n        response = await client.get(url, headers={\"User-Agent\": USER_AGENT})\n        response.raise_for_status()\n        return response.text\n\nasync def fetch_html_playwright(url: str) -> str:\n    async with async_playwright() as p:\n        browser = await p.chromium.launch(headless=True, args=[\"--no-sandbox\", \"--disable-dev-shm-usage\"])\n        context = await browser.new_context(user_agent=USER_AGENT)\n        page = await context.new_page()\n        await page.goto(url, wait_until=\"networkidle\", timeout=20000)\n        content = await page.content()\n        await browser.close()\n        return content\n\ndef extract_metadata(html: str) -> PageMetadata:\n    soup = BeautifulSoup(html, \"lxml\")\n    og_data = {}\n    for tag in soup.find_all(\"meta\"):\n        prop = tag.get(\"property\", tag.get(\"name\", \"\"))\n        if prop.startswith(\"og:\") or prop.startswith(\"twitter:\"):\n            og_data[prop] = tag.get(\"content\", \"\")\n\n    title = og_data.get(\"og:title\") or (soup.title.string if soup.title else None)\n    author = og_data.get(\"article:author\") or og_data.get(\"twitter:creator\")\n    pub_date = og_data.get(\"article:published_time\")\n\n    return PageMetadata(\n        title=title,\n        author=author,\n        published_date=pub_date,\n        site_name=og_data.get(\"og:site_name\"),\n        opengraph=og_data\n    )\n\ndef extract_content(html: str, url: str) -> tuple[str, str]:\n    # Primary Engine: Trafilatura\n    extracted = trafilatura.extract(\n        html,\n        url=url,\n        output_format=\"markdown\",\n        include_links=True,\n        include_images=False,\n        favor_recall=False\n    )\n    if extracted and len(extracted.strip()) > 200:\n        return extracted, \"trafilatura\"\n\n    # Fallback Engine: Readability + Markdownify\n    doc = Document(html)\n    summary_html = doc.summary()\n    markdown_output = md(summary_html, heading_style=\"ATX\").strip()\n\n    if len(markdown_output) > 100:\n        return markdown_output, \"readability\"\n\n    return \"\", \"none\"\npython\nfrom fastapi import FastAPI, HTTPException, status\n\napp = FastAPI(title=\"Article-to-Markdown Extraction API\", version=\"1.0.0\")\n\n@app.post(\"/api/v1/extract\", response_model=ExtractionResponse)\nasync def extract_article(payload: ExtractionRequest):\n    start_time = time.perf_counter()\n    url_str = str(payload.url)\n    engine_used = \"trafilatura\"\n\n    try:\n        if payload.force_playwright:\n            html = await fetch_html_playwright(url_str)\n            markdown, engine_used = extract_content(html, url_str)\n            engine_used = f\"playwright+{engine_used}\"\n        else:\n            # Tier 1 & 2: Fast HTTP path\n            try:\n                html = await fetch_html_fast(url_str)\n                markdown, engine_used = extract_content(html, url_str)\n            except Exception:\n                markdown = \"\"\n\n            # Tier 3: SPA / Fallback if empty\n            if not markdown or len(markdown.strip()) < 150:\n                html = await fetch_html_playwright(url_str)\n                markdown, engine_used = extract_content(html, url_str)\n                engine_used = f\"playwright_fallback+{engine_used}\"\n\n        if not markdown:\n            raise HTTPException(\n                status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,\n                detail=\"Failed to extract meaningful content from the target URL.\"\n            )\n\n        metadata = extract_metadata(html)\n        words = len(markdown.split())\n        metadata.reading_time_minutes = max(1, round(words / 200))\n\n        if payload.max_chars and len(markdown) > payload.max_chars:\n            markdown = markdown[:payload.max_chars] + \"\\n\\n[Content Truncated]\"\n\n        exec_duration = (time.perf_counter() - start_time) * 1000\n\n        return ExtractionResponse(\n            url=url_str,\n            markdown=markdown,\n            metadata=metadata,\n            engine_used=engine_used,\n            execution_time_ms=round(exec_duration, 2)\n        )\n\n    except HTTPException:\n        raise\n    except Exception as e:\n        raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))\n```\n\nWhen exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:\n\nLaunching a new Chromium instance via `async_playwright` on every request will instantly lead to Linux OOM crashes. In a production build:\n\n`BrowserContext` sessions per request.`--disable-dev-shm-usage` and `--no-sandbox` flags inside your Docker container.\nArticles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:\n\n``` php\nimport hashlib\n\ndef generate_cache_key(url: str) -> str:\n    normalized = url.split(\"?\")[0].rstrip(\"/\").lower()\n    return f\"extract:{hashlib.sha256(normalized.encode()).hexdigest()}\"\nFROM python:3.11-slim\n\nENV PYTHONUNBUFFERED=1 \\\n    PIP_NO_CACHE_DIR=1 \\\n    PLAYWRIGHT_BROWSERS_PATH=/ms-playwright\n\nWORKDIR /app\nRUN apt-get update && apt-get install -y --no-install-recommends \\\n    libxml2-dev libxslt-dev gcc libc-dev curl \\\n    && rm -rf /var/lib/apt/lists/*\n\nCOPY requirements.txt .\nRUN pip install -r requirements.txt\nRUN playwright install --with-deps chromium\n\nCOPY . .\nEXPOSE 8000\nCMD [\"uvicorn\", \"main:app\", \"--host\", \"0.0.0.0\", \"--port\", \"8000\", \"--workers\", \"2\"]\n```\n\nYou can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.\n\nIf you want the complete, production-ready implementation out of the box—including:\n\nYou can grab the complete source repository directly:\n\n`EARLYBIRD`", "url": "https://wpnews.pro/news/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and", "canonical_source": "https://dev.to/reigen/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-fastapi-and-playwright-9de", "published_at": "2026-10-05 08:36:29+00:00", "updated_at": "2026-10-05 08:49:56.127990+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools", "ai-infrastructure", "large-language-models", "mlops"], "entities": ["FastAPI", "Playwright", "Trafilatura", "Readability-lxml", "Redis", "HTTPX", "BeautifulSoup", "Chromium"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and", "markdown": "https://wpnews.pro/news/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and.md", "text": "https://wpnews.pro/news/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and.txt", "jsonld": "https://wpnews.pro/news/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-and.jsonld"}}