# Building a High-Throughput Article-to-Markdown API for LLM Ingestion with FastAPI and Playwright

> Source: <https://dev.to/reigen/building-a-high-throughput-article-to-markdown-api-for-llm-ingestion-with-fastapi-and-playwright-9de>
> Published: 2026-10-05 08:36:29+00:00

Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.

A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.

Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.

Most teams start with simple libraries like `BeautifulSoup` or `newspaper3k`. These quickly break down:

`<div id="root"></div>` shells to standard HTTP clients.
To balance latency, compute cost, and reliability, the optimal architecture uses a **tiered extraction waterfall**:

`trafilatura`. Latency: ~150-300ms.` readability-lxml` algorithm.

```
[Incoming Request: URL]
          │
          ▼
┌───────────────────┐
│ Redis Cache Check │ ──(Hit)──► Return Markdown & Metadata
└───────────────────┘
          │ (Miss)
          ▼
┌───────────────────┐
│ HTTPX Async Fetch │
└───────────────────┘
          │
    [Static HTML]
          │
          ▼
┌───────────────────┐
│    Trafilatura    │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
          │ (Failed / Empty)
          ▼
┌───────────────────┐
│ Readability-lxml  │ ──(Success: Length > Threshold)──► Parse Meta & Return
└───────────────────┘
          │ (Failed / SPA Detected)
          ▼
┌───────────────────┐
│ Playwright Worker │ ──(Render DOM)──► Re-extract via Trafilatura
└───────────────────┘
          │
          ▼
[Store in Cache & Return Markdown]
```

This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.

Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.

``` python
from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, Any
from datetime import datetime

class ExtractionRequest(BaseModel):
    url: HttpUrl
    force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering")
    max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits")

class PageMetadata(BaseModel):
    title: Optional[str] = None
    author: Optional[str] = None
    published_date: Optional[str] = None
    site_name: Optional[str] = None
    reading_time_minutes: int = 0
    opengraph: Dict[str, Any] = {}

class ExtractionResponse(BaseModel):
    url: str
    markdown: str
    metadata: PageMetadata
    engine_used: str
    execution_time_ms: float
python
import time
import httpx
import trafilatura
from readability import Document
from bs4 import BeautifulSoup
from markdownify import markdownify as md
from playwright.async_api import async_playwright

USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"

async def fetch_html_fast(url: str) -> str:
    async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:
        response = await client.get(url, headers={"User-Agent": USER_AGENT})
        response.raise_for_status()
        return response.text

async def fetch_html_playwright(url: str) -> str:
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"])
        context = await browser.new_context(user_agent=USER_AGENT)
        page = await context.new_page()
        await page.goto(url, wait_until="networkidle", timeout=20000)
        content = await page.content()
        await browser.close()
        return content

def extract_metadata(html: str) -> PageMetadata:
    soup = BeautifulSoup(html, "lxml")
    og_data = {}
    for tag in soup.find_all("meta"):
        prop = tag.get("property", tag.get("name", ""))
        if prop.startswith("og:") or prop.startswith("twitter:"):
            og_data[prop] = tag.get("content", "")

    title = og_data.get("og:title") or (soup.title.string if soup.title else None)
    author = og_data.get("article:author") or og_data.get("twitter:creator")
    pub_date = og_data.get("article:published_time")

    return PageMetadata(
        title=title,
        author=author,
        published_date=pub_date,
        site_name=og_data.get("og:site_name"),
        opengraph=og_data
    )

def extract_content(html: str, url: str) -> tuple[str, str]:
    # Primary Engine: Trafilatura
    extracted = trafilatura.extract(
        html,
        url=url,
        output_format="markdown",
        include_links=True,
        include_images=False,
        favor_recall=False
    )
    if extracted and len(extracted.strip()) > 200:
        return extracted, "trafilatura"

    # Fallback Engine: Readability + Markdownify
    doc = Document(html)
    summary_html = doc.summary()
    markdown_output = md(summary_html, heading_style="ATX").strip()

    if len(markdown_output) > 100:
        return markdown_output, "readability"

    return "", "none"
python
from fastapi import FastAPI, HTTPException, status

app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0")

@app.post("/api/v1/extract", response_model=ExtractionResponse)
async def extract_article(payload: ExtractionRequest):
    start_time = time.perf_counter()
    url_str = str(payload.url)
    engine_used = "trafilatura"

    try:
        if payload.force_playwright:
            html = await fetch_html_playwright(url_str)
            markdown, engine_used = extract_content(html, url_str)
            engine_used = f"playwright+{engine_used}"
        else:
            # Tier 1 & 2: Fast HTTP path
            try:
                html = await fetch_html_fast(url_str)
                markdown, engine_used = extract_content(html, url_str)
            except Exception:
                markdown = ""

            # Tier 3: SPA / Fallback if empty
            if not markdown or len(markdown.strip()) < 150:
                html = await fetch_html_playwright(url_str)
                markdown, engine_used = extract_content(html, url_str)
                engine_used = f"playwright_fallback+{engine_used}"

        if not markdown:
            raise HTTPException(
                status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
                detail="Failed to extract meaningful content from the target URL."
            )

        metadata = extract_metadata(html)
        words = len(markdown.split())
        metadata.reading_time_minutes = max(1, round(words / 200))

        if payload.max_chars and len(markdown) > payload.max_chars:
            markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]"

        exec_duration = (time.perf_counter() - start_time) * 1000

        return ExtractionResponse(
            url=url_str,
            markdown=markdown,
            metadata=metadata,
            engine_used=engine_used,
            execution_time_ms=round(exec_duration, 2)
        )

    except HTTPException:
        raise
    except Exception as e:
        raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))
```

When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:

Launching a new Chromium instance via `async_playwright` on every request will instantly lead to Linux OOM crashes. In a production build:

`BrowserContext` sessions per request.`--disable-dev-shm-usage` and `--no-sandbox` flags inside your Docker container.
Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:

``` php
import hashlib

def generate_cache_key(url: str) -> str:
    normalized = url.split("?")[0].rstrip("/").lower()
    return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}"
FROM python:3.11-slim

ENV PYTHONUNBUFFERED=1 \
    PIP_NO_CACHE_DIR=1 \
    PLAYWRIGHT_BROWSERS_PATH=/ms-playwright

WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
    libxml2-dev libxslt-dev gcc libc-dev curl \
    && rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install -r requirements.txt
RUN playwright install --with-deps chromium

COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
```

You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.

If you want the complete, production-ready implementation out of the box—including:

You can grab the complete source repository directly:

`EARLYBIRD`
