Feeding raw HTML into LLM context windows is one of the most expensive and inefficient mistakes in modern AI engineering.
A standard modern news or blog page easily spans 1.5MB to 4MB of raw DOM payload. When passed straight into an LLM or vector database, 90% of those tokens are spent on tracking scripts, serialized JSON-LD blobs, cookie banners, navigation menus, and inline CSS styles. This not only causes severe context bloat and escalates inference bills, but it also degrades retrieval-augmented generation (RAG) semantic search precision by polluting your vector space with boilerplate noise.
Here is how to design and build an enterprise-grade, self-hosted extraction microservice using FastAPI, Trafilatura, Readability, and an asynchronous Playwright fallback for SPA rendering.
Most teams start with simple libraries like BeautifulSoup or newspaper3k. These quickly break down:
<div id="root"></div> shells to standard HTTP clients.
To balance latency, compute cost, and reliability, the optimal architecture uses a tiered extraction waterfall:
trafilatura. Latency: ~150-300ms. readability-lxml algorithm.
[Incoming Request: URL]
β
βΌ
βββββββββββββββββββββ
β Redis Cache Check β ββ(Hit)βββΊ Return Markdown & Metadata
βββββββββββββββββββββ
β (Miss)
βΌ
βββββββββββββββββββββ
β HTTPX Async Fetch β
βββββββββββββββββββββ
β
[Static HTML]
β
βΌ
βββββββββββββββββββββ
β Trafilatura β ββ(Success: Length > Threshold)βββΊ Parse Meta & Return
βββββββββββββββββββββ
β (Failed / Empty)
βΌ
βββββββββββββββββββββ
β Readability-lxml β ββ(Success: Length > Threshold)βββΊ Parse Meta & Return
βββββββββββββββββββββ
β (Failed / SPA Detected)
βΌ
βββββββββββββββββββββ
β Playwright Worker β ββ(Render DOM)βββΊ Re-extract via Trafilatura
βββββββββββββββββββββ
β
βΌ
[Store in Cache & Return Markdown]
This setup allows 85% of standard web content to pass through the lightweight Python tier without launching Chromium, keeping memory usage stable and infrastructure costs near zero.
Below is the complete implementation of the dual-engine pipeline using FastAPI, Pydantic v2, and async execution.
from pydantic import BaseModel, HttpUrl, Field
from typing import Optional, Dict, Any
from datetime import datetime
class ExtractionRequest(BaseModel):
url: HttpUrl
force_playwright: bool = Field(default=False, description="Bypass fast path and force browser rendering")
max_chars: Optional[int] = Field(default=None, description="Truncate body content for token limits")
class PageMetadata(BaseModel):
title: Optional[str] = None
author: Optional[str] = None
published_date: Optional[str] = None
site_name: Optional[str] = None
reading_time_minutes: int = 0
opengraph: Dict[str, Any] = {}
class ExtractionResponse(BaseModel):
url: str
markdown: str
metadata: PageMetadata
engine_used: str
execution_time_ms: float
python
import time
import httpx
import trafilatura
from readability import Document
from bs4 import BeautifulSoup
from markdownify import markdownify as md
from playwright.async_api import async_playwright
USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36"
async def fetch_html_fast(url: str) -> str:
async with httpx.AsyncClient(timeout=10.0, follow_redirects=True) as client:
response = await client.get(url, headers={"User-Agent": USER_AGENT})
response.raise_for_status()
return response.text
async def fetch_html_playwright(url: str) -> str:
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True, args=["--no-sandbox", "--disable-dev-shm-usage"])
context = await browser.new_context(user_agent=USER_AGENT)
page = await context.new_page()
await page.goto(url, wait_until="networkidle", timeout=20000)
content = await page.content()
await browser.close()
return content
def extract_metadata(html: str) -> PageMetadata:
soup = BeautifulSoup(html, "lxml")
og_data = {}
for tag in soup.find_all("meta"):
prop = tag.get("property", tag.get("name", ""))
if prop.startswith("og:") or prop.startswith("twitter:"):
og_data[prop] = tag.get("content", "")
title = og_data.get("og:title") or (soup.title.string if soup.title else None)
author = og_data.get("article:author") or og_data.get("twitter:creator")
pub_date = og_data.get("article:published_time")
return PageMetadata(
title=title,
author=author,
published_date=pub_date,
site_name=og_data.get("og:site_name"),
opengraph=og_data
)
def extract_content(html: str, url: str) -> tuple[str, str]:
extracted = trafilatura.extract(
html,
url=url,
output_format="markdown",
include_links=True,
include_images=False,
favor_recall=False
)
if extracted and len(extracted.strip()) > 200:
return extracted, "trafilatura"
doc = Document(html)
summary_html = doc.summary()
markdown_output = md(summary_html, heading_style="ATX").strip()
if len(markdown_output) > 100:
return markdown_output, "readability"
return "", "none"
python
from fastapi import FastAPI, HTTPException, status
app = FastAPI(title="Article-to-Markdown Extraction API", version="1.0.0")
@app.post("/api/v1/extract", response_model=ExtractionResponse)
async def extract_article(payload: ExtractionRequest):
start_time = time.perf_counter()
url_str = str(payload.url)
engine_used = "trafilatura"
try:
if payload.force_playwright:
html = await fetch_html_playwright(url_str)
markdown, engine_used = extract_content(html, url_str)
engine_used = f"playwright+{engine_used}"
else:
try:
html = await fetch_html_fast(url_str)
markdown, engine_used = extract_content(html, url_str)
except Exception:
markdown = ""
if not markdown or len(markdown.strip()) < 150:
html = await fetch_html_playwright(url_str)
markdown, engine_used = extract_content(html, url_str)
engine_used = f"playwright_fallback+{engine_used}"
if not markdown:
raise HTTPException(
status_code=status.HTTP_422_UNPROCESSABLE_ENTITY,
detail="Failed to extract meaningful content from the target URL."
)
metadata = extract_metadata(html)
words = len(markdown.split())
metadata.reading_time_minutes = max(1, round(words / 200))
if payload.max_chars and len(markdown) > payload.max_chars:
markdown = markdown[:payload.max_chars] + "\n\n[Content Truncated]"
exec_duration = (time.perf_counter() - start_time) * 1000
return ExtractionResponse(
url=url_str,
markdown=markdown,
metadata=metadata,
engine_used=engine_used,
execution_time_ms=round(exec_duration, 2)
)
except HTTPException:
raise
except Exception as e:
raise HTTPException(status_code=status.HTTP_500_INTERNAL_SERVER_ERROR, detail=str(e))
When exposing this service in production pipelines (e.g., connecting n8n webhooks or feeding continuous Celery queues), consider these three safeguards:
Launching a new Chromium instance via async_playwright on every request will instantly lead to Linux OOM crashes. In a production build:
BrowserContext sessions per request.--disable-dev-shm-usage and --no-sandbox flags inside your Docker container.
Articles rarely mutate hourly. Adding an upstream Redis cache using a SHA-256 hash of the normalized canonical URL drastically drops response times to sub-10ms for repeated requests:
import hashlib
def generate_cache_key(url: str) -> str:
normalized = url.split("?")[0].rstrip("/").lower()
return f"extract:{hashlib.sha256(normalized.encode()).hexdigest()}"
FROM python:3.11-slim
ENV PYTHONUNBUFFERED=1 \
PIP_NO_CACHE_DIR=1 \
PLAYWRIGHT_BROWSERS_PATH=/ms-playwright
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
libxml2-dev libxslt-dev gcc libc-dev curl \
&& rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install -r requirements.txt
RUN playwright install --with-deps chromium
COPY . .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000", "--workers", "2"]
You can manually wire this FastAPI microservice into your infrastructure following the architectural design and code above.
If you want the complete, production-ready implementation out of the boxβincluding:
You can grab the complete source repository directly:
EARLYBIRD