cd /news/developer-tools/build-a-web-scraping-pipeline-for-ll… · home topics developer-tools article
[ARTICLE · art-117644] src=sourcefeed.dev ↗ pub= topic=developer-tools verified=true sentiment=· neutral

Build a Web Scraping Pipeline for LLMs with Crawl4AI

Crawl4AI 0.9.3, released 2026-08-31, enables building a web scraping pipeline that converts pages into clean Markdown for RAG and structured JSON for agents without LLM calls. The tutorial, verified on macOS 26 with Python 3.11 and Playwright 1.62.0, demonstrates crawling quotes.toscrape.com pages 1-3, using a CSS schema to extract quotes, authors, and tags, and writing outputs to Markdown and JSONL files. The pipeline uses a single CrawlerRunConfig for both jobs, ensuring each page is fetched once.

read5 min views1 publishedSep 1, 2026
Build a Web Scraping Pipeline for LLMs with Crawl4AI
Image: Sourcefeed (auto-discovered)

Crawl pages once, get clean Markdown for RAG and structured JSON for agents, with no LLM calls.

Mariana Souza

What you'll build #

A Python script that crawls a list of pages with Crawl4AI, strips navigation and boilerplate, and writes two outputs per run: clean Markdown files you can drop into a RAG index, and a JSONL file of structured records pulled from the same pages with a CSS schema. No LLM calls, no API keys.

Prerequisites #

Verified on macOS 26 (Apple Silicon) with:

  • Python 3.11 (Crawl4AI requires 3.10+; PyPI classifiers go up to 3.13, so avoid 3.14 for now)
  • Crawl4AI 0.9.3 (released 2026-08-31)
  • Playwright 1.62.0, pulled in as a dependency; crawl4ai-setup

downloads Chromium Headless Shell 151 (about 95 MB)

On Linux, Chromium needs system libraries. Run python -m playwright install --with-deps chromium

after the install step if crawl4ai-setup

complains. No accounts needed. The target site, quotes.toscrape.com, exists for scraping practice.

1. Install Crawl4AI and the browser #

python3.11 -m venv .venv && source .venv/bin/activate
pip install -U crawl4ai
crawl4ai-setup
crawl4ai-doctor

crawl4ai-setup

installs the headless browser and initialises a local SQLite cache under ~/.crawl4ai/

. crawl4ai-doctor

does a real test crawl of crawl4ai.com and should end with:

[COMPLETE] ● ✅ Crawling test passed!

2. Write the pipeline #

Save this as pipeline.py

. It uses one CrawlerRunConfig

for both jobs: markdown_generator

produces the cleaned Markdown, extraction_strategy

produces the JSON. Both run against the same fetched HTML, so each page loads once.

import asyncio
import json
from pathlib import Path

from crawl4ai import (
    AsyncWebCrawler,
    BrowserConfig,
    CacheMode,
    CrawlerRunConfig,
    DefaultMarkdownGenerator,
    JsonCssExtractionStrategy,
    PruningContentFilter,
)

URLS = [
    "https://quotes.toscrape.com/page/1/",
    "https://quotes.toscrape.com/page/2/",
    "https://quotes.toscrape.com/page/3/",
]
OUT = Path("out")

SCHEMA = {
    "name": "quotes",
    "baseSelector": "div.quote",
    "fields": [
        {"name": "text", "selector": "span.text", "type": "text"},
        {"name": "author", "selector": "small.author", "type": "text"},
        {
            "name": "tags",
            "selector": "div.tags a.tag",
            "type": "list",
            "fields": [{"name": "tag", "type": "text"}],
        },
    ],
}

async def main():
    OUT.mkdir(exist_ok=True)

    run_cfg = CrawlerRunConfig(
        cache_mode=CacheMode.BYPASS,
        excluded_tags=["nav", "footer", "header", "aside"],
        exclude_external_links=True,
        markdown_generator=DefaultMarkdownGenerator(
            content_filter=PruningContentFilter(threshold=0.45, threshold_type="dynamic"),
            options={"ignore_links": True},
        ),
        extraction_strategy=JsonCssExtractionStrategy(SCHEMA),
    )

    records = []
    async with AsyncWebCrawler(config=BrowserConfig(headless=True)) as crawler:
        results = await crawler.arun_many(URLS, config=run_cfg)
        for r in results:
            if not r.success:
                print(f"[FAIL] {r.url}: {r.error_message}")
                continue
            slug = r.url.rstrip("/").split("/")[-1]
            (OUT / f"{slug}.md").write_text(r.markdown.fit_markdown)
            items = json.loads(r.extracted_content)
            for item in items:
                item["tags"] = [t["tag"] for t in item["tags"]]
                records.append({"source_url": r.url, **item})
            print(f"[OK] {r.url}: {len(r.markdown.fit_markdown)} chars, {len(items)} quotes")

    with (OUT / "quotes.jsonl").open("w") as f:
        for rec in records:
            f.write(json.dumps(rec, ensure_ascii=False) + "\n")
    print(f"Wrote {len(records)} records to {OUT/'quotes.jsonl'}")

if __name__ == "__main__":
    asyncio.run(main())

The cleaning happens in three layers. excluded_tags

drops whole DOM regions before conversion, which is the cheapest win: on page 1 it cut raw Markdown from 4,375 to 1,619 characters. PruningContentFilter

then scores the remaining blocks by text density and link ratio and drops low scorers; its output lands in result.markdown.fit_markdown

, while the unfiltered version stays in raw_markdown

. ignore_links

strips anchor URLs from the Markdown so you don't pay tokens for them.

JsonCssExtractionStrategy

takes a schema: baseSelector

matches one element per record, and each field's selector

is scoped inside it. A list

field returns a list of dicts ([{"tag": "change"}, ...]

), which is why the script flattens it to plain strings before writing.

3. Run it #

python pipeline.py

arun_many

runs the three URLs concurrently in one browser. The whole run takes about 1.5 seconds on a laptop.

Verify it works #

Crawl4AI prints its own [FETCH]

, [SCRAPE]

and [EXTRACT]

progress lines. Below them you should see:

[OK] https://quotes.toscrape.com/page/1/: 1410 chars, 10 quotes
[OK] https://quotes.toscrape.com/page/3/: 1584 chars, 10 quotes
[OK] https://quotes.toscrape.com/page/2/: 3241 chars, 10 quotes
Wrote 30 records to out/quotes.jsonl

Order varies because the pages finish concurrently. Check the outputs:

ls out/
head -c 300 out/1.md
head -1 out/quotes.jsonl
1.md  2.md  3.md  quotes.jsonl
Login
“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.” by Albert Einstein (about)
Tags: change deep-thoughts thinking world
{"source_url": "https://quotes.toscrape.com/page/1/", "text": "“The world as we have created it is a process of our thinking. It cannot be changed without changing our thinking.”", "author": "Albert Einstein", "tags": ["change", "deep-thoughts", "thinking", "world"]}

The Markdown has no nav, footer, or link URLs; the JSONL has one record per quote with tags as a flat list. Each .md

file is ready to chunk and embed, and each JSONL line is ready to load as tool output or a metadata-filtered document.

Troubleshooting #

playwright._impl._errors.Error: BrowserType.launch: Executable doesn't exist at .../chromium_headless_shell-1234/...

You skipped crawl4ai-setup

, or ran it in a different virtualenv. Activate the venv you installed Crawl4AI into and run crawl4ai-setup

again. If it still fails, run python -m playwright install chromium

directly.

Host system is missing dependencies to run browsers

(Linux only) Chromium needs shared libraries that minimal images and CI runners don't ship. Run python -m playwright install --with-deps chromium

(needs sudo on most distros) and retry.

TypeError: the JSON object must be str, bytes or bytearray, not NoneType

result.extracted_content

is None

because no extraction_strategy

was set on the CrawlerRunConfig

you passed. Check you're passing config=run_cfg

to arun_many

, not a fresh config.

fit_markdown

is an empty string fit_markdown

is only populated when DefaultMarkdownGenerator

has a content_filter

. Without one, use result.markdown.raw_markdown

instead.

Next steps #

Tune PruningContentFilter

per site: raise threshold

toward 0.6 on link-heavy pages, or set min_word_threshold

to drop short nodes (on this site a value of 10 also drops the author lines, so test before trusting it). For query-driven trimming, swap in BM25ContentFilter(user_query="...")

from the same module. To discover pages instead of listing them, add a deep_crawl_strategy

such as BFSDeepCrawlStrategy(max_depth=2)

to the run config. And when a site's markup is too messy for CSS selectors, LLMExtractionStrategy

accepts a Pydantic model and does the same job with a model call per page. All of these are documented under docs.crawl4ai.com.

Sources & further reading #

Crawl4AI Installation and Setup— docs.crawl4ai.com - Crawl4AI Markdown Generation Basics— docs.crawl4ai.com - Crawl4AI Extracting JSON (No LLM)— docs.crawl4ai.com - Crawl4AI Multi-URL Crawling— docs.crawl4ai.com - Crawl4AI 0.9.3 on PyPI— pypi.org - Crawl4AI CHANGELOG— github.com

Mariana Souza· Senior Editor

Mariana covers the fast-moving world of machine learning and generative AI, with a particular focus on how these technologies are reshaping development workflows. When she isn't stress-testing the latest foundation models, she's usually at a local hackathon.

Discussion 0 #

No comments yet

Be the first to weigh in.

── more in #developer-tools 4 stories · sorted by recency
── more on @crawl4ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-a-web-scraping…] indexed:0 read:5min 2026-09-01 ·