cd /news/ai-tools/show-hn-pyscrappy-self-healing-web-s… Β· home β€Ί topics β€Ί ai-tools β€Ί article
[ARTICLE Β· art-98539] src=github.com β†— pub= topic=ai-tools verified=true sentiment=↑ positive

Show HN: PyScrappy, self-healing web scraping selectors plus an MCP server

PyScrappy, an AI-native web scraping toolkit that converts websites into structured, LLM-ready data, has been released as a Python library and MCP server, featuring self-healing selectors, JS rendering, and 20+ built-in scrapers. The toolkit, installable via pip, includes an MCP server for AI agents and a built-in agent for local models like Ollama, supporting tool calling.

read11 min views1 publishedAug 16, 2026
Show HN: PyScrappy, self-healing web scraping selectors plus an MCP server
Image: Michielbdejong (auto-discovered)

PyScrappy is an AI-native web scraping toolkit that turns websites into structured, LLM-ready data. Use it as a Python library or expose it as an MCP server for AI agents.

πŸ“– Documentation: pyscrappy.vercel.app

Generic scraperβ€” give it any URL, get back structured text, links, images, tables, and metadata** LLM-ready output**β€”.to_markdown()

turns any result into clean Markdown; also.to_json()

and.to_dataframe()

MCP serverβ€” expose the scrapers as tools for AI agents (Claude, Cursor, local LLMs, …)** JS rendering**β€” optional Playwright backend for JavaScript-heavy sites** Custom selectors**β€” pass CSS selectors to extract exactly what you need** Chainable**β€” navigate HTML directly with CSS/XPath,Selector

find_all

,find_by_text

, andfind_similar

(Scrapy/BeautifulSoup-style)Adaptive (self-healing) selectorsβ€” remember an element and relocate it by similarity when a site changes its markup, so scrapers don't silently break** Concurrent scraping**β€”scrape_many

/scrape_all

run scrapes in parallelProxy & scraping-API supportβ€” route through a proxy or ScraperAPI/ScrapeOps for blocked sites** TLS-fingerprint impersonation**β€”impersonate="chrome"

gets past anti-bot filters that block plain clients (optionalcurl_cffi

backend)Command-line extractβ€”pyscrappy extract <url> out.md

scrapes a URL straight to a file, no codeRetry & rate-limitingβ€” built-in exponential backoff and per-domain rate limiting** Type-safe**β€” full type hints,py.typed

marker20+ built-in scrapersβ€” Wikipedia, IMDB, stocks, news, GitHub, Amazon/IKEA, YouTube, andmore

pip install pyscrappy

Optional extras:

pip install 'pyscrappy[browser]'
playwright install chromium

pip install 'pyscrappy[dataframe]'

pip install 'pyscrappy[mcp]'

pip install 'pyscrappy[stealth]'

pip install 'pyscrappy[all]'

PyScrappy ships an MCP server that exposes its scrapers as tools, so an agent (Claude, Cursor, an OpenAI agent, a local LLM) can pull structured web data from any URL and hand it straight to the model:

AI agent  ──MCP tool call──▢  PyScrappy  ──fetch + extract──▢  Any website
   β–²                                                                β”‚
   └──────────────  clean Markdown / JSON  β—€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
pip install 'pyscrappy[mcp]'
claude mcp add pyscrappy pyscrappy-mcp

Then just ask: "use pyscrappy to summarize the latest headlines from bbc.com." See MCP server for the full setup and tool list.

Ollama can't talk MCP on its own, so normally you'd run a host (Goose, Cline, …) in between. PyScrappy skips that with a built-in agent that talks to Ollama directly and lets a local model call the scrapers as tools:

pip install 'pyscrappy[mcp]'                 # needs Python 3.10+
pyscrappy chat --model qwen2.5 "what's the current AAPL quote?"

It exposes the same 22 tools as the MCP server. The only requirement is a model that supports tool calling (Llama 3.1, Qwen 2.5, Mistral, …); how well it picks the right tool is up to the model. Point it at a remote Ollama with --host

, and pass -v

to see each tool call.

PyScrappy ships an optional Model Context Protocol server, so an AI agent (e.g. Claude) can call PyScrappy's scrapers as tools and get structured web data back.

pip install 'pyscrappy[mcp]'

The MCP extra installs the standalone fastmcp

package and requires Python 3.10 or newer. On Python 3.9 the core scraping library still works, but the MCP server is unavailable.

This installs the pyscrappy-mcp

command. It uses stdio by default for local MCP clients; Streamable HTTP and legacy SSE are available for remote deployments:

pyscrappy-mcp          # stdio (default)
pyscrappy-mcp --http   # Streamable HTTP
pyscrappy-mcp --sse    # legacy SSE

You can also run the stdio server with python -m pyscrappy.mcp

.

claude mcp add pyscrappy pyscrappy-mcp

Add to your claude_desktop_config.json

and restart the app:

{
  "mcpServers": {
    "pyscrappy": {
      "command": "pyscrappy-mcp"
    }
  }
}

Tip:Claude Desktop does not inherit your shellPATH

. Ifpyscrappy-mcp

is not found, use the absolute path to the command (e.g. the one printed bywhich pyscrappy-mcp

).

The server exposes 20+ tools. The most common ones are ** scrape_url** (any URL β†’ text, links, images, tables, metadata),

,

scrape_wikipedia

,

scrape_stock

, and

scrape_news

β€” plus many more covering image/YouTube/LinkedIn/Hacker News/book search, weather, crypto, currency, dictionary, Amazon/Newegg/IKEA/SoundCloud, IMDB, and Zomato/Uber Eats.

search_github

To see the full, live list, ask the agent to call the ** list_available_scrapers** tool, or from a shell:

python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"

The lookup_movie

tool needs a free OMDb API key. Pass it to the server through your MCP client config, e.g. for Claude Desktop:

{
  "mcpServers": {
    "pyscrappy": {
      "command": "pyscrappy-mcp",
      "env": { "OMDB_API_KEY": "your-key" }
    }
  }
}

Once registered, just ask the agent naturally, e.g. "use pyscrappy to get the latest headlines from bbc.co.uk and the AAPL stock quote."

PyScrappy ships 24 built-in scrapers, and every one that works without a proxy is also exposed as an MCP tool.

A few of them:

β€” scrape any URL with auto-extraction (text, links, images, tables, metadata)GenericScraper

Data / researchβ€”,WikipediaScraper

(Yahoo Finance),StockScraper

(RSS/Atom),NewsScraper

,GitHubScraper

, plus weather, crypto, currency, dictionary, image, LinkedIn-jobs, and book searchHackerNewsScraper

E-commerceβ€”,AmazonScraper

NeweggScraper

,IKEAScraper

Social / media / foodβ€”, SoundCloud, Zomato, Uber Eats (Instagram / Twitter / Spotify also ship, but are blocked and need a proxy)YouTubeScraper

…and many more. To see the full, live list:

python -c "from pyscrappy import list_scrapers; print(', '.join(sorted(list_scrapers())))"

** IMDBScraper** (

lookup_movie

) is the one exception that needs a key β€” a free OMDb

OMDB_API_KEY

(see the MCP configabove for how to pass it).

PyScrappy is extensible: you can add your own scrapers, and third parties can ship them as standalone pyscrappy-<name>

packages. A registered scraper works everywhere a built-in does, including the MCP server and the pyscrappy chat

agent, with no change to PyScrappy core.

In your own code β€” register with the decorator:

from pyscrappy import BaseScraper, register_scraper, get_scraper
from pyscrappy.core.models import ScrapeResult, ScrapeMetadata

@register_scraper("reddit")
class RedditScraper(BaseScraper):
    def scrape(self, subreddit: str, **kwargs) -> ScrapeResult:
        data = self.fetch_and_parse(f"https://old.reddit.com/r/{subreddit}/.json")
        return ScrapeResult(data=[...], metadata=ScrapeMetadata(scraper="reddit"))

get_scraper("reddit")().scrape(subreddit="python")

As a distributable package β€” advertise an entry point in your pyproject.toml

, and PyScrappy discovers it once your package is installed:

[project.entry-points."pyscrappy.scrapers"]
reddit = "pyscrappy_reddit:RedditScraper"

After pip install pyscrappy-reddit

, the scraper shows up in list_scrapers()

, and an AI agent can call it via the scrape_with

MCP tool β€” no core change required.

First-class MCP tools (optional). Add an mcp_tools

mapping and your scraper becomes a dedicated, typed MCP tool instead of only being reachable through the generic scrape_with

β€” its schema is derived from the method signature, so agents get proper named arguments:

@register_scraper("reddit")
class RedditScraper(BaseScraper):
    mcp_tools = {"search_reddit": "scrape"}   # tool name -> method

    def scrape(self, subreddit: str, sort: str = "hot") -> ScrapeResult:
        ...

See the plugin template for a complete, copyable starting point, and the plugin guide for the full walkthrough.

from pyscrappy import scrape

result = scrape("https://en.wikipedia.org/wiki/Web_scraping")

print(result.to_markdown())   # feed straight to an LLM

Prefer raw fields? Every result is a ScrapeResult

with .data

(a list of dicts):

print(result.data[0]["metadata"]["title"])
print(result.data[0]["text"]["word_count"])
python
from pyscrappy import GenericScraper

with GenericScraper() as gs:
    result = gs.scrape(
        url="https://news.ycombinator.com",
        selectors={"title": ".titleline a", "score": ".score"},
    )
    for item in result.data:
        print(item["title"], item.get("score", ""))

When you want to traverse markup directly (Scrapy/BeautifulSoup-style) rather than get back structured dicts, use Selector

:

from pyscrappy import Selector

page = Selector(html)                             # or navigate any HTML string
page.css(".title::text").getall()                 # CSS with ::text / ::attr(name)
page.xpath("//a/@href").getall()                   # XPath (elements, text(), @attr)
page.find_all("h2", class_="title")                # BeautifulSoup-style search
page.find_by_text("Add to cart", tag="button")     # search by text content

first = page.css(".product")[0]
first.css(".price::text").get()                    # chainable
first.find_similar()                               # sibling elements shaped like this one

css()

/ xpath()

return a SelectorList

with .get()

/ .getall()

/ .text()

. find_similar()

locates elements with the same tag and overlapping classes, handy for pulling every card/row once you've found one.

A hard-coded CSS selector silently breaks the day a site changes its markup. Adaptive selectors survive that: save a fingerprint of the element the first time, and if the selector later matches nothing, relocate it by structural and textual similarity instead of returning empty.

from pyscrappy import Selector

page = Selector(html_v1, url="https://shop.example.com")
price = page.css(".price", auto_save=True, adaptive_id="price").get()

page = Selector(html_v2, url="https://shop.example.com")
result = page.css(".price", adaptive=True, adaptive_id="price")
print(result.get(), "β†’ confidence:", result.adaptive_confidence)

How the relocation decides β€” and where it's stronger than a naive similarity match:

Weighted signals, not a flat average. A stableid

/data-*

hook counts far more than a sibling-tag list, so weak signals can't outvote strong ones.Anchor-relative. It remembers the nearest stable ancestor (an id'd /data-*

container) and depth, so it survives layout reshuffles that move absolute positions.Volatility-aware text. Prices, dates, and counts are down-weighted, so healing stays reliable on exactly the fields that change most between scrapes.Confidence-scored.SelectorList.adaptive_confidence

(0-100) tells you how sure the relocation was;threshold=

sets the minimum to accept.

Fingerprints persist in a small JSON store (~/.pyscrappy/adaptive.json

by default, or $PYSCRAPPY_HOME

), namespaced by site so the same adaptive_id

on two sites never collides. Adaptive is entirely opt-in: without adaptive=True

, a broken selector still just returns empty, exactly as before.

Every built-in scraper follows the same pattern β€” instantiate, scrape(...)

, read result.data

(or .to_dataframe()

/ .to_markdown()

):

from pyscrappy import WikipediaScraper

with WikipediaScraper() as ws:
    result = ws.scrape(query="Python (programming language)", mode="summary")
    print(result.data[0]["text"])

Each scraper has its own arguments (Wikipedia, stocks, IMDB, news, YouTube, Amazon/Newegg/IKEA, Uber Eats, and more β€” see the full list). For per-scraper arguments and examples, see the documentation.

Scrape a URL straight to a file without writing any code β€” the output format is inferred from the file extension:

pyscrappy extract https://example.com out.md      # clean Markdown
pyscrappy extract https://example.com out.json    # structured JSON
pyscrappy extract https://example.com out.txt     # extracted page text
pyscrappy extract https://example.com out.html    # raw fetched HTML

pyscrappy extract https://example.com items.txt --css-selector ".product"
pyscrappy extract https://example.com page.md --render-js
python
from pyscrappy import ScraperConfig, GenericScraper

config = ScraperConfig(
    timeout=20.0,            # request timeout in seconds
    max_retries=3,           # retry failed requests
    rate_limit=2.0,          # seconds between requests per domain
    proxy="http://...",      # proxy URL, or a list to rotate through
    scraper_api=None,        # route via a scraping-API service (see below)
    headless=True,           # browser runs headless
    render_js="auto",        # auto-detect if JS rendering is needed
    cache_ttl=0,             # response cache TTL in seconds (0 = disabled)
    impersonate=None,        # e.g. "chrome" to spoof a browser's TLS fingerprint (see below)
)

with GenericScraper(config) as gs:
    result = gs.scrape(url="https://example.com")

Some sites (e.g. eBay, Instagram, Twitter/X, Spotify) block direct automated requests. PyScrappy supports two ways to get through them.

A proxy (or a rotating list) β€” applies to both the HTTP and browser backends:

from pyscrappy import ScraperConfig, AmazonScraper

config = ScraperConfig(proxy="http://user:pass@host:port")

config = ScraperConfig(proxy=["http://p1:8080", "http://p2:8080"])

A scraping-API service (ScraperAPI, ScrapeOps, ScrapingBee) β€” routes requests through the service, which handles proxies and anti-bot challenges for you:

config = ScraperConfig(scraper_api={
    "provider": "scraperapi",   # or "scrapeops", "scrapingbee"
    "api_key": "YOUR_KEY",
    "render_js": True,           # optional
})

with AmazonScraper(config) as scraper:
    result = scraper.scrape(query="laptop")

This is the reliable way to use the scrapers marked "needs proxy" above.

TLS-fingerprint impersonation β€” many anti-bot systems block a plain HTTP client by its TLS/JA3 fingerprint before serving any content. Set impersonate

to mimic a real browser's fingerprint and get past that class of block without a headless browser:

from pyscrappy import ScraperConfig, GenericScraper

config = ScraperConfig(impersonate="chrome")   # or "chrome124", "safari", "firefox"

with GenericScraper(config) as gs:
    result = gs.scrape("https://example.com")

Impersonation currently applies to the synchronous path only; setting it on an async client raises a clear error. All the usual retry, rate-limiting, caching, and robots handling still apply.

Scraping is I/O-bound, so running several scrapes at once parallelizes the network waits. scrape_many

runs one scraper over many inputs; scrape_all

runs a mix of scrapers together. Both preserve input order.

from pyscrappy import scrape_many, scrape_all, AmazonScraper, WikipediaScraper, NewsScraper

results = scrape_many(AmazonScraper, [{"query": "laptop"}, {"query": "phone"}])

results = scrape_all([
    lambda: WikipediaScraper().scrape(query="Python"),
    lambda: NewsScraper().scrape(feed_url="https://rss.nytimes.com/services/xml/rss/nyt/World.xml"),
])

Set cache_ttl

to a positive number of seconds to cache successful GET responses. Repeated requests for the same URL (and query params) within the TTL are served from cache, skipping both the network and the rate limiter. Caching is disabled by default (cache_ttl=0

).

from pyscrappy import WikipediaScraper
from pyscrappy import ScraperConfig

config = ScraperConfig(cache_ttl=300)   # cache for 5 minutes

with WikipediaScraper(config) as ws:
    ws.scrape(query="Python")   # fetched over the network
    ws.scrape(query="Python")   # served from cache

The cache is in memory and shared across scraper instances in the same process (so it also speeds up repeated calls through the MCP server), and is cleared when the process exits. Call HttpClient.clear_cache()

to empty it manually.

Required: httpx

, beautifulsoup4

, lxml

Optional: playwright

(JS rendering), pandas

(DataFrames), fastmcp

(MCP server, Python 3.10+)

All contributions welcome. See Issues.

This package is for educational and research purposes.

── more in #ai-tools 4 stories Β· sorted by recency
── more on @pyscrappy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/show-hn-pyscrappy-se…] indexed:0 read:11min 2026-08-16 Β· β€”