Writing a scraper usually means inspecting the page, finding CSS selectors, handling edge cases, and then fixing it all again when the site changes its layout. For a lot of jobs that's overkill: you just want these five facts from these pages.
So I built an extractor that works the other way around. You describe the data, an LLM reads the page, and you get JSON back.
{
"startUrls": [{ "url": "https://github.com/apify/crawlee" }],
"fields": {
"name": "string",
"description": "string",
"license": "string",
"primary_language": "string"
}
}
{
"url": "https://github.com/apify/crawlee",
"success": true,
"data": {
"name": "crawlee",
"description": "Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers...",
"license": "Apache License 2.0",
"primary_language": "JavaScript"
}
}
No selectors, no code, and the same input works on any other repository page, or on a completely different site if you change the fields.
null otherwise. Missing is better than made up.string, number, integer, boolean or array per field, or a full JSON Schema for nested results like "all jobs on this page with title, location and remote flag".
If you need millions of pages from one site with a fixed layout, a classic selector-based scraper is cheaper. This tool shines when pages vary, layouts change, or you only need a handful of fields from many different sites.
AI Web Data Extractor on Apify. New Apify accounts get free monthly credit, so you can try it on your own URLs for free. It can also be called from code or used as a tool by AI agents through the Apify MCP server.
Feedback welcome, especially examples where it gets something wrong.