{"slug": "scrape-any-website-into-json-by-just-listing-the-fields-you-want", "title": "Scrape any website into JSON by just listing the fields you want", "summary": "A developer built an AI Web Data Extractor on Apify that turns website scraping into a schema-first task: users list the fields they want, an LLM reads the page, and the tool returns JSON without CSS selectors or custom code. The extractor supports string, number, integer, boolean and array field types as well as full JSON Schema for nested results, returning null rather than fabricating missing values, and can be invoked from code or used as a tool by AI agents through the Apify MCP server. The developer notes classic selector-based scrapers remain cheaper for millions of pages with a fixed layout, while this tool targets varied or changing pages and small field sets across many sites.", "body_md": "Writing a scraper usually means inspecting the page, finding CSS selectors, handling edge cases, and then fixing it all again when the site changes its layout. For a lot of jobs that's overkill: you just want *these five facts* from *these pages*.\n\nSo I built an extractor that works the other way around. You describe the data, an LLM reads the page, and you get JSON back.\n\n```\n{\n  \"startUrls\": [{ \"url\": \"https://github.com/apify/crawlee\" }],\n  \"fields\": {\n    \"name\": \"string\",\n    \"description\": \"string\",\n    \"license\": \"string\",\n    \"primary_language\": \"string\"\n  }\n}\n{\n  \"url\": \"https://github.com/apify/crawlee\",\n  \"success\": true,\n  \"data\": {\n    \"name\": \"crawlee\",\n    \"description\": \"Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers...\",\n    \"license\": \"Apache License 2.0\",\n    \"primary_language\": \"JavaScript\"\n  }\n}\n```\n\nNo selectors, no code, and the same input works on any other repository page, or on a completely different site if you change the fields.\n\n`null` otherwise. Missing is better than made up.`string`, `number`, `integer`, `boolean` or `array` per field, or a full JSON Schema for nested results like \"all jobs on this page with title, location and remote flag\".\nIf you need millions of pages from one site with a fixed layout, a classic selector-based scraper is cheaper. This tool shines when pages vary, layouts change, or you only need a handful of fields from many different sites.\n\n[AI Web Data Extractor on Apify](https://apify.com/tidytools/ai-web-data-extractor). New Apify accounts get free monthly credit, so you can try it on your own URLs for free. It can also be called from code or used as a tool by AI agents through the Apify MCP server.\n\nFeedback welcome, especially examples where it gets something wrong.", "url": "https://wpnews.pro/news/scrape-any-website-into-json-by-just-listing-the-fields-you-want", "canonical_source": "https://dev.to/tidytools/scrape-any-website-into-json-by-just-listing-the-fields-you-want-5bbi", "published_at": "2026-09-28 20:53:48+00:00", "updated_at": "2026-09-28 21:20:01.116439+00:00", "lang": "en", "topics": ["ai-tools", "ai-agents", "agent-protocols", "structured-data"], "entities": ["Apify", "Crawlee", "AI Web Data Extractor", "Apify MCP server", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/scrape-any-website-into-json-by-just-listing-the-fields-you-want", "markdown": "https://wpnews.pro/news/scrape-any-website-into-json-by-just-listing-the-fields-you-want.md", "text": "https://wpnews.pro/news/scrape-any-website-into-json-by-just-listing-the-fields-you-want.txt", "jsonld": "https://wpnews.pro/news/scrape-any-website-into-json-by-just-listing-the-fields-you-want.jsonld"}}