Real Estate Scraper– Token-efficient structured property crawler Kodomoppoi released Real Estate Scraper, an open-source pipeline that uses LLM Structured Outputs to discover real estate websites, crawl listing pages, and extract structured property data without site-specific scrapers. The tool features AI pre-curation, DOM token condensation (~75% noise reduction), and country-aware extraction, with support for Google Gemini (4,000,000 TPM) and Groq Cloud models, and is designed to handle portals in any language. An intelligent, token-efficient pipeline designed to automatically discover real estate websites, crawl listing pages, and extract structured property data using LLM Structured Outputs. VIDEO SHOWCASE : https://youtu.be/xcjRcRTZt-I?si=thM-5ZTkQCM7sZmI https://youtu.be/xcjRcRTZt-I?si=thM-5ZTkQCM7sZmI /Kodomoppoi/Real-estate-Scrapper/blob/main/docs/screenshots/dashboard overview.png Interactive Real Estate Explorer with KPI metrics, AI curated portals, and property listing tables. ⚠️ Important Note on Crawl Settings, Model Selection & Rate Limits TPM : Exponential Crawl Scaling: The total number of crawled pages, AI token expenditure, and overall execution time increaseexponentiallywith theCrawl Settings AI Curated Sites × Pages per Site . For Testing: It is strongly recommended to set1 Curated Siteand1 Page per Site to verify location results quickly and conserve API quota. 1 / 1 For Full Scrapes: Increase to higher limits e.g., 2–5 sites, 2–3 pages once you confirm portal accessibility. AI Provider Capacity & Extraction Yield TPM Limits : Google Gemini : Features a 4,000,000 TPM limit and 1M context window. It effortlessly extracts gemini-3.6-flash Recommended for Maximum Volume 20 to 25 complete listings per pagein seconds.Groq Cloud Free Tier : Provides ultra-fast LPU inference, but large 70B/120B models e.g. openai/gpt-oss-120b have a tight ~6,000 TPM cap that can rate-limit full-page extraction down to only 1–2 listings per request.For Groq, use high-throughput modelslike qwen/qwen3.6-27b , groq/compound-mini , or openai/gpt-oss-20b 20,000+ TPM . Traditional web scrapers rely on brittle CSS/XPath selectors that constantly break whenever real estate portals change layouts, obfuscate class names, or render dynamic JavaScript feeds. By leveraging an AI semantic extraction engine with Pydantic structured outputs, this pipeline extracts clean, structured property data across any portal worldwide in any language without writing or maintaining site-specific scrapers. Location & Filters e.g. Ipanema, Rio de Janeiro / Brasil │ ▼ 1. Direct Listing Discovery DuckDuckGo Engine - Scaled Candidates │ ▼ 2. AI Pre-Curation & Index Matching Preserves exact deep routes & filters noise │ ▼ 3. Dual-Engine Crawler with Early Site Abandonment Validates Page 1 first │ ▼ 4. DOM Token Condensation ~75% Noise Reduction │ ▼ 5. Country-Aware Structured Extraction Suites, Amenities, Highlights, Financing │ ▼ 6. Fail-Fast Resiliency, Pandas Deduplication & CSV/JSON Export Problem Solved : Real estate homepages are landing pages with complex search forms, while hardcoding portal URLs breaks across cities. The discovery engine performs targeted natural-language queries e.g., apartamentos a venda em Ipanema Rio de Janeiro , retrieving live deep listing search URLs dynamically without form automation. Problem Solved : Asking an LLM to rewrite or generate URLs causes link hallucinations or truncates deep paths to root domains e.g. returning zapimoveis.com.br/ . By numbering candidate URLs and having the LLM select 1-based integer indexes 1, 2 , exact deep paths are preserved with 100% fidelity. Problem Solved : Crawling multi-page routes on dead or anti-bot blocked portals wastes network time and AI tokens. Page 1 is validated first; if 0 listings are found or access is blocked, all remaining pages for that site are aborted immediately. Problem Solved : Raw webpage HTML contains 80,000+ characters of SVGs, cookie banners, navigation menus, and ads. The regex cleaner strips noise and filters text to retain only property-relevant signals R$ , m² , quartos , amenities , fitting within fast LLM token windows. Problem Solved : Property listings are unstructured and written in diverse regional formats. The Pydantic schema extracts normalized attributes price , area m2 , bedrooms , suites , amenities , financing accepted localized to the target country's official language. Problem Solved : Rate limits and quota exhaustion previously led to long wait loops. The engine enforces immediate fail-fast handling on critical errors and compiles extracted records into deduplicated Pandas DataFrames for instant CSV utf-8-sig and JSON export. 1. Clone the repository git clone https://github.com/Kodomoppoi/Real-estate-Scrapper.git cd Real-estate-Scrapper 2. Run the application streamlit run app.py Configure your API key directly in the Web Dashboard sidebar or create a .env file in the project root: Google Gemini Recommended - Free Tier available GEMINI API KEY=AIzaSy... LLM MODEL=gemini-3.6-flash Or OpenAI OPENAI API KEY=sk-... LLM MODEL=gpt-4o-mini - Real-Time Terminal Activity : Live streaming terminal box embedded directly inside the browser showing search, crawling, and AI steps. - In-App API Key Manager : Test and save Gemini, OpenAI, Groq, or OpenRouter API keys directly from the sidebar. - KPI Metrics Cards : Total listings, estimated average market price, median area $m^2$ , and top neighborhood. - Interactive Listings Table : Client-side keyword search, neighborhood filters, bedroom filters, price range filters, and direct links to original ads. - Extra Details Tab : Amenity frequency rankings, financing status breakdown, and Price-per-$m^2$ calculation rankings. - One-Click Export : Export consolidated datasets to CSV Excel compatible with utf-8-sig and JSON. Run the automated test suite with pytest: pytest tests/