cd /news/artificial-intelligence/real-estate-scraper-token-efficient-… · home topics artificial-intelligence article
[ARTICLE · art-116498] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Real Estate Scraper– Token-efficient structured property crawler

Kodomoppoi released Real Estate Scraper, an open-source pipeline that uses LLM Structured Outputs to discover real estate websites, crawl listing pages, and extract structured property data without site-specific scrapers. The tool features AI pre-curation, DOM token condensation (~75% noise reduction), and country-aware extraction, with support for Google Gemini (4,000,000 TPM) and Groq Cloud models, and is designed to handle portals in any language.

read4 min views1 publishedAug 31, 2026
Real Estate Scraper– Token-efficient structured property crawler
Image: Michielbdejong (auto-discovered)

An intelligent, token-efficient pipeline designed to automatically discover real estate websites, crawl listing pages, and extract structured property data using LLM Structured Outputs. #

VIDEO SHOWCASE : https://youtu.be/xcjRcRTZt-I?si=thM-5ZTkQCM7sZmI

Interactive Real Estate Explorer with KPI metrics, AI curated portals, and property listing tables.

⚠️ Important Note on Crawl Settings, Model Selection & Rate Limits (TPM):

Exponential Crawl Scaling: The total number of crawled pages, AI token expenditure, and overall execution time increaseexponentiallywith theCrawl Settings(AI Curated Sites

×Pages per Site

).

For Testing: It is strongly recommended to set1 Curated Siteand1 Page per Site (to verify location results quickly and conserve API quota.1 / 1

)For Full Scrapes: Increase to higher limits (e.g., 2–5 sites, 2–3 pages) once you confirm portal accessibility.

AI Provider Capacity & Extraction Yield (TPM Limits):

Google Gemini (: Features a 4,000,000 TPM limit and 1M context window. It effortlessly extractsgemini-3.6-flash

) [Recommended for Maximum Volume]20 to 25 complete listings per pagein seconds.Groq Cloud (Free Tier): Provides ultra-fast LPU inference, but large 70B/120B models (e.g.openai/gpt-oss-120b

) have a tight ~6,000 TPM cap that can rate-limit full-page extraction down to only 1–2 listings per request.For Groq, use high-throughput modelslikeqwen/qwen3.6-27b

,groq/compound-mini

, oropenai/gpt-oss-20b

(20,000+ TPM).

Traditional web scrapers rely on brittle CSS/XPath selectors that constantly break whenever real estate portals change layouts, obfuscate class names, or render dynamic JavaScript feeds.

By leveraging an AI semantic extraction engine with Pydantic structured outputs, this pipeline extracts clean, structured property data across any portal worldwide in any language without writing or maintaining site-specific scrapers.

Location & Filters (e.g. Ipanema, Rio de Janeiro / Brasil)
       │
       ▼
1. Direct Listing Discovery (DuckDuckGo Engine - Scaled Candidates)
       │
       ▼
2. AI Pre-Curation & Index Matching (Preserves exact deep routes & filters noise)
       │
       ▼
3. Dual-Engine Crawler with Early Site Abandonment (Validates Page 1 first)
       │
       ▼
4. DOM Token Condensation (~75% Noise Reduction)
       │
       ▼
5. Country-Aware Structured Extraction (Suites, Amenities, Highlights, Financing)
       │
       ▼
6. Fail-Fast Resiliency, Pandas Deduplication & CSV/JSON Export

Problem Solved: Real estate homepages are landing pages with complex search forms, while hardcoding portal URLs breaks across cities. The discovery engine performs targeted natural-language queries (e.g.,apartamentos a venda em Ipanema Rio de Janeiro

), retrieving live deep listing search URLs dynamically without form automation.

Problem Solved: Asking an LLM to rewrite or generate URLs causes link hallucinations or truncates deep paths to root domains (e.g. returningzapimoveis.com.br/

). By numbering candidate URLs and having the LLM select 1-based integer indexes ([1, 2]

), exact deep paths are preserved with 100% fidelity.

Problem Solved: Crawling multi-page routes on dead or anti-bot blocked portals wastes network time and AI tokens. Page 1 is validated first; if 0 listings are found or access is blocked, all remaining pages for that site are aborted immediately.

Problem Solved: Raw webpage HTML contains 80,000+ characters of SVGs, cookie banners, navigation menus, and ads. The regex cleaner strips noise and filters text to retain only property-relevant signals (R$

,

,quartos

,amenities

), fitting within fast LLM token windows.

Problem Solved: Property listings are unstructured and written in diverse regional formats. The Pydantic schema extracts normalized attributes (price

,area_m2

,bedrooms

,suites

,amenities

,financing_accepted

) localized to the target country's official language.

Problem Solved: Rate limits and quota exhaustion previously led to long wait loops. The engine enforces immediate fail-fast handling on critical errors and compiles extracted records into deduplicated Pandas DataFrames for instant CSV (utf-8-sig

) and JSON export.

git clone https://github.com/Kodomoppoi/Real-estate-Scrapper.git
cd Real-estate-Scrapper

streamlit run app.py

Configure your API key directly in the Web Dashboard sidebar or create a .env

file in the project root:

GEMINI_API_KEY=AIzaSy...
LLM_MODEL=gemini-3.6-flash

OPENAI_API_KEY=sk-...
LLM_MODEL=gpt-4o-mini

Real-Time Terminal Activity: Live streaming terminal box embedded directly inside the browser showing search, crawling, and AI steps. - In-App API Key Manager: Test and save Gemini, OpenAI, Groq, or OpenRouter API keys directly from the sidebar. - KPI Metrics Cards: Total listings, estimated average market price, median area ($m^2$ ), and top neighborhood. - Interactive Listings Table: Client-side keyword search, neighborhood filters, bedroom filters, price range filters, and direct links to original ads. - Extra Details Tab: Amenity frequency rankings, financing status breakdown, and Price-per-$m^2$ calculation rankings. - One-Click Export: Export consolidated datasets to CSV (Excel compatible withutf-8-sig

) and JSON.

Run the automated test suite with pytest:

pytest tests/
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @kodomoppoi 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/real-estate-scraper-…] indexed:0 read:4min 2026-08-31 ·