Reading any marketplace product page from a Chrome extension: JSON-LD first, selectors second A developer building AI Product Page Optimization, a Chrome extension that reads an open product page and returns marketplace-ready title, description, and image suggestions in about 30 seconds, described the extraction pipeline behind it: parse application/ld+json Product blocks first, then fall back to URL-scoped, platform-specific DOM selectors, then apply each marketplace's field constraints. The implementation detects page language so output matches the listing locale and never auto-pastes, routing extracted data and suggestions to a dashboard for human review before publishing. If you have ever tried to build a tool that "reads" a product page — an Amazon listing, a Shopify storefront, an eBay item, an Etsy shop page — you already know the problem: every marketplace renders the same conceptual data title, bullets, description, images, price in completely different DOM shapes. We ran into this while building AI Product Page Optimization https://aiproductpageoptimization.com , a Chrome extension that reads the product page you already have open and returns marketplace-ready title, description, and image suggestions in about 30 seconds. This post is the technical breakdown of how the reading step works. This article is disclosed as AI-assisted; the extraction design and constraints below are from our actual implementation. Most marketplaces embed structured data in the page. Before touching a single DOM selector, we look for application/ld+json blocks and parse them for Product -shaped objects: js const ldjson = Array.from document.querySelectorAll 'script type="application/ld+json" ' .map s = { try { return JSON.parse s.textContent ; } catch { return null; } } .filter Boolean ; Why JSON-LD first: Product schema block rarely moves. name , images are an image array, and offers carry price/currency regardless of which marketplace you are on. But JSON-LD is never the whole story. Marketplaces routinely put richer or more current data in the DOM than in their structured data: bullet points that are not in any schema field, variant-specific copy, A/B-tested titles. After JSON-LD, we apply platform-specific DOM selectors to fill the gaps — and the key word is scoped : we detect the marketplace from the URL first, then load only that platform's selector set. The four surfaces we read today: Reading the page is half the job. The other half is knowing what each marketplace will accept when you write the optimized copy back: One more detail that surprised us: output language follows the detected page language. Cross-border catalogs mix English with other locales, and returning English suggestions for a Japanese listing page is worse than useless — it breaks the seller's workflow. Two things worth knowing if you build this on Manifest V3: We deliberately never auto-paste anything. Extracted data and generated suggestions land on a web dashboard with a before/after comparison, history, and optional share links — the seller copies and publishes only what they approve. Optimization is credit-based text 1 credit, AI-enhanced image 2, keeping an original image is free; a typical product page uses about 5 credits, and new accounts start with 5 free credits . If you are building anything that reads product pages — price trackers, feed tools, listing optimizers — the pattern is the same: structured data first, scoped selectors second, platform constraints third, and a human review step before anything goes back to the marketplace. Try the extension: aiproductpageoptimization.com https://aiproductpageoptimization.com