If you have ever tried to build a tool that "reads" a product page — an Amazon listing, a Shopify storefront, an eBay item, an Etsy shop page — you already know the problem: every marketplace renders the same conceptual data (title, bullets, description, images, price) in completely different DOM shapes.
We ran into this while building AI Product Page Optimization, a Chrome extension that reads the product page you already have open and returns marketplace-ready title, description, and image suggestions in about 30 seconds. This post is the technical breakdown of how the reading step works. (This article is disclosed as AI-assisted; the extraction design and constraints below are from our actual implementation.)
Most marketplaces embed structured data in the page. Before touching a single DOM selector, we look for application/ld+json blocks and parse them for Product-shaped objects:
const ldjson = Array.from(
document.querySelectorAll('script[type="application/ld+json"]')
).map((s) => {
try { return JSON.parse(s.textContent); } catch { return null; }
}).filter(Boolean);
Why JSON-LD first:
Product schema block rarely moves.name, images are an image array, and offers carry price/currency regardless of which marketplace you are on.
But JSON-LD is never the whole story. Marketplaces routinely put richer or more current data in the DOM than in their structured data: bullet points that are not in any schema field, variant-specific copy, A/B-tested titles.
After JSON-LD, we apply platform-specific DOM selectors to fill the gaps — and the key word is scoped: we detect the marketplace from the URL first, then load only that platform's selector set.
The four surfaces we read today:
Reading the page is half the job. The other half is knowing what each marketplace will accept when you write the optimized copy back:
One more detail that surprised us: output language follows the detected page language. Cross-border catalogs mix English with other locales, and returning English suggestions for a Japanese listing page is worse than useless — it breaks the seller's workflow.
Two things worth knowing if you build this on Manifest V3:
We deliberately never auto-paste anything. Extracted data and generated suggestions land on a web dashboard with a before/after comparison, history, and optional share links — the seller copies and publishes only what they approve. Optimization is credit-based (text 1 credit, AI-enhanced image 2, keeping an original image is free; a typical product page uses about 5 credits, and new accounts start with 5 free credits).
If you are building anything that reads product pages — price trackers, feed tools, listing optimizers — the pattern is the same: structured data first, scoped selectors second, platform constraints third, and a human review step before anything goes back to the marketplace.
Try the extension: aiproductpageoptimization.com