Shopify competitor analysis: how to validate a new market with Apify and Claude A new guide from Apify demonstrates how to validate a new market for a Shopify store using its web-scraping platform and Anthropic's Claude AI, using a fictional running socks brand expanding into Latin America as a case study. The process involves extracting competitor data with Apify Actors, analyzing it with Claude, and answering five key questions to avoid costly expansion mistakes. Merchants rarely lose money in new markets because their product is bad. Usually, it's because they committed blindly without proper research. Imagine stocking up on inventory, signing partnership deals, and launching ads, only to find out that your pricing is wrong for the market, or your sizing chart is confusing, and a local brand already dominates your exact positioning. What you need is enough precise data, collected and analyzed extensively, before a single dollar is invested. In this guide, you’ll run a Shopify competitor analysis for a fictional store that sells running socks and is expanding into Latin America. The analysis will pull Shopify data scraped with Apify, plus live import duty rates, consumption taxes, low-value parcel thresholds, and exchange rates sourced from the web. Claude runs Apify Actors through Apify's Model Context Protocol MCP server, then analyzes the datasets they produce - pulling smaller results into the chat and larger ones in as file exports - and builds a live report as follows: - Extract competitor catalogs and customer reviews to get a snapshot of the market. - Search the web for current exchange rates and cross-border fees to establish true costs. - Translate and cluster actual buyer complaints for each region. - Filter the combined data through a validation framework. - Configure your Shopify store for the market that proves most viable. Case study: Cadence & Co., a Shopify store expanding into Latin America Cadence & Co. is a fictional US Shopify store selling performance running socks priced at $16 a pair. The socks are made in Vietnam at a production cost of $5.50 per pair plus roughly $1.20 in freight costs. Where the socks are made matters here. Trade agreements attach to the physical goods, not the seller. If these socks were manufactured in the United States, they could qualify for duty-free entry or lower import taxes under regional trade deals. But because they are produced in Vietnam, they receive no such preferential treatment. These numbers sit close to where real US performance sock brands price, and they are the first thing you should swap for your own once you start working through the slides in your finished report. Running socks are a great product to test a framework like this on because they are utilitarian, and runners are usually a vocal audience eager to leave reviews. Cadence & Co. has shortlisted three markets: Mexico, Chile, and Brazil. On paper, Brazil looks like the obvious pick. It’s the largest e-commerce market in the region with an enormous running population. But “obvious” is exactly the kind of reasoning this process puts to the test. Why most merchants fail when expanding into new markets Going on gut feeling is the costliest mistake you can make when expanding. So is picking the wrong metrics, or committing too early, before you know who you're competing against, or whether your supplier price even leaves room for profit at local prices. Every successful expansion runs on three distinct layers: Intelligence. You must know exactly who your competitors are, what they charge, what their customers complain about, and whether there is a product gap you can fill. Storefront. This is what you show the customer, presented in their local currency and language, and fully adapted to their expectations. Fulfillment. You need a local delivery strategy that matches what the market considers normal, because no customer base used to a two-day delivery will wait three weeks for socks. Validating a new market means answering five questions This tutorial is focused on the intelligence layer. It’s the most critical phase of expansion, yet it is almost always sidelined. Apify is purpose-built for this step, offering ready-made Actors that extract the live competitor pricing, stock levels, and customer reviews needed to validate a market before you invest. But data is only valuable if you know what to look for. Five questions decide whether a market is worth entering: - Is there real demand here? A catalog gives you three signals. Products selling out mean buyers are arriving faster than competitors can restock. Full shelves and no new launches mean the opposite. A steady stream of new listings means competitors expect the demand to hold. - Who else is selling this? Bulk catalog scrapes of competitor stores tell you exactly who you are up against before committing to a marketing strategy. - What are they charging? Competitor pricing tells you whether the market has a cheap shelf, a premium shelf, or both. - What do customers complain about? Reviews surface the exact features and expectations you need to match or beat, in the buyer's own words. - Will my margin hold locally? Duties, taxes, and shipping can erase a price that looked profitable on paper. What you're building A Shopify competitor intelligence pipeline that utilizes multiple Apify Actors via its MCP server, like RAG Web Browser https://apify.com/apify/rag-web-browser to discover and verify local Shopify competitor stores in Mexico, Chile, and Brazil, and Shopify Products Scraper https://apify.com/trovevault/shopify-products-scraper to extract live prices, discounts, variants, and stock levels from their catalogs. Then, a platform-specific review scraper collects actual buyer feedback in Spanish and Portuguese. Claude ties this together by translating and clustering complaints, benchmarking prices in local currency, calculating true landed costs using real-time trade data, and building a visual interactive dashboard. Prerequisites To follow along, you need an Apify account https://console.apify.com/sign-up and a Claude account https://claude.ai/ with access to the Claude desktop app so you can add connectors. Step 1: Connect Apify to Claude through MCP Instead of manually running every Actor and downloading datasets for this tutorial, you can connect Claude directly to Apify. This connection uses the Model Context Protocol MCP https://docs.apify.com/integrations/mcp , an open standard that allows AI applications to call external tools. Since Apify runs an MCP server, a one-time setup transforms your chosen Actors into tools that Claude can run on your behalf, and the results flow directly back into your chat. Here’s how to set it up: - Go to mcp.apify.com https://mcp.apify.com/ and sign in with your Apify account. - Click Add Actors under Preloaded Actors and select three: the Shopify Products Scraper https://apify.com/trovevault/shopify-products-scraper/ for catalogs, the Shopify Reviews Scraper https://apify.com/mighty monk/shopify-reviews-scraper/ to detect which review app each store runs, and the RAG Web Browser https://apify.com/apify/rag-web-browser/ for store discovery. - You’ll add a fourth Actor later: the Judge Me Review Scraper https://apify.com/stanvanrooy6/judge-me-scraper/ , which pulls the actual review text. - Scroll up and copy the MCP server URL. - Open Customize in Claude and select Connectors . - Click the + icon, choose Add custom connector , and give it a name. - Paste the URL, click Add , then Connect , and authorize your account. - Keep in mind that free Claude plans allow only one custom connector. - Click the + icon below the message box in your chat and toggle the new connector on. - Turn off any other Apify connectors so Claude is not juggling overlapping tool names. Step 2: Discover your competitor set per market A localized search narrows results down to the specific markets you're building for in Latin America, so the key search words within the prompt are in Spanish and Portuguese. First prompt: product search For best results, it’s important to use terms and phrases a real shopper in each market would actually type into a search engine: Using RAG Web Browser, find Shopify stores selling running socks in three markets. Search these queries: "tienda calcetines running México", "calcetines de compresión correr comprar México", "tienda calcetines running Chile", "calcetines running comprar Chile", "loja meias de corrida Brasil", "meias de compressão corrida comprar Brasil". Record the exact query that surfaced each domain at the moment you collect it, and do not reconstruct that column afterwards. Then verify which candidates actually run Shopify by calling trovevault/shopify-products-scraper on the candidate domains with maxProducts set to 1, and drop any domain that returns nothing. Return only a table of verified stores with three columns: domain, store name, and the query that found it. Favor local specialist brands and regional sportswear stores that carry socks as a category. Do not run any queries beyond the six above, and tell me the verified count per market rather than padding it. No commentary. This first prompt verified six stores: three in Mexico, two in Chile, and one in Brazil. That’s a start for Mexico but nowhere near a competitor set for others, which is what the second prompt is for. Second prompt: search Shopify's URL structure Shopify storefronts share a /collections/ URL path that ordinary retailers don't, and searching for it with the site: operator pins down platform and country in a single search: Using RAG Web Browser, run these three searches and record which one surfaces each result as you collect it: "/collections/" calcetas running compresión site:mx , "/collections/" calcetines running compresión site:cl , and "/collections/" meias de corrida compressão site:com.br . Collect every distinct domain, exclude marketplaces, global brand storefronts, and any store already on my list. Verify each new candidate with trovevault/shopify-products-scraper at maxProducts 1, take the currency from that response rather than inferring it from the domain, and drop anything that returns nothing. Return the new candidates with vendor name, currency, market, and the query that found them. This pass did a better job and increased the store count from six to 24. But the probe proves a store runs on Shopify, not that it’s a competitor; a third prompt ensures that Claude filters for direct competitors in the running socks category: Third prompt: filter verified stores down to real competitors A candidate only counts if it actually sells running or sports socks. Exclude marketplaces, global brand storefronts, dev and unavailable stores, and any retailer whose category is adjacent rather than overlapping, including children's clothing, medical supply, and other sports. Where you can't tell what a store sells, don't sample the front of its catalog. Fetch its sock collection directly, or scrape enough of the catalog to search productType and title against the sock stems, and tell me which method you used per store. Step 3: Scrape the competitor catalogs After verifying the store list, pull every catalog into one run: Split the verified stores into two groups: specialist sock and running brands, and broad sportswear retailers that carry socks among many other categories. Run trovevault/shopify-products-scraper on the specialists with maxProducts 1000, and on the broad retailers with maxProducts 2000. Leave the proxy disabled for both. Report both dataset IDs, how many products each store returned, and flag any store that returned exactly its cap, since that store was truncated and needs a re-run at a higher limit. A store that returns exactly its cap was most likely truncated; you may use the max figure as it is or re-run only that store at a higher limit. Step 4: Export the catalogs and summarize them The tool Claude uses to read Apify datasets can select columns but can’t filter rows, so filtering thousands of products down to socks inside the conversation would mean every row traveling into Claude’s context first. Export them into the chat instead and let Claude analyze them as standalone files: - In Apify Console, open Storage , then Datasets . Your two runs sit at the top. - Click into each and glance at the product titles. The specialists dataset is almost all socks; the broad retailers one is full of shoes, shirts, and shorts. - Open the Actions menu, choose Rename , and use something like cadence-specialists-latam and cadence-broad-latam . - Click Export , choose CSV, and download it. - Upload both CSVs into your Claude conversation. Then run: These CSVs are Shopify catalog scrapes from three markets. Filter them to running sock products only, matching title, productType, and tags against the stems sock, calcetin, calceta, and meia, so that singular and plural forms both match. Assign each store to Mexico, Chile, or Brazil by its domain. Then give me one summary table per market containing: number of stores contributing sock products, number of sock products found, currency, median priceMin in local currency, price range, share on sale, average discount depth where compareAtPrice exists, share fully out of stock, and how many were published in the last 90 days. Break each market's median and price range out by store tier as well, specialist versus broad retailer, alongside the blended figure, and tell me how far apart the two tiers sit. Do all counting and arithmetic in code, not by reading rows. Flag any market with fewer than 25 sock products as too thin to draw a median from. Keep every figure in local currency; do not convert to USD. Step 5: Mine the reviews The catalog told you what competitors sell and for how much. Reviews will reveal product gaps and buyer complaints that might give you an edge over them. First, survey which review app each store runs Find out which review platform each store uses and how many collect reviews at all: Run mighty monk/shopify-reviews-scraper against one product URL per store on my list, with concurrency 2, requestDelayMs 2000, and the residential proxy enabled. Return a table of store, review platform detected, and total review count. Do not collect review text yet. This output tells you which review Actor to rent and how much of the market you'll be able to capture. In this run's survey, 8 of the 24 stores collected no reviews at all, 11 ran Judge.me http://Judge.me , 2 ran Loox, 1 ran Yotpo, and 2 ran Shopify's native reviews app. Rent the Actor for the platform that dominates To extract the actual reviews, you need to query each platform’s widget API with the correct shop identifier and handle pagination. Dedicated Apify Actors manage this complex routing for you on platforms like Judge.me http://judge.me/ , Loox, Yotpo, and Okendo so you can save time and focus entirely on analyzing your data. Here, the survey pointed to Judge.me http://Judge.me , so: - In Apify Console, open the Judge Me Review Scraper for Shopify stores https://apify.com/stanvanrooy6/judge-me-scraper/ stanvanrooy6/judge-me-scraper and click Rent Actor . It's a flat $19-a-month fee rather than per-result billing. - Run it from Console rather than through Claude. Rental Actors don’t appear in the MCP Actor picker until activated, and even then it is simpler to configure this one directly. - Test it on a single store before running the rest. Run it site-wide The Actor takes a store domain, not product URLs, and pulls the store's entire review dataset. - Take the domains your survey flagged as Judge.me http://Judge.me . These are the same verified domains from Step 2, so paste them straight from that list. - Raise the review limit well above what you expect. - Run it, then check each result count against the cap you set. Export each run's dataset as JSON or CSV from Storage , and upload the files to your Claude conversation. Filter the dataset before clustering Before clustering, apply two filters to drop rating-only submissions and isolate the seller's full product catalog. These files are Judge.me http://Judge.me review exports from Shopify stores in three markets. First, drop any review whose body is a fuzzy match for its product name, since those are rating-only submissions where the Actor backfilled the body with the title. Use fuzzy matching rather than exact, since some bodies are the title plus a size or variant. Second, keep only reviews left on running or sports sock products, matching the product name against the stems sock, calcetin, calceta, and meia, and excluding kids' novelty, character, boxer, and underwear lines. Report how many rows each filter removed and the surviving review count per store, so I can see where the sample is thin. Then assign each review to a market and a price tier, translate the Spanish and Portuguese, and cluster the complaints per market and per tier: the top three to five themes each, how often each appears, how many stores it spans, and two verbatim quotes in the original language with an English translation. Weight toward reviews at three stars and below, and separate product complaints from fulfillment and service complaints. Do all counting in code. A product complaint exposes what to improve, while an operational complaint reveals where competitors are letting people down, so you can step in and offer better service. Review text is user-generated content from sites you don't control, and you're feeding it into a model with tools attached. Text from the open web can carry malicious instructions aimed at the model rather than a reader. Keep this conversation's tool access narrow, enable only the Apify connector, and switch off anything reaching your email, files, or databases, so the worst outcome of a poisoned review is a wrong summary rather than an action taken against something that matters. Step 6: Run the analysis and build the dashboard Everything now resides in two compact tables within your conversation, ensuring the final prompt doesn’t need to access external datasets. Using only the market summaries and complaint clusters above, run the four-question expansion framework for Cadence & Co. Production cost is $5.50 per pair, freight is $1.20 per pair, and the socks are made in Vietnam. Convert every figure to USD using these rates, which I pinned today: rate per USD for each currency . State the rates and the date in the output. First, list every store's typical price per pair in USD, sorted low to high, ignoring which market each belongs to. Divide any multipack by its pair count so packs and singles are comparable. Identify the price clusters and any gap between them, and say where our own price falls. Then run the framework at three price points per market, the cheap cluster, our own price, and the premium cluster: compute landed cost including current import duties and taxes, strip the consumption tax from the shelf price, and give contribution per pair. State the duty and tax rates you used, their source, and flag any you are uncertain about. Then judge the positioning gap, demand structure, and local delivery for each market, and produce a verdict table with pass, partial, or fail per question. Where the evidence does not support a judgement, say so rather than inferring one. Render it as a self-contained visual dashboard artifact, and use the artifact's persistent storage so weekly snapshots can be added later and compared. Because the aggregates were computed once in Step 3 from files you scoped, the dashboard is built on numbers with a traceable origin rather than figures reconstructed from memory. An agent asked to remember a figure may hallucinate one, so give it a file to compute from instead. Here is what that prompt produced for Cadence, exchange rates pinned July 24, 2026. What the analysis found 1. Nobody sells a single pair anywhere near $16 When you line up every store with five or more running socks by their per-pair price, a clear market structure emerges. Six stores cluster tightly between $6.64 and $11.49, while seven others cluster between $18.17 and $39.02. No price point sits in the $6.68-wide gap between them, which might be an opportunity for Cadence at its $16 price point. Claude provides three readings on why that gap exists and a recommendation: 2. Evaluating profit margin against regional import duties and taxes Because Cadence manufactures in Vietnam, it receives no trade agreement benefits in these countries. Imports face a 35% duty into Mexico, a 6% duty into Chile, and rely on Brazil’s temporary low-value parcel exemption. To analyze profitability, Claude calculated landed cost production plus freight and duty against net revenue shelf price stripped of local consumption taxes . When tested across three distinct price points, the constraints become immediately clear: Selling at the low-end shelf price of $10.77 returns less per pair than the $9.30 you'd make selling at home in the US; your $16.00 price point is viable but at reduced margins, while selling above $25.49, on the premium tier, returns a healthy margin. The report is built to be dynamic, so you can slot your own operational costs and final sale price in the slide below, and it will automatically recalculate feasibility based on those new constraints: 3. What the reviews could and could not tell you Across 11,643 reviews scraped from five stores, complaint rates ranged from 0.11% to 4.86%. This reveals a data-integrity issue rather than a performance gap. The Judge.me http://judge.me/ review platform allows merchants to curate which reviews are published, so a low complaint rate most likely reflects an aggressive publication policy, not genuine customer satisfaction. For instance, the Mexican data pool published zero critical reviews, while Brazil yielded only a single sock review across the entire country mining Loox or Yotpo data may yield different results for Brazil . When filtering the entire dataset down to explicit complaints about running socks, only six reviews remain, all originating from a single Chilean retailer highlighting issues separated into two main categories: Operational issues 5 reviews : Addressed order fulfillment issues, including incorrect items shipped, unhelpful customer service, and missing packages. Product quality 1 review : Noted an issue with sizing, stating the socks were too large. Although the sample size is small, one critical signal stands out. Within the Chilean store, running socks drew a 17.6% complaint rate compared to just 3.5% for novelty socks. Because both product lines share the same warehouse, delivery team, and customer support line, this gap points to an inherent product-line flaw rather than an operational one. 4. Is anyone actually buying? Analyzing running socks data alone shows Mexico leads consumer demand with 26.8% of its 340-product running catalog fully sold out, whereas Chile represents the weakest demand at a 13.3% sell-out rate, albeit balanced by a high 15.1% new-product launch rate. Brazil sits in the middle with a 14.2% sell-out and a 9.5% launch rate across 169 products, though market depth is limited to just three core stores. 5. The scorecard The market analysis recommends prioritizing Chile for expansion, followed by Mexico, based on a four-pillar scorecard analyzing pricing, positioning, demand, and delivery. While Chile offers favorable delivery and margins, Mexico provides superior demand, with Brazil currently ranked lowest due to weak data, high risk, and measurement limitations. The report recommends a low-cost ad test in Santiago to validate the $16 price point before committing fully. For a detailed analysis of the market evaluation, refer to the full report https://cadence-latam-socks-report.netlify.app/ 15 . Conclusion For a few dollars, Apify found and scraped the product catalogs of 24 Shopify competitors across three Latin American countries. Claude then uncovered hidden pricing gaps, modeled landed costs across three scenarios based on actual cross-border fees, and separated high-demand products from high-margin opportunities. Cadence & Co. entered Chile already knowing what the retail shelf looked like, where the market gaps were, and exactly what local runners complain about. None of it required custom code, just four ready-built Actors for the tasks: RAG Web Browser for target stores, Shopify Products Scraper for catalogs, Shopify Reviews Scraper to detect each store's review platform, and a Judge.me http://judge.me/ for scraping the reviews. This pipeline scales instantly. You can swap domains to audit markets in Europe or Asia. Or run a separate analysis for other products. With 55,000+ Actors spanning marketplaces, ad libraries, and review platforms https://apify.com/store/categories , your only constraint is your curiosity, not your data. FAQs Does this only work for Shopify competitors? The catalog technique relies on Shopify’s public products.json endpoint, so it only works on Shopify stores. To capture the entire market across all platforms and marketplaces, use Apify’s E-commerce Scraping Tool https://apify.com/apify/e-commerce-scraping-tool . The underlying analysis framework remains exactly the same. Is scraping public Shopify data allowed? The Actors in this build access publicly available data through the endpoints Shopify deliberately exposes. You remain responsible for using the data in line with applicable laws and the target stores' terms. Can I point the scraper at one specific competitor instead of a whole market? Yes. The store list is the input, so a list of one works, and scraping a single known competitor on a schedule is a very effective way to use this pipeline. How do I know the competitor list my agent returned is real? Check the list twice before you build on it. Every store should have returned at least one product when you probed it, which is proof that it runs Shopify. Every store should also carry the search query that found it, recorded the moment it was collected.