We ran the same ChatGPT prompt three times in a row. Here is what changed. A developer ran the same ChatGPT prompt three times and found that only 4 of 16 domains appeared consistently across all runs, with Mailchimp missing entirely from one run. The top-ranked brands like Klaviyo and Omnisend held their positions, but lower-ranked brands were unstable. The developer warns that single-run measurements can be misleading and recommends multiple runs for reliable brand visibility tracking. We wanted to know how much a single ChatGPT answer can be trusted as a measurement. So we ran the same prompt through ChatGPT more than once, with nothing changed between runs, and recorded what came back each time. We used DataForSEO's ChatGPT LLM Scraper on 2026-08-16. For each prompt we ran the exact same query three times, minutes apart, same settings every time. Nothing about the prompt or the account changed between runs. Any difference in the answers is the model, not us. The prompt was "best email marketing app for Shopify," run three times. Across the three runs, 16 distinct domains showed up somewhere in the answers. Only 4 of those 16 appeared in all three runs. | Brand | Runs appeared in | Position | |---|---|---| | Klaviyo | 3 of 3 | 1st, every time | | Omnisend | 3 of 3 | 2nd, every time | | Mailchimp | 2 of 3 | 4th twice, absent once | Klaviyo and Omnisend held their spots exactly. Mailchimp did not. It sat at position 4 in two runs and was missing from the answer entirely in the third. Nothing about the prompt changed between that run and the other two. If you had only run this prompt once, and it happened to be the run without Mailchimp, you would have concluded Mailchimp does not appear for this query. Run it again and that conclusion is wrong. To see if this held up outside one prompt, we ran 4 prompts about the HTML-to-PDF API category, 3 runs each, for 12 observations total. 66 distinct domains showed up across those 12 observations. Only 5 domains appeared in every run of the specific prompt they showed up on. 4 of 8 brands we were tracking appeared in some runs of a prompt and were missing from others. The pattern from the email marketing prompt was not a one-off. The names at the top of an answer tend to hold their position. Everything below that is less settled than a single run makes it look. Running all 12 observations for the HTML-to-PDF category cost $0.004 per call, $0.048 in total. Checking whether an answer is stable is not expensive. Not checking is the part that costs something, because it produces a wrong number with the same confidence as a right one. A tool that runs a prompt once and reports "brand visibility: 67%" is reporting a single sample as if it were a measurement. Based on what we saw here, that number would land differently depending on which of the three runs happened to be sampled. The top of a ranking is stable. Klaviyo held position 1 in every run. Omnisend held position 2 in every run. That part, one run would have told you correctly. The marginal brand is not stable. Mailchimp came and went. In the wider check, more than half the brands we tracked were inconsistent from run to run. That instability sits exactly where most brands actually are, not at the top, and it is exactly what a business paying for this kind of tracking wants to know: are we in or out, and how often. A single run cannot answer that question. It can only tell you what happened once. Measured on 2026-08-16 using DataForSEO's ChatGPT LLM Scraper. Same prompts, same settings, three runs each, minutes apart.