My AI visibility score was 78%. Then I fixed how I measured it and it was 25% A developer at Reidify Design found that their studio's AI visibility score of 78% was inflated due to a flawed testing methodology. By separating blind discovery questions from named fact-checking questions into distinct sessions, the corrected score dropped to 25%, revealing inconsistent visibility across ChatGPT, Perplexity, and Gemini. The developer emphasizes the importance of verifying methodology to avoid misleading results. In August, I published a number on my studio’s website: we had been named unprompted in 14 out of 18 blind answers across ChatGPT, Perplexity, and Gemini. It was wrong. And the way it was wrong turned out to be more valuable than the number itself. I wrote a set of questions and pasted them into a single session on each engine. Most of the questions never mentioned my studio by name. Two of them did, because I also wanted to check whether the engines had the facts right. The problem seems obvious now. Every question in that paste shared the same context window. Once the studio’s name appeared in one question, it became available to the engine while answering the others. So I was not really measuring whether an engine could discover and recommend my studio unprompted. I was measuring whether it could read the information already sitting in front of it. At first, I missed the flaw completely. A near-perfect result was exciting, and I wanted to believe it. But then I noticed something strange: text and details from one answer were being carried into others. That was the giveaway. The prompts were contaminating each other, which meant the test was not genuinely blind. I repeated the test using two completely separate sessions for each engine. One session contained only the blind discovery questions. The other contained the named fact-checking questions and was opened separately in a fresh incognito window. Across ChatGPT, Perplexity, and Gemini, this gave me 36 genuinely blind answers. The corrected result was nine mentions out of 36 answers: 25%. That was a significant drop from the original 14 out of 18. More importantly, the visibility was not consistent across all three engines. The studio was effectively performing on only one engine out of three, rather than being discovered uniformly across ChatGPT, Perplexity, and Gemini. The engines that scored zero still knew the correct facts when I asked about the studio directly. The information was available. The problem was that those engines were not using the studio’s own website when deciding what to surface unprompted. Instead, they appeared to rely more heavily on third-party roundups and external mentions. That distinction matters. An engine knowing your brand exists is not the same as the engine choosing to recommend it. Keep blind discovery tests and named fact-checking tests in completely separate sessions. Do not allow the brand name, website, or identifying details to enter the context window before the blind questions are answered. Most importantly, verify the methodology before celebrating an increase in your score. A strong result built on a contaminated test is not visibility. It is context. The full protocol, the twelve questions verbatim and every result including the misses are published at reidify.design/research/measuring-ai-visibility https://reidify.design/research/measuring-ai-visibility .