{"slug": "chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post", "title": "ChatGPT Changed How It Picks Sources While You Were Reading My Last Post", "summary": "OpenAI removed the `result_source` field from ChatGPT's network traffic on 21 July and added a fifth retrieval pipe, `bing`, that is being rolled out to some accounts but not others, according to a teardown by Suganthan Mohanadasan. The same Bob Vila URL was fetched via Bright Data's `bright` pipe for Mohanadasan and via `bing` for GEO researcher David Konitzny, showing the rollout is cohort-gated. The change retires the previous pipe labels and moves `supporting_websites` onto citation items, with FanoutFox tracking the format changes.", "body_md": "# ChatGPT Changed How It Picks Sources While You Were Reading My Last Post\n\nA reader found a source pipeline I'd never seen, so I went back into ChatGPT's network traffic 10 days after the first teardown. Some of it had changed underneath me, I caught 2 of my own over-reaches, and I found the layer that shows who you lose citations to.\n\n## The 22 July update\n\nThis post ends with a warning that everything in it carrying a proper noun has a shelf life measured in days. Eight days after publish, the shelf gave out. On 21 July OpenAI deleted `result_source`\n\nfrom the payloads, the field every pipe label below hangs off. The next day the shape moved again, and thinking answers now pack their sources into per-domain groups under `search_result_groups`\n\n, with the reference id switching from a string to an object.\n\nThe honest state of each layer, so you know what you’re reading. The pipe labels are gone from fresh traffic, deleted rather than renamed, which retires the census, the cohort map and the bing hunt below from live experiments to history. The runner-up layer survived and moved house; `supporting_websites`\n\nnow rides on the citation items instead of the retrieval entries. Fan-out queries, sources, snippets and citations all still read fine once you know the new shapes. Spelunking by hand on a fresh conversation, search the stored JSON for `search_result_groups`\n\nand expect entries of `type: \"search_result\"`\n\n; a search for `result_source`\n\ncomes back empty.\n\nThe mechanism didn’t stop existing when its label stopped shipping. Commercial fetchers still do ChatGPT’s reading, you’ve just lost the window showing which vendor got routed your way. [FanoutFox](https://fanoutfox.com/) reads all three generations of the format, its [changelog](https://fanoutfox.com/changelog/) is the running record between teardowns, and everything below stands as what it always was, a dated snapshot of July’s plumbing. This post’s own rule applies, date every claim, including mine.\n\nSomeone tagged me on a [Linkedin post](https://www.linkedin.com/posts/davidkonitzny_bing-is-now-officially-a-result-source-in-share-7478896640515354624-ZQRk/?utm_source=share&utm_medium=member_desktop&rcm=ACoAAA5HxykBBrrELSaiCU5K68YkjhUBR7L7CPg) by David Konitzny, a GEO researcher.\n\nHe posted a screenshot of his own network tab showing a `result_source`\n\nvalue I’d never seen in 2 days of staring at this stuff. `bing`\n\n.\n\nBing, sitting right there in ChatGPT’s traffic reproducible across multiple prompts on his account.\n\nMy first thought was that I’d missed it. I hadn’t. It’s better than that.\n\n## The same page came through 2 different pipes\n\nDavid’s screenshot showed `bing`\n\non pages from bobvila.com, a lawn mower review.\n\nSo I did the most direct experiment available. I asked my ChatGPT about Bob Vila’s electric mower rankings and aimed it at the exact page in his screenshot, then pulled the traffic.\n\n`bobvila.com/reviews/best-electric-mower/`\n\nreached me through `bright`\n\n. Bright Data, the scraper from [Part 1](https://suganthan.com/blog/how-chatgpt-picks-sources/). Every one of the 20 results in that thread came through `bright`\n\n.\n\nThe same URL, in the same week, arrived at David’s browser stamped `bing`\n\nand at mine stamped `bright`\n\n.\n\nI went wider to make sure. A census across my last 30 conversations counted `bright`\n\n558 times, `labrador`\n\n21, `serp`\n\n16, and `bing`\n\nexactly 0. I then grepped the 1.2MB of feature-flag config ChatGPT caches in my browser for anything bing-related. Nothing.\n\nThe rollout isn’t even being served to my account.\n\nSo `bing`\n\nis real, it’s new since late June, and it’s cohort-gated. OpenAI added a 5th retrieval pipe and is switching it on for some accounts and not others, which is a strange sentence to type about the most-watched product on the internet, and also completely normal engineering.\n\n### The AI SEO/GEO takeaway\n\n#### Sort out your Bing indexation\n\nIf you land in a `bing`\n\ncohort, your Bing index suddenly matters. Pages missing from Bing’s index are missing from those users’ ChatGPT, so Bing Webmaster Tools plus a sitemap check is cheap insurance. It’s cohort-gated and could be pulled at any time, so file it under watch, not panic.\n\n## The runner-up layer\n\nThe best new find of the re-test is a field Part 1 dismissed in a single breath. `supporting_websites`\n\nsat empty in every June capture, so I catalogued the name and moved on. It’s populated now, and it’s quietly the most useful thing in the whole payload.\n\n**The wire now shows you, per claim, exactly who beat you and who you’re beating.** Not at the vague level of “competitor X gets cited more”, but this sentence, this claim, you were the runner-up, here’s the page that won.\n\nWhen ChatGPT cites a source for a claim, that citation now carries an array of other pages that supported the same claim but didn’t win the visible slot. Runner-up citations, each with its own `result_source`\n\n, sitting invisibly under the winner.\n\n*(Since 21 July the field rides on the citation items themselves rather than the retrieval entries, and the per-runner-up pipe label went with the rest of the pipe layer. Same data, new address; the update at the top has the map.)*\n\nThey come in 2 flavours.\n\nThe first is the domain fold from Part 1 made visible. Ask about Emirates baggage rules and the cited emirates.com page carries other emirates.com pages underneath it, the siblings that got folded into the domain’s 1 slot, including a German-locale duplicate of the exact same page. I watched a Range Rover overview page win a spec claim with the brand’s own `/electric-range`\n\npage demoted to support beneath it, and a canada.ca visa page beat a sibling canada.ca page the same way.\n\n* The 20-thin-pages problem isn’t a theory anymore*. You can watch your own pages lose to each other. If 20 thin pages cover 1 topic, 19 of them are losing to the 20th, so build 1 strong page per claim, not a pile of weak ones.\n\nThe second flavour is competitive, and that’s the one that gave me proper Spidey sense goosebumps.\n\nIn a thread about crawler verification, the cited page was `ahrefs.com/robot`\n\n, and sitting underneath it as support were both `semrush.com/bot`\n\nand `mj12bot.com`\n\n, all 3 fetched through `bright`\n\n. One claim, 3 rivals, 1 visible winner and 2 invisible runner-ups.\n\nIt crosses the tiers too. In a lawn mower thread, bhg.com won the citation through the licensed `labrador`\n\npipe with protoolreviews.com supporting it through the `bright`\n\nscraper, a licensed source and a scraped one backing the exact same claim, one on screen and one hidden. Elsewhere a Cloudflare blog post won with The Verge’s coverage of the same story tucked beneath it, a scraped source beating a licensed one for the same sentence.\n\nI asked “best keyword clustering tool in 2026” and read the winners and the runner-ups straight off the console.\n\nFor GEO this is a genuinely new instrument. When the visibility tools start surfacing this field, and they will, “second place per claim” becomes a metric you can work against.\n\n### The AI SEO/GEO takeaway\n\n#### Read your own runner-up data\n\nRun your money queries, pull the stored conversation, and study how the winning page states the claim you lost, in the exact words that won it. That’s a per-claim competitive audit no paid tool offers yet, and it’s the whole reason I ended up building the extension at the end of this post.\n\n## The prompt you can’t read\n\nOne more find, and I’m including it precisely because I can’t fully open it.\n\nEvery conversation now carries 2 hidden system messages, invisible in the interface, flagged in the metadata as `identity_prompt`\n\nand `sources_and_filters_prompt`\n\n.\n\nSit with that second name for a moment. ChatGPT injects a **dedicated prompt** about **sources and filtering** into every conversation before it answers.\n\nThe instructions for how it should treat sources exist as a discrete, named object, and I can prove the slot is there, 64 instances across my last 30 conversations. What I can’t do is read it. The content ships empty to the browser in both the live stream and the stored conversation. OpenAI strips the text server-side and sends only the envelope. I can prove the safe exists. I can’t open it because that text never leaves OpenAI’s servers.\n\nBut I can’t read the instruction and still watch it run.\n\nLook at what the fan-out actually types. It doesn’t just reword your question, it appends the kind of source it wants. “Emirates DIFC city check-in cost fee queue Dubai 2026 **Emirates official**”. “** AhrefsBot official**”. “** official crawler IP ranges**”. That “official” suffix turned up 17 times across my census, bolted onto queries for facts. That’s the sealed source-filtering prompt executing in the open. For a factual claim it’s told to go looking for the official source, and you can watch it go looking.\n\nThe implication is bigger than the find itself in my opinion.\n\n**Source selection is a policy, not a behaviour.** It lives in an instruction OpenAI can rewrite any day, without retraining a model or telling anyone. That’s the mechanism behind half this post, and it’s why the plumbing moved within 10 days of Part 1.**Nobody can honestly sell you “ChatGPT’s ranking factors”.** The real instructions are a named object that never leaves OpenAI’s servers. Everyone claiming to know them is guessing, and now you can say so with a straight face.**You can still read the policy, just not the text.** Every regularity in this series, facts routed to official pages, opinion to reviews and Reddit, “official” quietly appended to searches, is the prompt being executed in front of you. Watch the behaviour, treat each pattern as a sentence in a memo you’re reconstructing from the outside, and when the behaviour changes, assume the memo got a new draft and update your playbook with it.\n\n### The AI SEO/GEO takeaway\n\n#### Be the official page for your own facts\n\nYou just watched the prompt hunt for “official” sources on factual queries, so be the official source. Facts route to official pages, opinion routes to reviews and Reddit, so keep your pricing, specs and docs in plain HTML on pages that are unmistakably yours, and let third-party coverage fight the opinion battle for you. Own the facts. If you won’t publish enterprise prices, at least say “Starts at $5000” instead of “Contact sales”, because a random Reddit comment about your product costing $20,000 gets stated as fact by ChatGPT when your own page won’t say otherwise.\n\n## 3 accounts, 3 different ChatGPTs\n\nBoth of those finds came off my account, so here’s the sobering counterweight, and it decides how much you can trust any of this.\n\nAnother reader, Simone De Palma, had already challenged Part 1 from the opposite direction. He’s on free ChatGPT in Italy, and on his account every single publisher citation carries `labrador`\n\n, including small Italian sites that are nowhere near the “licensed national newspaper” club I described. He’d never seen `bright`\n\nor `oxylabs`\n\nat all and politely wondered where my tiers were coming from.\n\nPut the 3 of us side by side.\n\nAccount | What result_source shows |\n|---|---|\n| Simone, free tier, Italy | `labrador` on everything, no scraping vendors ever |\n| Me, Plus, UAE | `bright` dominant (558 of 595), thin `labrador` , no `bing` |\n| David, Germany | `bing` appearing, reproducible across prompts |\n\nSame product, same field, 3 completely different pictures. And it drifts over time on a single account too. In my June capture `oxylabs`\n\nwas live and `serp`\n\nhad gone quiet. 10 days later `oxylabs`\n\nhas vanished from my traffic entirely and `serp`\n\nis back. Nobody announced any of this. The plumbing just moved.\n\nWhich means I owe Part 1 a correction. I described `labrador`\n\nas a licensed tier you can’t join unless you own a national newspaper. The licensing deals are real and documented, but the tier reading was me standing in one spot and mistaking my view for the map. On free accounts `labrador`\n\nappears to be the default pipe for everyone, tiny Italian sites included. On my paid account it narrows to major media while Bright Data does the heavy lifting. The field doesn’t grade quality. It names which fetcher OpenAI routed your page through, for that user, in that cohort, that week.\n\n### The AI SEO/GEO takeaway\n\n#### Don’t extrapolate from one account\n\nFree tier, paid tier and different countries see different retrieval mixes for the same page. Audit your visibility from a single account and you’re auditing a cohort, not the product. Vary the tier and the geo before you declare a win or a crisis, and be just as sceptical of anyone selling you one screenshot as the truth.\n\n## I re-tested every claim, and most of the specifics had already moved\n\nGoing viral with a teardown of a system that ships weekly is a good way to develop paranoia.\n\nSo instead of writing a victory lap, I spent an evening re-running the whole thing. A 30-conversation census, 4 fresh test queries, and 2 live stream captures, checking each claim from the first post against what the wire says now.\n\nHow to read this.Everything sorts into 3 buckets and I label them as I go.The platform moved, things that genuinely changed in the traffic between 23 June and 4 July.I over-reached, things Part 1 stated more confidently than the data deserved, which I’m correcting.It held, claims I re-tested that survived. Structural facts still come from single clean captures. Anything with a count is still 1 account and directional.\n\nPart 1 claim | 10 days later |\n|---|---|\n| 4 retrieval pipes (serp, labrador, bright, oxylabs) | Platform moved. Now at least 5. `bing` is rolling out per cohort, `oxylabs` has left my mix, `serp` came back. |\n| labrador is a licensed club you can’t join | I over-reached. Real licensing, wrong tier theory. It’s the free-tier default and a per-cohort routing decision. |\nThe `text` bucket skips the web entirely | Held. Confirmed live, with a footnote below about where that field actually lives. |\nThe fan-out runs `site:` probes and reformulates queries | Held, and improved. The stored fan-out now shows it appending source-type hints, “Emirates official”, “AhrefsBot official”. It doesn’t just rewrite your question, it specifies what kind of site should answer it. |\n| Citations bind per claim, dedupe by domain | Held, and upgraded. The dedupe mechanism became visible in the re-test. |\n| News citations can’t be resolved to sources | Platform moved. The news pool now persists in the stored conversation. OpenAI shipped a fix between my captures. |\n| No ranking score anywhere in the traffic | I over-reached, slightly. True for web results. The local pipeline has a `ranking_score` slot, but it came back null in both my tests, so no exposed score, just a narrower claim. |\n| Local surfaces only 2 places | Weakest claim standing. The config that said 2 went dark (more below) and the map payload carries 12 to 28 places even when few render. Treat Part 1’s line with care. |\n| You’re 1 of 573 experiments | Platform moved. The feature-flag config is now hash-obfuscated. The gates are still there, the names and count are no longer readable. |\n\n2 smaller confessions for completeness. Part 1 said a per-result `weight`\n\nfield existed, then my re-test declared it gone, and both were sloppy. It’s a float on messages, not results, and my first search pattern only matched whole numbers. And the labrador snippets I called “near full article” are * now even longer*, averaging 1,217 characters against Bright Data’s 153, so that asymmetry strengthened.\n\nThe product layer also woke up. In June, product cards only rendered on queries ChatGPT classified as shopping. In the re-test, an ordinary “best lawn mowers of 2026” question came back with full merchant cards, priced in dirhams through a UAE reseller, because the commerce feed now geo-localises and fires on plain questions. The Levi’s Jeans 500 placeholder from Part 1’s bonus section is earning its keep.\n\nOne technical footnote that will save replicators some pain. A few of Part 1’s fields, `turn_use_case`\n\nand the per-URL moderation checks among them, never appear in the stored conversation you pull from the API. They exist only in the live response stream while the answer renders. My re-test initially declared them dead, and they’re not, they just live in a layer you have to catch in the moment. If you replicate this and a field seems missing, check which layer you’re looking at before writing your own correction post.\n\n### The AI SEO/GEO takeaway\n\n#### If you sell, mind the commerce feed\n\nProduct cards now fire on plain “best X” questions and localise pricing to the buyer’s market, reseller listings and all. Feed hygiene and per-market merchant presence just stopped being a shopping-only problem and started mattering on informational queries.\n\n## Build for what holds, not for what I measured\n\nThe uncomfortable part isn’t that Part 1 needed corrections. It’s the speed. Between my capture on 23 June and David’s screenshot on 3 July, OpenAI added a retrieval vendor, benched another, started persisting the news pool, opened the commerce feed to plain queries, populated a dormant field, and obfuscated the config that used to leak its experiment count. 10 days.\n\nThat pace sorts every GEO claim into 2 piles, and the sorting matters more than any single finding.\n\nThe durable pile is mechanisms. ChatGPT reformulates your query and steers it toward source types, classifies intent before deciding whether to search at all, binds citations to individual claims, dedupes by domain, reads your official pages for facts and everyone else’s for opinion, and rewards being cleanly scrapable. Those behaviours survived every re-test, they show up in some form on the other engines too, and building for them is building on rock. Most of the levers in this post fall straight out of them: consolidate to 1 page, be the official source, read your runner-up data. One more is worth its own heading, because it’s the one people quietly get wrong.\n\n### The AI SEO/GEO takeaway\n\n#### Stay reachable by the scrapers, not just the branded bots\n\nWhich vendor fetches your pages is perishable. The mechanism underneath isn’t. Commercial scraping networks are one of ChatGPT’s retrieval pipes, and on my paid account they carried nearly everything. A firewall rule that blanket-blocks scrapers can quietly cut you out of an entire cohort’s pipe, so don’t geo-block your own pages and check your WAF rules, the Cloudflare policies especially. Here’s [my note on firewall blocks](https://suganthan.com/notes/firewall-blocks-ai-crawlers/), and 2 free tools I built to check it the way a crawler actually sees you.\n\n[AI Crawler access checker](https://suganthan.com/free-seo-tools/ai-crawler-access-checker/). When someone asks ChatGPT or Perplexity about your topic, a crawler fetches your pages to build the answer. Three things quietly stop it: robots.txt disallowing a bot, a firewall challenge a bot can’t solve, or a server too slow before it gives up. This checks a domain the way a crawler sees it and shows you what’s getting through.\n\n[499 Timeout risk checker](https://suganthan.com/free-seo-tools/499-checker/). When ChatGPT, Claude or Perplexity needs your page mid-answer, a live fetcher requests it right then, and it doesn’t wait around. Too slow and the fetcher abandons the request, your server logs a 499, and the answer gets built from someone else’s page. This times any URL the way those fetchers experience it, cold cache included.\n\nThen there’s the perishable pile. Everything with a proper noun or a number in it. Which vendor fetches your pages, what percentage came through which pipe, which tier a field implies, what a config value is set to. That pile has a **shelf life measured in days**, and the industry habit of laminating a screenshot into a permanent “how ChatGPT works” slide is how agencies end up presenting a June diagram of plumbing that got re-piped in July.\n\nSo date every claim you rely on, including these. If a slide, a tool or a consultant tells you how ChatGPT picks sources and there’s no capture date on it, treat it as folklore.\n\n**Build on the mechanisms. Treat every vendor name and every percentage as weather, including mine.**\n\n## Check it yourself in 2 minutes\n\nThe [Part 1](https://suganthan.com/blog/how-chatgpt-picks-sources/) recipe still works as a method, DevTools, Network tab, read the stored conversation. The search string changed on 21 July though. `result_source`\n\nreturns nothing on fresh conversations, so search for `search_result_groups`\n\ninstead and expect entries of `type: \"search_result\"`\n\n. Part 1’s console snippet now prints an empty pipe table, and with it went the cohort mapping; the bing hunt closed early, not because anyone answered it but because OpenAI switched off the light.\n\nJust remember the layer rule from above.\n\nThe pipe labels persist in stored conversations from before 21 July.\n\nThe query classification and moderation fields only exist in the live stream, so catch them while the answer is still rendering or they’re gone.\n\nDoing that by hand for every answer is exactly as tedious as it sounds, so I built the whole recipe into a Chrome extension.\n\nI call it FanoutFox.\n\nFanoutFox reads your own ChatGPT session and lays out the fan-out queries, the sources with their stored snippets, what got cited against the runner-ups parked underneath, and which brands the answer named. On captures from before the format change it still shows the retrieval pipe on each source.\n\nIt hooks the live stream too, so it catches the fields that vanish once the answer finishes rendering. Everything in this post, in 1 click, without the Console, and it survived the drift week because its parser reads all three generations of the format. It’s free and it stays in your browser. While the store build waits out Google’s review queue, the current version installs from [fanoutfox.com/download/](https://fanoutfox.com/download/), or run the recipe above by hand if you’d rather read the raw wire yourself.\n\n## What’s next\n\nWhile validating all this I finally captured the one surface the first post never touched.\n\nDeep Research.\n\nIt runs under its own internal codename, fired 38 searches off a single question of mine, and the pattern of what it cites is different enough from regular ChatGPT that it changes what kind of page you should build for it.\n\nIt also travels over a channel my original capture setup couldn’t even see, which is a story in itself.\n\nThat teardown is the next post.\n\n*Baseline captured 23 and 24 June 2026, re-tested with fresh captures on 4 July 2026, on my own logged-in ChatGPT account. Same rules as Part 1. Structural findings are read straight from the traffic and solid at a single capture. Counts are 1 account and directional. Credit where it’s due, David Konitzny found the bing pipe, Simone De Palma’s free-tier data broke my tier theory.*\n\nEnjoying this? Get the next one in your inbox.\n\nI email when I publish something new. Usually 2 to 4 emails a month. Unsubscribe anytime.\n\nJoin 1,210+ readers. Unsubscribe anytime.\n\n[Suganthan Mohanadasan](/about/)\n\nNorwegian entrepreneur with 20+ years in SEO. Co-founder of Keyword Insights and Snippet Digital. Based in Dubai.", "url": "https://wpnews.pro/news/chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post", "canonical_source": "https://suganthan.com/blog/how-chatgpt-picks-sources-part-2/", "published_at": "2026-07-14 00:00:00+00:00", "updated_at": "2026-08-10 12:18:50.001551+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-infrastructure"], "entities": ["OpenAI", "ChatGPT", "Bright Data", "FanoutFox", "David Konitzny", "Suganthan Mohanadasan", "Bob Vila"], "alternates": {"html": "https://wpnews.pro/news/chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post", "markdown": "https://wpnews.pro/news/chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post.md", "text": "https://wpnews.pro/news/chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post.txt", "jsonld": "https://wpnews.pro/news/chatgpt-changed-how-it-picks-sources-while-you-were-reading-my-last-post.jsonld"}}