An AI search citation decay study is travelling through the press this week under a stark headline: 57% of domains cited in ChatGPT and Google are never cited again. The figure is real — it appears in a vendor study of the Australian insurance category by Somantra AI, a Sydney-based AEO platform — but the sentence it usually travels in is not the sentence the study’s own tables support.
This matters beyond one press release. Citation-decay claims are becoming budget arguments: if most cited domains really do vanish after one citation, content strategy, measurement and reporting all change. Before a number like this reaches a slide deck, it deserves the same treatment any measurement deserves — read the tables, recompute the arithmetic, and check what population each percentage describes.
So that is what this review does. We read every public document the vendor has published about the study, recomputed every figure that its own tables allow, ran a bounded survey of how the finding is being covered, and assessed the study against ten specific, checkable questions. The result is the table in Section 05 — built to be fair enough that the study’s own authors could read it without objection, because several rows credit the study for disclosures a typical vendor release never makes.
- 01The August 17 item is not a new study.The study was first released August 3, 2026, and its report page is dated July 28. The item published Monday, August 17 is a follow-up press release repackaging one finding — the “content graveyard” — from the same dataset. No new data was added.
- 02The methodology is published; the fair criticism is narrower.A substantial report page discloses the dataset schema, named formulas, eleven insight sections, monthly breakdowns and the study’s own limitations. What is missing is different: no peer review, no raw data, no query set — so nothing is independently auditable.
- 03The survival table sums to 10,446 domains, not 28,725.The table behind the 57.2% and 2.7% figures totals 10,446 — exactly matching the study’s own ChatGPT-only plus both-platform domain counts. Which denominator the percentages describe is an open question the vendor could settle in one sentence.
- 04Scope rides with the number, every time.One country (Australia), one vertical (insurance), seven non-consecutive observed months — April and June 2026 are missing from the study’s own month list, unexplained on any page we reviewed. The study’s fine print states the scope; the syndicated headline drops it.
- 05Coverage is thin and wire-heavy.Our bounded three-query survey found 8 wire-service pickups and exactly 2 independent editorial pieces. The unscoped version of the number travels mainly through wire copy.
01 — The TimelineWhat was actually published — and when. #
Start with the date, because it is itself part of how the number is travelling. On Monday, August 17, 2026, a GlobeNewswire release appeared under the headline claim that 57% of domains cited in ChatGPT and Google are never cited again. Coverage written from that release reads as if a study landed that day. It did not. The underlying study — “AI Search Ranking Factors for Insurance Brands: The Australian Citation Report” — was first announced on August 3, 2026, and the vendor’s own full report page carries a July 28 date in its page metadata. The August 17 item is a follow-up press release repackaging one finding — “Insight 5: The Content Graveyard” — from the same dataset. Not a new study, not new data.
The dataset itself, as the vendor states it: 2,437,107 citation records across 28,725 unique domains, collected from ChatGPT and Google, restricted to the Australian insurance category, across seven observed months between November 2025 and July 2026. Those headline counts internally check out — the study’s platform split (4,924 domains appearing only in ChatGPT records, 18,279 only in Google, 5,522 in both) sums to 28,725 exactly. The survival percentages are where the arithmetic gets interesting, and Section 04 walks through why.
2,437,107 records
The vendor-stated total across ChatGPT and Google, restricted to the Australian insurance category, spanning seven observed months between November 2025 and July 2026.
The headline population
Vendor-stated split: 4,924 domains appear only in ChatGPT records, 18,279 only in Google, 5,522 in both. The three counts sum to 28,725 exactly — this part of the arithmetic checks out.
ChatGPT share of records
346,172 of the 2,437,107 records are ChatGPT; 2,090,935 (85.8% of 2,437,107) are Google. The schema’s “google.” platform value carries no qualifier — AI Overviews, AI Mode or organic results is not stated.
gappedseven-month window — has usually fallen off entirely.
02 — MethodologyHow this review was produced. #
A methodology critique earns nothing if its own method is vague, so here is ours, stated precisely enough to redo. Everything in this post rests on documents anyone can fetch, plus arithmetic anyone can re-run in two minutes against the same public pages.
What was collected. Every public-facing document Somantra AI has published about this study: both GlobeNewswire releases (August 3 and August 17, 2026), the full report page at somantra.ai, and the vendor’s homepage for product-methodology context. Plus independent material: TechTimes’ August 17 article, Insurance Business Australia’s July 22 article on the same underlying dataset, and DeltaV Digital’s separately run AI-citation study and its methodology page.
As-of date. All documents reflect their published state as of August 17, 2026, at the time of writing. The study’s own data period is November 2025 – July 2026 (seven observed months, non-consecutive).
Recomputation rule. Every statistic in the vendor’s report that could be recomputed from a table the vendor itself published was recomputed by hand and checked against the prose claim built on it. All three discrepancies found are shown in Section 04 — none is hidden, and none is inflated.
Coverage survey. Three targeted queries run against a general web search index on the assessment date, each returning roughly the first five to eight results; every distinct result was classified by hand. This is a bounded sample, not an exhaustive census — full counts in Section 07.
What was excluded. The vendor’s client case-study page (a testimonial, not a methodology source), and social-platform discussion (not surveyed — no claim is made about it either way).
Known limitations. We have no access to the raw data, so every check is internal-consistency arithmetic on published tables, not an independent re-measurement. The coverage survey’s search method returns partial results per query; absence findings are narrowed to exactly what was checked.
03 — Credit FirstWhat the study does disclose. #
Be fair before being critical: this study is substantially more documented than a typical vendor press release. The report page is a genuine methodology document — it discloses the record schema (month, year, platform, domain, URL, position, title, description, source), names the formulas behind its proprietary metrics, publishes eleven insight sections with supporting sub-tables and monthly breakdowns, and flags several of its own limitations. Its most important self-disclosed caveat: observed ChatGPT citation volume swung roughly 29x between the smallest month (4,682 records in February 2026) and the largest (136,494 in January 2026), a difference the report attributes to changes in “the observed query set” and explicitly warns should not be read as brand momentum. That is real methodological honesty, and it is worth crediting.
The accurate criticism is therefore narrower than “the methodology is unpublished” — it is published. What is missing is independence: the study is not peer-reviewed, no raw data or API access is offered, and the actual query set is not released anywhere we could find. The schema itself leaves one material ambiguity: the platform field’s values are chatgpt.standard
and google.
— the ChatGPT value carries a sub-type qualifier, the Google value does not, and no page we reviewed states whether “google.” means AI Overviews, AI Mode or conventional organic results. For a study framed entirely around AI search citations, that is a load-bearing ambiguity. “Documented in detail” and “independently checkable” are different properties; only the first one holds here.
04 — ArithmeticThe denominator question. #
Here is the finding original to this review, and it comes straight from the vendor’s own published tables. The headline pairs “2,437,107 citation records across 28,725 domains on ChatGPT and Google” with the survival figures — 57.2% of domains cited in exactly one sampled month, 2.7% cited in all seven — in the same breath, implying the percentages describe all 28,725 domains. But the report’s own “Domain Survival Rates” table sums to exactly 10,446 domains (279 + 438 + 1,481 + 2,278 + 5,970). And 10,446 is not a random number. It exactly equals the study’s own ChatGPT-only domain count (4,924) plus its both-platform count (5,522). It also exactly equals the domain total in a second, independently summing table on the same page — the “Citation Concentration Tiers” table, which totals 346,172 citations, the report’s own stated ChatGPT-only record count.
Three separate arithmetic paths converge on the same narrower population, which is strong evidence that the survival percentages describe the roughly 10,400 domains that appeared in ChatGPT’s citation records — not the full 28,725-domain, both-platform population the headline names. But no sentence on any page we reviewed states the survival table’s population explicitly, so we phrase this precisely: it is an open question which denominator the 57.2% describes, most likely a scoping detail lost between the report and the press release rather than anything more. It is the kind of question the study’s authors could answer in one sentence — and until that sentence exists, “of domains cited in ChatGPT and Google” is a claim the vendor’s own tables do not clearly support at that scope.
| Stated in the study | Its own underlying numbers | Our recomputation | Reading |
|---|---|---|---|
| “Complete guide” content is 3.5x more common among one-citation domains than among seven-month survivors | Guide format: 2.0% of survivor domains vs 7.3% of one-citation domains | 7.3 ÷ 2.0 = 3.65x | Minor rounding gap in the vendor’s own material — small, but the kind of thing independent auditing would catch. |
| Brand homepage citations rose “12.7x” from March to July 2026 | Homepages: 1.4% of ChatGPT citations in March 2026, 12.8% in July 2026 | 12.8 ÷ 1.4 ≈ 9.14x (+11.4 points) | A larger discrepancy. Peripheral to the graveyard claim, but it sits on the same page and bears on the report’s self-checked arithmetic generally. |
| The 57.2% / 2.7% survival split covers “domains cited in ChatGPT and Google” | Survival table sums to 10,446 domains; the ChatGPT-scoped concentration table also sums to 10,446; ChatGPT-only (4,924) + both-platform (5,522) = 10,446 | Population appears to be the ChatGPT-side 10,446, not 28,725 | The open denominator question. Three arithmetic paths agree; no vendor sentence names the population either way. |
published in detailand
independently audited: a reader with a calculator found all three in an afternoon, and a pre-publication reviewer would have too.
05 — The AssetThe complete ten-question assessment. #
This table is the review. Every row asks one specific, checkable question of the study, states what the published material actually shows, and gives a plain verdict on whether that is sufficient to support the syndicated headline — “57% of domains cited in ChatGPT and Google are never cited again” — as phrased. Rows where the study comes off well are in here too; a critique that only keeps its wins is marketing, not measurement. All figures are vendor-stated unless marked as our recomputation.
| # | Question | What the published material shows | Supports the unscoped headline? |
|---|---|---|---|
| Scope and the denominator (rows 1–2) | |||
| 1 | Does the 57.2% describe the full 28,725-domain population the headline names? | Both releases pair “28,725 domains on ChatGPT and Google” with the survival figures. But the study’s own Domain Survival Rates table sums to 10,446 domains, which exactly equals its ChatGPT-only (4,924) plus both-platform (5,522) domain counts and matches a second ChatGPT-scoped table on the same page. No sentence we found names the survival table’s population. | Open question. The internal tables agree to the digit on a narrower population; one clarifying sentence from the authors would settle it. |
| 2 | Does the finding generalise beyond Australian insurance? | The study scopes itself to “the Australian insurance category,” and the follow-up release’s closing line states: “Findings reflect one category over a seven-month observation window and describe correlation rather than established causation.” | No — and the study’s own fine print says so. The gap is between that fine print and the headline’s unscoped phrasing. |
| Sampling and measurement design (rows 3–6) | |||
| 3 | Can monthly sampling distinguish “never cited again” from “not sampled again”? | The study’s own monthly table shows observed ChatGPT citation volume swinging roughly 29x — 4,682 records in February 2026 vs 136,494 in January 2026 — attributed to changes in “the observed query set,” with the study’s own warning that raw monthly counts “should not be interpreted as pure brand momentum.” | No. A domain whose one citation fell in a low-volume month had mathematically less chance of re-observation. This is the study’s own disclosed caveat, applied to the one claim the release copy did not apply it to. |
| 4 | Are the “seven observed months” consecutive? | No. The stated period is November 2025 – July 2026, but the study’s own month list is November, December, January, February, March, May and July — April 2026 and June 2026 are absent, with no reason stated on any page we reviewed. | No. “Seven months” reads as a continuous run; it is seven non-consecutive months in a nine-month span. A domain cited only in April or June would not register at all. |
| 5 | Does domain-level rollup conflate very different sites? | The disclosed schema captures citations at URL level, rolled up so a domain counts as cited in a month if any one page is cited once. The study’s own concentration table shows 8,771 of the 10,446 ChatGPT-side domains hold just 1–9 total citations each across the entire dataset. | Partially disclosed. The rollup is a real, checkable design choice — but a one-page site and a many-page publisher land in the same “vanished” bucket, which is not what the framing suggests. |
| 6 | Can a seven-month window support “never cited again”? | The study’s own release language is careful here: findings “describe correlation rather than established causation” and reflect “one category over a seven-month observation window.” Nothing claims the pattern holds beyond July 2026. | No. Accurate only as “not observed cited again within this gapped seven-month window.” The syndicated headline drops that qualifier; the release’s own body text does not. |
| Confounds and data quality (rows 7–8) | |||
| 7 | Is the format finding confounded by who publishes which formats? | The study’s own source-type table shows comparison sites earn 19.2% of their citations from comparison-format content vs 2.4% for insurers — the winning formats cluster on a business type structurally built around them. An independently run cross-industry study (DeltaV Digital) found no single format dominated universally: listicles took 61% of citations in B2B technology services, homepages 55% in local services, program pages 53% in higher education. | Not fully controlled. The correlation is real inside this vertical; format winners vary by industry elsewhere. |
| 8 | Were the 38 flagged spam domains included in the survival counts? | Independent reporting (Insurance Business Australia, July 22, 2026) covered a separate finding from the same dataset: 38 of the 28,725 domains flagged as spam or “grey-area parasite SEO operators,” accounting for 1.97% of ChatGPT citations vs 0.10% of Google citations — a roughly 19-fold platform difference — with one flagged domain alone drawing 5,366 ChatGPT citations. Neither release nor the report page states whether these 38 sit inside the survival counts. | Unresolved on the pages reviewed. A fair question the study’s own text does not answer either way. |
| Independence and verifiability (rows 9–10) | |||
| 9 | Does the vendor’s commercial interest change how to read the numbers? | The study’s publisher sells AEO/GEO brand-visibility tracking built on the same citation-tracking approach, plus a free brand-audit product; the follow-up release’s recommendations map directly onto what that product measures. One of the two independent pieces we found made the same observation in its own disclosure paragraph. | Not disqualifying on its own — vendor research is normal and does not make the numbers false — but the interest belongs in the same sentence as the number. |
| 10 | Is any of it independently verifiable? | The report page is genuinely detailed: exact counts, a record schema, named formulas, eleven insight sections and self-disclosed limitations. Not published anywhere we reviewed: the query list or count, the citation-detection method, what the schema’s “google.” platform value covers, any raw-data or API access, or any form of third-party review. | No. “Documented in detail” and “independently checkable” are different properties, and only the first is true here — which is still more than most vendor studies manage. |
06 — The VerdictsWhat the data supports — and what it cannot. #
A fair reading leaves real findings standing. The dataset is large, the single-vertical scoping may well be a deliberate design choice rather than a flaw — a controlled look at one category beats a shallow look at twenty — and a directional finding inside one vertical is genuinely useful to anyone operating in it. The distinction that matters is between what this dataset can carry and what the syndicated sentence asks it to carry.
Majority churn inside this dataset
Within Australian insurance, across the study’s seven non-consecutive observed months, a majority of the domains in its survival table were observed cited in only one month. That is a real, citable pattern — with its scope attached.
Format correlates with persistence here
On the study’s own numbers, discount (10.6% of survivors vs 5.4% of one-hit domains), comparison (8.0% vs 4.3%) and FAQ formats (1.4% vs 0.8%) persist at roughly 1.8–2.0x the one-citation rate in this vertical.
A web-wide vanishing claim
Nothing here supports a general claim about “AI search” broadly — one country, one vertical, and an open question about whether the percentages even cover both platforms. Cross-industry data shows format winners vary by industry.
Permanence
“Never cited again” is accurate only as “not re-observed within a gapped seven-month window whose sampled volume swung roughly 29x month to month.” The study’s own text calls the pattern correlational.
The interpretive point runs deeper than one study. Citation measurement is becoming an industry, and vendor datasets are — for now — most of what exists, which is why our meta-analysis of 54 AI-citation studies reads that field across many datasets rather than one. The trend that matters is not that vendors publish studies — it is that syndication strips qualifiers faster than vendors add them. The scoped, careful sentence in the release body and the unscoped sentence in the headline were written by the same organisation on the same day, and only one of them travelled.
07 — SyndicationHow the number is travelling. #
Coverage claims deserve the same discipline as any other claim, so here is ours with its method attached. We ran three targeted queries against a general web search index on the assessment date — each returning roughly the first five to eight results — and classified every distinct result by hand. That is a bounded sample, not an exhaustive census, and the counts below should be read that way.
Coverage survey results · classified by hand, as of the time of writing
Source: Digital Applied coverage survey — three targeted queries, bounded method stated in Section 02The picture is neither “widely covered” nor “ignored” — it is thin, wire-heavy syndication with two independent treatments. The most encouraging detail: TechTimes’ piece scoped the finding correctly and disclosed the vendor’s commercial interest unprompted, and Insurance Business Australia had already examined the same dataset from an entirely different angle nearly four weeks before the follow-up release. The unscoped version of the number travels almost exclusively through wire copy. One narrow absence finding, stated as narrowly as it was checked: a site-restricted search of three major SEO trade outlets on the assessment date returned no results for the study — not found via this method on this date at those three outlets, which is not a claim that no major SEO outlet has covered it.
“Disclosure: Somantra is both the author of the study and a commercial provider of AEO and GEO services and brand audit tools... Readers should weigh the findings accordingly.”— TechTimes editorial disclosure, August 17, 2026
08 — The Transferable PartThe same questions, asked of any citation study. #
A checklist only means something if it cuts both ways, so here it is applied — briefly — to a study this site has covered favourably. SparkToro’s zero-click study found 68% of US Google searches end without a click, and we covered it in June without running this kind of audit at the time. Scope: that study’s own headline carries its market qualifier — “68% of US Google searches” — in the same breath as the number, which is exactly the standard the Somantra headline misses. Commercial interest: SparkToro is itself a commercial audience-research vendor with a direct stake in a “clicks are disappearing” narrative — a structurally similar incentive — though its study is co-branded with a second, independent data partner rather than self-authored end to end. We did not re-run a full methodology audit of it for this post; the point is that the lens applies to studies whose conclusions we like, too.
The independent comparison point deserves its own precision. DeltaV Digital’s separately run study — 21,075 AI-engine responses across eight industries and five platforms, April 14 to July 13, 2026 — corroborates the directional “format matters” finding: comparison pages earned only 4.1% of its total citations but the highest per-retrieval citation rate of any format (1.87 vs a 1.29 portfolio average, 45% above). But it measures cross-sectional citation share over a fixed 90-day window; it does not track the same domain’s citation status month over month, so it does not test — and cannot confirm — the “cited once, then never again” persistence claim. It also publishes its own limitations, which is the habit this whole genre needs. We did not find a second independently run study that directly tests longitudinal citation persistence, within this review’s search budget — stated as an absence in what we could locate, not proof none exists.
what population is the denominator— and does the study’s own table sum to it? What is the scope — market, vertical, platform — and does it ride in the same sentence as the number? Is the observation window continuous, and can its sampling distinguish “gone” from “not sampled”? Who profits if the finding is believed? And could anyone outside the organisation reproduce the number? The Somantra study fails none of these for existing — it is more disclosed than most — but the headline built on it fails the first two.
The practical alternative to arguing about other people’s denominators is owning your own: track which of your pages AI surfaces actually cite, month over month, with a method you control. Our 100-point brand citation audit checklist is the self-serve version of that work, and our agentic SEO practice runs it as a service — measurement first, format changes second, and never on the strength of a single vendor headline.
09 — ConclusionRead the tables before you repeat the number. #
A better-than-average vendor study, carrying a worse-than-average headline.
The Somantra study deserves a fairer fate than its own press release gives it. The report page is a genuine methodology document — schema, formulas, monthly tables, self-disclosed limitations — and inside its stated scope it supports a real finding: in Australian insurance, across seven non-consecutive observed months, most domains in its survival table were observed cited only once. That is worth knowing if you operate in that market.
What the data does not support is the sentence doing the travelling. The survival table sums to 10,446 domains against a 28,725-domain headline — an open question the vendor could close in one sentence — the “seven months” have two silent gaps, the stated multipliers do not quite survive recomputation from their own inputs, and nothing is independently auditable. None of that is an accusation; all of it is checkable; and every check in this post can be re-run by anyone with the same public pages and a calculator.
Looking forward, expect more of this genre, not less — citation measurement is where rank tracking was two decades ago, with vendors racing to define the metrics they will then sell. The durable skill is not knowing which studies to trust; it is knowing which questions to ask of all of them. The ten in Section 05 are a reusable start.