# The State of AI Disclosure 2026: what 1,088 EU sites' chat widgets say

> Source: <https://disclosureproof.com/research/state-of-ai-disclosure/>
> Published: 2026-08-25 20:46:06+00:00

# The State of AI Disclosure 2026

Eight days after Article 50 of the EU AI Act began to apply, we swept our locked cohort of 1,142 detector-flagged, EU-facing sites — 1,088 scanned, 54 excluded by robots.txt and counted, each rendered in a real browser, homepage only, read-only — on a staffed Monday, early afternoon in EU time. Corrected for the detector's measured error, roughly 11% of the 12,285-site frame shows a real chat launcher (the frame and its correction were measured 27–28 July; the surfaces below, on the scan date). Most of what happens at the first interaction is invisible to our automated visitor: on 78% of confirmed widgets it could not read the first message at all, and on the 174 surfaces it could read, a disclosure that the chat is automated was detected on 19 — about one in ten. This page is a single post-application snapshot: it publishes no before/after delta and no compliance verdicts, and every exclusion is counted below.

≈11% of 12,285 EU-facing candidate sites show a real chat launcher once measured detection error is corrected (95% CI 9.0–14.3); the raw detector said 9.4%. Frame and correction measured 27–28 July 2026 — this figure reads no sweep data and is not dated to the scan day.

## How to read these numbers

Every figure on this page is an aggregate. No individual site is named, and none will be:
the point is the state of the ecosystem, not a wall of shame. We hold that a raw detector
count without its precision and recall is not a finding — and by that standard exactly
one rate here is published *corrected* for its detector's measured error, with the
error rates themselves in the open: the prevalence headline. The other numbers below are
raw, and are labelled as such where they appear. The widget detector's error *was*
measured, but it is not corrected out of the coverage figure, whose denominator it
contaminates in a direction we name in that section. The *disclosure* detector
— the one deciding whether a first message discloses AI — has no measured
precision or recall at all: no judging pass in this study ever read disclosure wording, so
the three disclosure-state counts are raw detector output whose error rate is unknown rather
than known to be small. Where we report disclosure, a finding of “not detected”
means our scanner, loading the site the way *this study's automated visitor* does
— which is not the way a human visitor does, and this study ran no human arm, so it
cannot tell you how far apart the two are — could not observe a disclosure at the
widget's first interaction: a statement about what was observable, never a legal conclusion
about any organisation. We report three outcomes only:
**detected**, **not detected**, and **could not
verify** — and we publish the could-not-verify rate instead of hiding it.
[Why we never say “compliant”](/why-we-never-say-compliant/) applies
to this study exactly as it applies to a report.

## How many EU-facing sites run a chat widget?

The corrected figure above comes from auditing our own detector, not from trusting it. Both admission rules were checked against what a blind judge could see in a real-browser screenshot of the rendered page — that capture is the ground truth here, and it is not the same thing as what a site actually runs, nor what a human visitor would find — and the raw detector count is adjusted by the measured rates below. The two errors run in opposite directions and do not cancel; the correction and its confidence interval carry both.

One defect in that interval's history is disclosed here rather than left for a reader to
find. The bootstrap behind the interval this study first drafted drew the same sampling
variance twice, so the interval it produced — 7.95–15.42 — was labelled
95 % but sized closer to 99 %. The estimator was frozen under precommitment until
the second sweep completed, and the one-line fix was applied after it, on 10 August 2026;
the correctly sized 95 % interval is the one printed above, and the point estimate did
not move. Note which way the fix cut, because it ran *against* this page rather than
for it: the load-bearing claim here is that the corrected figure is not distinguishable from
the raw detector's, and the narrower interval makes that claim harder to sustain, not
easier. The raw figure remains inside the corrected interval — the claim survives the
fix, by less room than the draft interval suggested. Both intervals are stated here so the
correction is checkable rather than silent.

| Correction component | Measured on | Rate (95% CI) |
|---|---|---|
| Generic-launcher admits showing a real launcher | 49 decidable | 47% (34–61) |
| Vendor-fingerprint admits showing a real launcher | 52 decidable | 73% (60–83) |
| “No widget” sites showing a launcher anyway (recall gap) | 322 decidable | 6.8% (4.6–10.1) |
| Inert-SDK sites showing a launcher anyway | 41 decidable | 12% (5–26) |

## By widget vendor

Vendor shares are of sweep-confirmed widgets. One caveat travels with this table:
**“the vendor's SDK is present” and “a visitor can chat” are
different claims.** A substantial share of the sites whose vendor SDK we
fingerprinted showed no chat our scanner could reach. We cannot tell you why, and we will
not guess: that gap is a single undecomposed quantity mixing our own detector's false
positives, inert or disabled SDKs, staffed hours, login gates and several days of site churn
between the two measurements. This study forbids itself from republishing it as any one of
those causes — in particular as an after-hours or weekend penalty — and its size
is not distinguishable from the detector precision published above.

Two limits bear on this table specifically. First, a consequence of the frame,
pre-registered before the data existed: candidates are taken from the top of a popularity
ranking, which over-represents enterprise sites relative to the long tail, so SMB-leaning
vendors are under-represented here against their real install base. The unknown / custom
cell is precisely the one a rank-truncated frame inflates, because the vendors that
concentrate below the rank cut never enter the study at all. Second, the SDK-versus-reachable
caveat above is scoped to the *fingerprinted* rows, while the detector contamination
sits almost entirely in the other class: the sites matching no vendor fingerprint are the
ones our admission audit found least likely to be chat widgets in the first place.
“Unknown or custom” is therefore a statement about what our fingerprints matched
— it is not a measurement of how much of the ecosystem is custom-built, and correcting
it for measured precision would move it down. Vendor cells below the minimum-cell threshold
are grouped as “other / n too small” rather than published individually; the
country section's note on how the pre-registered threshold is applied covers this table
too.

| Vendor | Sites | Share of confirmed |
|---|---|---|
| unknown / custom | 570 | 72% |
| zendesk | 60 | 8% |
| livechat | 33 | 4% |
| other (13 vendors, each n < 30) | 131 | 16% |

## Can an automated visitor reach the first message at all?

On 78.1% (95% CI 75.1–80.8) of the 794 sites where the sweep confirmed a chat widget, our automated visitor could not read the first-interaction surface. Raw share on the detector-confirmed denominator, scanned Monday 10 August 2026 (graded 11:01–12:12 UTC).

This number is usually buried as a denominator footnote; we measured it on purpose and publish it as a headline. For every confirmed widget the sweep tried, read-only, to dismiss the consent layer, click the launcher and read the first message, and recorded where that failed. Consent walls, launchers our scanner could not click, panels that never open, panels whose text could not be read: each failure cause is counted below, including the ones that are our instrument's limits rather than the site's.

Three qualifications on that accounting, none of them buried. **The consent step is
attempted but its outcome is not recorded per site**, so neither a reader nor an
author can audit which unreadable rows really died on a cookie wall — only the ones
that failed loudly enough to be bucketed appear as consent failures, on a scanner whose own
source calls cookie walls the primary blocker to reading a widget's first message.
**The bucketing itself carried a defect in this study's first sweep**: rows
were assigned by matching error-message wording, so panels that opened but whose text could
not be read under unexpected wording fell into “other”, and the drafted funnel
failed simple subtraction. The aggregation code was frozen under precommitment until the
second sweep completed and was corrected after it, on 10 August 2026; buckets now key on
the recorded interaction outcome, and the table below is generated by the corrected code.
And **each failure cause below carries an attribution** — instrument-side,
site-side, or mixed — because one cause this study first drafted was labelled as a
fact about the site when the code that emitted it reads it as a fact about us: rows where a
vendor SDK was fingerprinted but our launcher selector no longer matched are a stale recipe
of ours, not an absent widget. They are labelled that way below, and the interaction-stage
instrument-side subtotal is printed as its own row. Read that subtotal as a *floor*
on our own share of this headline, not the total: the blind audit below shows the largest
failure bucket is dominated by a different error of ours — admissions that were never
chat widgets — which interaction-stage tags cannot see, and which is why that bucket
is tagged mixed rather than site-side.

One known contamination is stated rather than hidden: our admission audit found that a
share of “confirmed” widgets are not chat widgets at all, and those sites land in
the “panel did not open” row — so the denominator here is
*detector-confirmed*, not verified widget presence. That is not a small caveat on a
large number. The interval printed with this headline is **sampling error only**,
and sampling error is not the dominant uncertainty here: the denominator contamination is
larger, runs in one direction, and is *not* inside that interval. Correcting for it
would move this figure down. The correction pre-registered for exactly this purpose —
a blind re-judging of the failure-bucket screenshots this study already holds — has
now been run against this sweep's rows, and its result is published here rather than left
as an open action item: the sweep’s own above-fold screenshots of all 361 “panel did not open” sites were re-judged blind (judges saw the image only — no site, vendor, or error metadata; 20 planted duplicates agreed 19/20 on the launcher question). 260 were decidable: 67.7% (95% CI 61.8–73.1) showed no chat launcher visible in the capture — accessibility widgets, carousel arrows, trust badges and contact links our detector had admitted as chat — and 32.3% (26.9–38.2) showed a real launcher whose panel nonetheless did not open. Two bounds travel with that rate: a launcher below the fold or rendered after the screenshot reads as absent, so the no-launcher share is biased upward by an unmeasured amount; and the 101 undecidable screenshots are dominated by consent layers covering the corners where launchers sit — an exclusion that is not neutral. Both are stated in the method log record beside the judgements. The headline above is still printed
*raw*, on the detector-confirmed denominator, exactly as pre-registered —
re-basing it on the audited denominator is an analysis change that goes through adversarial
review before it is published, not a hand adjustment — but the size and direction of
its largest known error term are now measured and on the page, not disclosed and
unresolved.

Whatever an automated visitor cannot reach, **this** automated check cannot
verify. That sentence is bounded to this instrument deliberately, because the unbounded
version — that no automated check could do better — is one this study falsified
in its own workshop. A single fix to how our scanner clicks launchers moved the readable
share on the same subsample of sites by around eight points — most of that movement was
in our code rather than in the web, though the log is explicit that this comparison ran
across an instrument change and cannot fully separate the two (method log §13a: 9
readable gained against 2 lost, 137 sites). That cross-code movement is also this study's
*only* noise estimate — there is no test-retest floor of any kind, as the
provenance table below states. Our own aggregation code refuses to publish a
per-vendor readable column for this reason, in its own words, because it “would rank
OUR recipes while reading as a ranking of vendors”. Aggregating that quantity across
vendors does not stop it being a property of our recipes, and the headline above is its
aggregate. Read it as the ceiling on what *this scanner* could see — an
indication of the difficulty an automated check faces, never a measurement of the limit of
automated checking.

| Stage | Sites | Share of confirmed widgets |
|---|---|---|
| Widget confirmed by the sweep | 794 | 100% |
| Chat panel opened | 259 | 33% (29–36) |
| First-interaction surface readable | 174 | 22% (19–25) |
| Unreadable: Panel did not open after the click — mixed | 361 | 45% (42–49) |
| Unreadable: Panel opened but its text was unreadable — instrument-side | 85 | 11% (9–13) |
| Unreadable: Launcher covered by another overlay — site-side | 54 | 7% (5–9) |
| Unreadable: Vendor recipe stale (SDK fingerprinted; our launcher selector did not match) — instrument-side | 42 | 5% (4–7) |
| Unreadable: Panel selector matched only the closed launcher — mixed | 34 | 4% (3–6) |
| Unreadable: Launcher found but not clickable — site-side | 25 | 3% (2–5) |
| Unreadable: Launcher covered by a consent/cookie layer — site-side | 14 | 2% (1–3) |
| Unreadable: Other — unattributed | 5 | 1% (0–1) |
| Unreadable subtotal attributable to the instrument at the interaction stage (rows tagged instrument-side above; mixed rows are not counted). A lower bound: it excludes the admission contamination the blind audit above measures inside the panel-did-not-open row | 127 | 16% (14–19) |

## By country

Country is the site's ccTLD. The rate shown is the share of confirmed widgets whose
first message was readable. Cells under 30 are reported as “n too small” per the
pre-registered minimum-cell rule — applied here to *confirmed widgets*, which is
the denominator of the published rate. The rule as pre-registered set that threshold on
cohort sites; this note is where that change is recorded, and the vendor table above applies
the same threshold.

**This table ranks our scanner at least as much as it ranks countries.**
The readable rate is far higher on sites where we fingerprinted a known vendor than on sites
admitted by the generic rule, for the mundane reason that a known vendor is what gives our
scanner a panel recipe to follow. That split is large: readable rate runs far lower on
no-vendor admits than on vendor-fingerprint admits. The per-country share of unknown / custom
widgets is not published in the table below, so a reader cannot check that correlation against
what is printed here — it is asserted from the underlying rows, not shown. The table is published because it was
pre-registered, and cutting a pre-registered table after seeing which way it fell would be
worse than publishing it with this warning attached — but read the ordering as a map
of our recipe coverage per country before reading it as anything about the countries. Rank
depth was tested as an alternative explanation and does not account for the spread.

| ccTLD | Confirmed widgets | Surface readable (95% CI) |
|---|---|---|
| .de | 202 | 11% (8–17) |
| .fr | 99 | 20% (13–29) |
| .it | 62 | 18% (10–29) |
| .pl | 53 | 38% (26–51) |
| .nl | 47 | 15% (7–28) |
| .se | 35 | 29% (16–45) |
| .cz | 33 | 30% (17–47) |
| all others (each n < 30) | 263 | 28% (23–33) |

## Disclosure at first interaction — on the readable subset only

**Read the scope before the number.** This table covers only the sites
whose first-interaction surface we could actually read — the readable share shown
above — and those sites are not a random sample: they skew toward simpler widgets
and weaker consent walls. It is reported as the state of the *readable* web, not
extrapolated to the rest, and that selection is exactly why the coverage number above is a
headline finding rather than a footnote.

**What this table is not: evidence about Article 50 as a duty.**
Article 50(1) obliges a provider to inform a person that they are interacting with an AI
system, and the rule pack's failing clause accordingly requires that the chat was
*confirmed* to be automated. Confirming automation means probing a chat for an
automated reply, and this study never does that to a site it does not own — the probe
is hard-disabled for every third-party site, as a safety rule. The consequence is structural
and belongs on the page rather than in a log: **no site in this study could have been
graded as failing Article 50(1), and none was.** Every disclosure finding here is
at most a hedged warning, which is what the rule pack itself says in its own words —
automation could not be confirmed by the probe, so the finding stays hedged rather than a
failure. Two further blind spots run the same way. The instrument cannot see
Article 50(1)'s carve-out for cases where it is obvious to a reasonably well-informed
person that they are dealing with an AI: disclosure is read from the first-interaction text
alone, so a widget whose launcher plainly reads “AI assistant” but whose greeting
does not repeat it is counted below as no disclosure detected — an error running
against the sites doing the most visible thing. And Article 50(1) addresses providers,
while the duties falling on the site that deploys a widget are 50(2) and 50(4). Nothing on
this page allocates a duty to anyone.

**How the three cells are drawn.** “Could not verify” here means
the wording was inconclusive on a surface we *did* read — a different thing from
the coverage figure above, which counts surfaces we could not read at all. The split between
“could not verify” and “no disclosure detected” turns partly on a
minimum character count for the first-interaction text — 120 characters — which
lives in the detector code rather than in the versioned rule pack, so a sealed report citing a
pack version does not pin it. “No disclosure detected” is also a union of
two different pack outcomes, one of them covering chats that present as human-staffed
— where, if humans genuinely staff the chat, no disclosure is required at all. This
page cannot tell you how that cell divides, because the aggregation query did not retain the
field separating them. And, as stated at the top: these three counts are raw detector output.
The disclosure detector's precision and recall were never measured, so unlike the prevalence
headline these cells carry no correction and no published error rate.

| Outcome on the readable first-interaction surface | Sites | Share of readable |
|---|---|---|
| Disclosure detected | 19 | 11% (7–16) |
| No disclosure detected | 73 | 42% (35–49) |
| Could not verify (wording inconclusive) | 82 | 47% (40–55) |

## What we could not measure — published, not buried

There is no before/after comparison on this page: the baseline sweep taken on 1 August was retired by our own decision on 2 August, so this study holds no pre-application measurement of these surfaces and publishes none. There is no test-retest noise floor either — the pre-registered same-day retest never ran on baseline day, and its registered substitute never ran on run-2 day; the only noise estimate we hold is a cross-code bound (≈4 points on widget presence, ≈8 on readable surfaces, 137 sites, measured Monday 27 July across an instrument change), and nothing on this page calls that a noise floor. Of 794 confirmed widgets, 259 opened a panel and 174 yielded readable first-interaction text; the disclosure counts cover those 174 alone. Of the 361 sites whose panel never opened, a blind re-judging of the captured screenshots found no chat launcher visible in the capture on two-thirds of the decidable cases — consistent with the admission false-positive classes our own audits measured, and counted in the coverage section rather than buried. The sweep egressed from the United States (measured on scan day: US, Miami PoP); a widget that geofences or staffs differently for EU visitors would look different from inside the EU, and we did not measure that. The scan ran on a Monday where the registered plan said Saturday — a deviation recorded as one, with its reason, in the method log; that log, like the rest of the bundle, is not yet published.

## Methodology

The sampling frame, the exclusion rules and the judgement protocol were pre-registered
before the data existed, and are versioned. Several pre-registered requirements were
nonetheless not met; those are named in the section above rather than dropped. The method log
itself is not currently published — see the note under the instrument table. In brief:
candidates are the highest-ranked pay-level domains in the Tranco list (XN23N (2026-07-27))
whose public suffix is an EU-27 ccTLD or .eu. That is a *rank-truncated slice* of that
list, not every such domain: the frame stops at a fixed rank cut, so sites and vendors
concentrated below the cut are outside this study entirely. Each candidate was rendered in a
real browser (widgets inject after JavaScript, so a static check would systematically miss
client-rendered vendors); sites whose robots.txt disallowed our scanner — or whose
robots.txt could not be reached — were excluded and counted. The chat panel is opened
read-only: the scanner never types into, sends, or otherwise writes to any site's chat.

The discovery funnel — the stage *before* the sweep table below, at which
the prevalence denominator is built — is: **16,000** candidates taken at
the rank cut (of 128,124 eligible pay-level domains), **13,441** rendered,
minus **708** that redirected off the frame and **448** duplicate
final sites, giving the **12,285**-site frame the prevalence headline is
computed over. Two of its exclusions are named here because they run against this study's
own interest, and both are discovery-stage counts — not the single-digit rows of the
sweep funnel below. Candidates that answered the discovery pass with a bot challenge are
excluded outright rather than counted as having no widget — so the sites most likely
to be running defended commercial deployments are absent from the denominator, not scored
zero inside it. And the 708 candidates that redirected off the frame were dropped after
rendering; as a group they carried a *higher* widget rate than the sites retained.
Both nudge the prevalence denominator the same way, and neither is corrected for. The
arithmetic effect is small. We state the counts because a funnel that discards a large
share of its candidates before the first published row should not report only the part
that survived.

| Stage | Sites |
|---|---|
| Submitted to the sweep | 1142 |
| Excluded: robots.txt disallows scanning | 4 |
| Excluded: robots.txt unreachable (treated as no) | 50 |
| Excluded: invalid / unresolvable | 0 |
| Capture blocked (bot protection / redirect) | 1 |
| Scan did not complete (cause not attributed; excluded) | 22 |
| Scanned to completion | 1065 |
| Captured outside pre-registered hour band (excluded) | 1 |
| Graded inside the pre-registered band | 1064 |
| Widget confirmed at sweep time | 794 |
| First-interaction surface readable | 174 |

### Instrument & provenance

**What you can and cannot check today.** The repository snapshot, the method
log and the underlying data are **not published**. The commit, cohort hash and
timestamp anchors above are therefore checkable only by someone who already holds the bundle;
no dataset download, tarball or DOI exists, and nothing on this page can be independently
recomputed from a public source today. Publishing the bundle is an open action item, not a
completed one, and the claims above should be read with that discount applied. Even with the
bundle in hand, the sweep aggregates are not recomputable by an outside reader: the
aggregation script reads from a private database that does not ship with the snapshot. One
practical note for whoever does get the bundle — verify the manifest against the commit
named above, not against a later tree: documentation inside the bundle has been edited since
the anchor was taken, which is expected and is described in the precommitment note.

## Limitations

EU-facing is a ccTLD proxy: candidates are the highest-ranked pay-level domains whose public suffix is an EU-27 ccTLD or .eu — a rank-truncated slice that misses .com EU businesses and everything below the rank cut.

The sweep ran from a United States vantage (measured on scan day: US, Miami PoP). Geofenced or EU-differentiated widgets may present differently from inside the EU; the baseline run’s egress was never recorded and never can be.

Homepage only, one page per site: a widget that lives on a support or checkout page is out of scope by design.

The prevalence correction’s precision and recall inputs are measured on judged samples with stated confidence intervals; the recall term dominates the corrected interval’s width.

Ground truth is a blind judge reading an above-fold real-browser screenshot — not what a site runs, and not what a human visitor exploring the page would find. A launcher below the fold or rendered late reads as absent.

“Unclear” judgements are excluded from every measured rate, and the exclusion is not neutral: unclears are dominated by consent walls that hide exactly the corners where launchers sit.

The disclosure lexicon covers 23 of the 24 official EU languages plus Catalan — Irish is absent — and a greeting in an uncovered language publishes as “no disclosure detected”, not as “could not assess”. The pre-registered language-fallback outcome was never implemented.

No human arm: this study cannot say how far an automated visitor’s view sits from a human visitor’s.

Single post-application snapshot: the retired 1 August baseline is not a publishable comparator, so there is no before/after delta, and there is no test-retest noise floor of any kind.

The automation probe is hard-disabled on third-party sites (it would write into a stranger’s support queue), so no site here could have been graded as failing Article 50(1); disclosure findings are hedged observations, never legal conclusions.

**Where does your own site stand?** The scanner that produced these numbers runs free, one URL at a time, evidence included — on the current code and the current rule pack, both of which move on, so a scan run today need not reproduce a figure on this page.

[Run the free scan →](/scan/)

Cite as: DisclosureProof, “The State of AI Disclosure
2026”. Scan dates: Monday 10 August 2026. The aggregate tables on this page are licensed
CC BY 4.0 — reuse with attribution and a link. That licence covers what is
printed here and nothing more: no underlying dataset, repository snapshot or per-site record
is published, and per-site records never will be.
Questions, corrections, or press: [hello@disclosureproof.com](mailto:hello@disclosureproof.com).
