# The EU AI Act training-data census, corrected: 13 of 21, not 2 of 21

> Source: <https://dev.to/pennyforgehq/the-eu-ai-act-training-data-census-corrected-13-of-21-not-2-of-21-41n9>
> Published: 2026-10-09 06:43:34+00:00

Nine days after [we counted "only 2 of 21"](https://dev.to/pennyforgehq/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for-414h) AI providers publishing the document the EU AI Act asks for, we went back and re-verified every row in the census. The corrected number: **13 of 21 providers now publish a dedicated, template-conformant "Public Summary of Training Content"** under Article 53(1)(d).

That's a big correction. Before we celebrate the providers or beat up on ourselves, it's worth being precise about what actually happened — because part of the delta is real drift, and part is us having looked too shallowly.

## 
  
  
  The document, briefly

The EU's template (C(2025) 8311 final) asks GPAI providers to publish a public summary with exactly three sections: (1) general information (provider, model IDs, EU market-placement date, per-modality data size brackets), (2) a list of data sources (public, licensed, crawled, user, synthetic), and (3) data processing aspects (TDM opt-outs, filtering, etc.). The obligation has applied since 2025-08-02; pre-existing models have until 2027-08-02.

In the issue #1 sweep we graded what a first-time visitor sees: the model card, the research page, the landing page. **That undercounted.** Most of the 10 providers we then promoted keep their Art. 53(1)(d) documents one click deeper — on a legal page, a trust center, or a transparency-report page that the card never links to.

## 
  
  
  The 10 promotions (all verified 2026-10-09)

Ten providers moved up to T3 (dedicated, template-structured summary):

- 
**OpenAI** — per-model "Public Summary of Training Content" PDFs (GPT-5.2 checked live: v1, 2026-07-30, OpenAI Ireland, placed 2025-12-11, >10T tokens; the document says it covers GPT-5.2 and all subsequent releases in the lifecycle). Six or more per-model PDFs, indexed at help.openai.com.
- 
**Google** — an 8-page "Public Summary of Training Content for General-Purpose AI Models — Gemini 3 model family" (v1, 2026-07-02, Google Ireland, placed Nov 2025, all four modalities), hosted on the AI Transparency Report's GCS bucket. A Gemma 4 report (2026-07-31) is listed too.
- 
**Anthropic** — a "Training Data Summaries" section on the Trust Center with**ten per-model documents** (Opus 4.7→5.5, Sonnet 5/5.5, Mythos/Fable lines, plus an AB 2013 summary). The Opus 5.5 document (v1, 2026-09-22) is template-conformant and names low-resource languages explicitly (Basque, Breton, Korean).
- 
**Meta** — an explicit Art. 53(1)(d) reporting program: "EU AI Act Transparency Reports published by Meta Platforms Ireland Limited pursuant to Article 53(1)(d) of Regulation (EU) 2024/1689", with per-model "Transparency Report" documents (Muse Glimmer v1, 2026-08-10 checked). Honest caveat: it's a thin 2-pager (a third-party tracker grades it C/D+).
- 
**Microsoft** — per-model "Data Summary" PDFs for the MAI series (MAI-Thinking-1 checked live: v1.0, 2026-08-12, Microsoft Ireland, EU rep contact listed, placed 2026-08-12).
- 
**ByteDance** — dedicated "Public Summary of Training Data Content" PDFs per model from seed.bytedance.com/transparency (Seed 2.0 Pro checked live: v1.0, 2026-07-31, EU authorised rep Mikros Ireland, placed March 2026; Seedream 5.0 Pro and Seedance 2.0/2.5 listed).
- 
**xAI** — "Public Summary of Training Content for Grok 4.5" (V1, 2026-07-08, xAI LLC, EU rep in Tallinn, placed 2026-07-14); Grok 4.6 (2026-08-12) also listed on x.ai/legal.
- 
**Tencent** — "Public Summary of Training Content for Hy3" (V2, 2026-08-20, OriGen Tech Singapore, EU rep Tencent International Service Europe B.V., Amsterdam, placed 2026-07-06). This closes the census's original CN-language blind spot.
- 
**Cohere** — "Public Summary of Training Content for Command A+" (v1, 2026-07-31, Cohere Germany GmbH, placed 2026-05-20, 48 languages, data to April 2026).
- 
**Aleph Alpha** — "Kolibri 1 — Sufficiently Detailed Summary" (v1.0, 2026-10-03): 8 pages, all three sections substantive, 8 named public datasets and 7 named synthetic-generator models.

Add the three that were already at T3 before the correction (DeepSeek, Mistral, NVIDIA) and you get 13 of 21.

## 
  
  
  The DeepSeek mystery, resolved

Last week we couldn't resolve DeepSeek's summary file: their CDN is an S3 bucket whose index 404s and whose filenames appear to rotate. This week the answer arrived from a third-party tracker's link list: the exact file `DeepSeek_Template_for_the_Public_Summary_of_Training_Content_for_GeneralPurpose_AI_model.pdf` **is live** on cdn.deepseek.com. It covers DeepSeek-V4 (incl. V4-Pro and V4-Flash), v1, last updated 2026-08-04, placed on the EU market 2026-04-24 — and it names an EU authorised representative: **Prighter Group**. Lesson: the bucket isn't unstable for the *current* filename, just for the old ones.

## 
  
  
  The remaining 8

- 
**T2+ (detailed reports/cards, no template summary):** Alibaba Qwen (Qwen3 tech report, 36T tokens; the Qwen3.8 flagship cards are the place to watch), Baidu ERNIE 5.0 (tech report + release post, 2.4T unified multimodal), IBM Granite 4.0, Moonshot Kimi K2 (15.5T tech report), Cerebras-GPT.
- 
**T2:** Amazon (Nova 2 Lite "AI Service Card" — one training-data paragraph).
- 
**T2-gated:** Stability AI (card inside a login-gated Hugging Face repo).
- 
**T1-2:** Zhipu GLM (research page only).

## 
  
  
  The gap moved. It's now about hosting, not existence.

The most interesting finding of the re-verification is that the compliance gap has *changed shape*. In October 2025 the question was "does anyone publish?"; now the question is "can you still find it in 90 days":

- 
**Cohere** hosts its summary behind a**signed S3 URL that expires after 7 days** . Our first-party link from August is already dead; the document survives via the stable docs page that links to it (and via third-party archives).
- 
**Tencent's** CDN returned an obscure HTTP 567 to EU fetchers all day; the document only resolved through a dated archive copy.
- 
**DeepSeek's** old filenames 404 even as the current one works — no redirects, no index.
- 
**Meta's** per-model report is 2 pages.

None of that disqualifies the document today. All of it will matter in 2027, when the pre-existing-models deadline hits and rightsholders start auditing.

## 
  
  
  The closest incumbent (and why we cross-reference it)

The AI Accountability Lab (aial.ie) runs "GPAI Training Transparency": 78 discovered summaries with live URLs, dated archives and A+–F grades. It's the closest thing to an incumbent of what we're building — and its dated archive copies served as our verification substitutes wherever first-party hosting was flaky (Tencent, Cohere, Anthropic, Meta). We treat it as a discovery cross-check, not ground truth: every census row now records in `url_source` exactly which route verified it (fetched-direct, or live-anchor-fetched + dated-archive-checked).

## 
  
  
  The machine-readable part

The census is now a small public repo with a diff tool: [github.com/m0nk111-qwen-agent/gpai-transparency-census](https://github.com/m0nk111-qwen-agent/gpai-transparency-census). The first real 24-hour diff (2026-10-08 → 2026-10-09) shows 17 of 21 rows changing, including the 10 promotions. That diff — not the absolute count — is the number we'll actually track going forward: **how many providers move per day**.

## 
  
  
  Honest caveats

- The "2 of 21" in issue #1 was the *card-level* truth at publish time (what a first-time sweep sees), stated as such in the repo's README; this post is the correction, not a contradiction.
- "Template-conformant" is a structural check (the three sections present and filled), not a quality grade — Meta's 2-pager passes structurally and is thin in practice. A quality dimension is the next iteration.
- Chinese providers were swept with a degraded search lane plus direct domain checks; Qwen3.8, ERNIE 5.0 and MiniMax H3 are the best next-checks if anyone wants to sharpen that corner.

*Pennyforge (one-person studio) · per-provider evidence as linked · AI-assisted research, studio-owned*
