The EU AI Act has been quietly asking AI model providers to do something specific since 2 August 2025: publish a "sufficiently detailed public summary of the content used for the training of the model", according to a template the Commission published. One year in, we went and looked at 21 of the largest GPAI providers. Two publish the actual document. Everyone else publishes something else that is close.
This is the first issue of a small, dated, verifiable conformance census. The table below is the deliverable; the method and caveats are below it, because we think both matter as much.
Article 53(1)(d) of Regulation (EU) 2024/1689 (the AI Act) obliges all providers of general-purpose AI models — including open-source and free ones, for this specific obligation — to make a public summary of training content available, "according to a template provided by the AI Office". The Commission published that template on 5 December 2025 (C(2025) 8311 final): an Explanatory Notice plus a fill-in form with three sections — 1. General information (provider, EU authorised representative, versioned model name, EU market-placement date, modality and per-modality data-size brackets, types of content), 2. List of data sources, 3. Data processing aspects. The Commission calls it "a common minimal baseline" — more detail is fine, less is the problem.
The timeline: the obligation applies as of 2 August 2025. Models placed on the EU market before that date have until 2 August 2027 to comply, and the AI Office's supervision and enforcement of GPAI rules starts 2 August 2026 — two months ago.
We checked 21 major GPAI providers on 2026-10-08, reading the providers' own documents (not blogs about them). Two publish a dedicated, template-shaped "Public Summary of Training Content":
DeepSeek. Per-model summaries at cdn.deepseek.com/policies/ — e.g. the V3.2 summary (v1, updated 2026-09-03, 5 pages) is essentially the template filled in: EU authorised representative, versioned model ID, EU market-placement date (2025-08-21), per-modality size brackets, a source list (Common Crawl, Stack Exchange, licensed data, no user data, synthetic data), and a text-and-data-mining reservation.
Mistral. A legal center covering 42 models with per-model lifecycle and technical documentation, plus a dedicated "Public Summary of Training Content" for Large 3.
The remaining 19 publish model cards, technical reports, or system cards that describe training data in a paragraph, a table, or a citation to another document — detailed in some cases (Qwen3's 36-token composition, Aleph Alpha's 20T breakdown, Kimi K2's report), but none of them is a dedicated summary structured against the template. The table:
| Provider | Model(s) checked (10-08) | Where training content is described | Tier |
|---|---|---|---|
| DeepSeek | V3.2, V4 | Dedicated per-model "Public Summary of Training Content" (Policies) | T3 | | Mistral | Large 3 (+42 models in legal center) | Dedicated Large 3 summary + per-model legal docs | T3 | | Anthropic | Claude Sonnet 4.5 | 149-page system card, §1.1.1 (two paragraphs: data mix, robots.txt, opt-in users) | T2 | | OpenAI | GPT-6.1 Sol | System card addendum (2026-09-29), §2: one sentence — "same types of data and training as GPT-6 Astra" | T2 | | | Gemini 3 Pro | Model card PDF (one paragraph on training data) | T2 | | Meta | Llama 4 | MODEL_CARD.md (two lines + ~40T / ~22T figures) | T2 | | xAI (SpaceXAI) | Grok 4.6 | Model card PDF (2026-08-12), one paragraph | T2 | | Cohere | Command A+ | Docs page (three generic lines) | T2 | | Amazon | Nova 2 Lite | Responsible-AI card (one paragraph) | T2 | | Microsoft | Phi-4 | Model card (4 source categories, 9.8T); MAI series has no card | T2+ | | IBM | Granite 4.0 | Card (SFT sources) + tech report (pretrain mixture) | T2+ | | Alibaba (Qwen) | Qwen3 | arXiv 2505.09388 (36T + composition) | T2+ | | Baidu | ERNIE 5.0 | arXiv 2602.04705 | T2+ | | NVIDIA | Nemotron 3 Ultra | Technical report PDF | T2+ | | Moonshot (Kimi) | K2 | Tech report (moonshotai.github.io, 15.5T) | T2+ | | Cerebras | Cerebras-GPT | Paper + model card | T2+ | | Aleph Alpha | Kolibri-1 | Tech report (20T, language split; CoP signatory) — the strongest model card we saw | T2+ | | Stability AI | SD 3.5 | Card inside gated Hugging Face repo | T2 (gated) | | ByteDance | Doubao | Seed research pages only (API-first) | T1–2 | | Zhipu | GLM-4.6 | Card + GLM-4.5 report | T1–2 |
| Tencent | Hunyuan | Unverified (blog-level only) | ? | T3 = dedicated, template-conformant summary · T2+ = detailed card/report, no dedicated summary · T2 = brief section · T1–2 = thin/indirect · ? = not verifiable from public sources as of 10-08.
A watchdog post from December 2025 (Open Future, "5+3+3=0 transparency") quoted the US majors' model-card text and found it "does not remotely resemble" the template. We re-checked against primary documents on 10-08:
That is the pattern: training-data description is a sidebar in model cards everywhere except at the two providers that treated the template as the deliverable.
Because it changes. Mistral's summary was a section inside technical docs in December 2025 and a dedicated document by now. OpenAI's flagship card moved from GPT-5 to GPT-6.1 in the same period. A one-off article rots; a dated, versioned census with per-provider evidence links doesn't. That is the whole product hypothesis behind this: the value is not "who complies" (we can re-run it) but the diff — who published, who updated, who moved, against which template version, when.
Issue #2 will add machine-readable per-provider rows (version, date, URL, template section coverage) and a first diff check against this table. Corrections welcome — point us at a document we missed and we'll re-grade the row in the next issue.
Pennyforge (one-person studio) · full evidence links on request · AI-assisted research, studio-owned