# Only 2 of 21 AI Providers Publish the Document the EU AI Act Asks For

> Source: <https://dev.to/pennyforgehq/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for-414h>
> Published: 2026-10-08 19:09:27+00:00

The EU AI Act has been quietly asking AI model providers to do something specific since 2 August 2025: publish a "sufficiently detailed public summary of the content used for the training of the model", according to a template the Commission published. One year in, we went and looked at 21 of the largest GPAI providers. Two publish the actual document. Everyone else publishes *something else that is close*.

This is the first issue of a small, dated, verifiable conformance census. The table below is the deliverable; the method and caveats are below it, because we think both matter as much.

Article 53(1)(d) of Regulation (EU) 2024/1689 (the AI Act) obliges **all** providers of general-purpose AI models — including open-source and free ones, for this specific obligation — to make a public summary of training content available, "according to a template provided by the AI Office". The Commission published that template on 5 December 2025 (C(2025) 8311 final): an Explanatory Notice plus a fill-in form with three sections — **1. General information** (provider, EU authorised representative, versioned model name, EU market-placement date, modality and per-modality data-size brackets, types of content), **2. List of data sources**, **3. Data processing aspects**. The Commission calls it "a common minimal baseline" — more detail is fine, less is the problem.

The timeline: the obligation applies as of **2 August 2025**. Models placed on the EU market before that date have until **2 August 2027** to comply, and the AI Office's supervision and enforcement of GPAI rules starts **2 August 2026** — two months ago.

We checked 21 major GPAI providers on **2026-10-08**, reading the providers' own documents (not blogs about them). Two publish a dedicated, template-shaped "Public Summary of Training Content":

**DeepSeek.** Per-model summaries at cdn.deepseek.com/policies/ — e.g. the V3.2 summary (v1, updated 2026-09-03, 5 pages) is essentially the template filled in: EU authorised representative, versioned model ID, EU market-placement date (2025-08-21), per-modality size brackets, a source list (Common Crawl, Stack Exchange, licensed data, no user data, synthetic data), and a text-and-data-mining reservation.

**Mistral.** A legal center covering 42 models with per-model lifecycle and technical documentation, plus a dedicated "Public Summary of Training Content" for Large 3.

The remaining 19 publish model cards, technical reports, or system cards that describe training data in a paragraph, a table, or a citation to another document — detailed in some cases (Qwen3's 36-token composition, Aleph Alpha's 20T breakdown, Kimi K2's report), but none of them is a dedicated summary structured against the template. The table:

| Provider | Model(s) checked (10-08) | Where training content is described | Tier | 
|---|---|---|---|
| DeepSeek | V3.2, V4 | Dedicated per-model "Public Summary of Training Content" (Policies) | T3 | 
| Mistral | Large 3 (+42 models in legal center) | Dedicated Large 3 summary + per-model legal docs | T3 | 
| Anthropic | Claude Sonnet 4.5 | 149-page system card, §1.1.1 (two paragraphs: data mix, robots.txt, opt-in users) | T2 | 
| OpenAI | GPT-6.1 Sol | System card addendum (2026-09-29), §2: one sentence — "same types of data and training as GPT-6 Astra" | T2 | 
|  | Gemini 3 Pro | Model card PDF (one paragraph on training data) | T2 | 
| Meta | Llama 4 | MODEL_CARD.md (two lines + ~40T / ~22T figures) | T2 | 
| xAI (SpaceXAI) | Grok 4.6 | Model card PDF (2026-08-12), one paragraph | T2 | 
| Cohere | Command A+ | Docs page (three generic lines) | T2 | 
| Amazon | Nova 2 Lite | Responsible-AI card (one paragraph) | T2 | 
| Microsoft | Phi-4 | Model card (4 source categories, 9.8T); MAI series has no card | T2+ | 
| IBM | Granite 4.0 | Card (SFT sources) + tech report (pretrain mixture) | T2+ | 
| Alibaba (Qwen) | Qwen3 | arXiv 2505.09388 (36T + composition) | T2+ | 
| Baidu | ERNIE 5.0 | arXiv 2602.04705 | T2+ | 
| NVIDIA | Nemotron 3 Ultra | Technical report PDF | T2+ | 
| Moonshot (Kimi) | K2 | Tech report (moonshotai.github.io, 15.5T) | T2+ | 
| Cerebras | Cerebras-GPT | Paper + model card | T2+ | 
| Aleph Alpha | Kolibri-1 | Tech report (20T, language split; CoP signatory) — the strongest model card we saw | T2+ | 
| Stability AI | SD 3.5 | Card inside gated Hugging Face repo | T2 (gated) | 
| ByteDance | Doubao | Seed research pages only (API-first) | T1–2 | 
| Zhipu | GLM-4.6 | Card + GLM-4.5 report | T1–2 | 
| Tencent | Hunyuan | Unverified (blog-level only) | ? | 

**T3** = dedicated, template-conformant summary · **T2+** = detailed card/report, no dedicated summary · **T2** = brief section · **T1–2** = thin/indirect · **?** = not verifiable from public sources as of 10-08.

A watchdog post from December 2025 (Open Future, "5+3+3=0 transparency") quoted the US majors' model-card text and found it "does not remotely resemble" the template. We re-checked against primary documents on 10-08:

That is the pattern: training-data description is a sidebar in model cards everywhere except at the two providers that treated the template as the deliverable.

Because it changes. Mistral's summary was a section inside technical docs in December 2025 and a dedicated document by now. OpenAI's flagship card moved from GPT-5 to GPT-6.1 in the same period. A one-off article rots; a dated, versioned census with per-provider evidence links doesn't. That is the whole product hypothesis behind this: the value is not "who complies" (we can re-run it) but the **diff** — who published, who updated, who moved, against which template version, when.

Issue #2 will add machine-readable per-provider rows (version, date, URL, template section coverage) and a first diff check against this table. Corrections welcome — point us at a document we missed and we'll re-grade the row in the next issue.

*Pennyforge (one-person studio) · full evidence links on request · AI-assisted research, studio-owned*
