{"slug": "only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for", "title": "Only 2 of 21 AI Providers Publish the Document the EU AI Act Asks For", "summary": "A conformance census of 21 major general-purpose AI providers found that only DeepSeek and Mistral publish a dedicated \"Public Summary of Training Content\" structured against the template the European Commission published on 5 December 2025 under Article 53(1)(d) of the EU AI Act. The remaining 19 providers — including Anthropic, OpenAI, Google, Meta, xAI, Cohere, Amazon, Microsoft, IBM, Alibaba and Baidu — describe training data only in model cards, system cards or technical reports, which the census classifies as non-conforming substitutes. The obligation has applied since 2 August 2025, with the AI Office's supervision and enforcement of GPAI rules beginning 2 August 2026.", "body_md": "The EU AI Act has been quietly asking AI model providers to do something specific since 2 August 2025: publish a \"sufficiently detailed public summary of the content used for the training of the model\", according to a template the Commission published. One year in, we went and looked at 21 of the largest GPAI providers. Two publish the actual document. Everyone else publishes *something else that is close*.\n\nThis is the first issue of a small, dated, verifiable conformance census. The table below is the deliverable; the method and caveats are below it, because we think both matter as much.\n\nArticle 53(1)(d) of Regulation (EU) 2024/1689 (the AI Act) obliges **all** providers of general-purpose AI models — including open-source and free ones, for this specific obligation — to make a public summary of training content available, \"according to a template provided by the AI Office\". The Commission published that template on 5 December 2025 (C(2025) 8311 final): an Explanatory Notice plus a fill-in form with three sections — **1. General information** (provider, EU authorised representative, versioned model name, EU market-placement date, modality and per-modality data-size brackets, types of content), **2. List of data sources**, **3. Data processing aspects**. The Commission calls it \"a common minimal baseline\" — more detail is fine, less is the problem.\n\nThe timeline: the obligation applies as of **2 August 2025**. Models placed on the EU market before that date have until **2 August 2027** to comply, and the AI Office's supervision and enforcement of GPAI rules starts **2 August 2026** — two months ago.\n\nWe checked 21 major GPAI providers on **2026-10-08**, reading the providers' own documents (not blogs about them). Two publish a dedicated, template-shaped \"Public Summary of Training Content\":\n\n**DeepSeek.** Per-model summaries at cdn.deepseek.com/policies/ — e.g. the V3.2 summary (v1, updated 2026-09-03, 5 pages) is essentially the template filled in: EU authorised representative, versioned model ID, EU market-placement date (2025-08-21), per-modality size brackets, a source list (Common Crawl, Stack Exchange, licensed data, no user data, synthetic data), and a text-and-data-mining reservation.\n\n**Mistral.** A legal center covering 42 models with per-model lifecycle and technical documentation, plus a dedicated \"Public Summary of Training Content\" for Large 3.\n\nThe remaining 19 publish model cards, technical reports, or system cards that describe training data in a paragraph, a table, or a citation to another document — detailed in some cases (Qwen3's 36-token composition, Aleph Alpha's 20T breakdown, Kimi K2's report), but none of them is a dedicated summary structured against the template. The table:\n\n| Provider | Model(s) checked (10-08) | Where training content is described | Tier | \n|---|---|---|---|\n| DeepSeek | V3.2, V4 | Dedicated per-model \"Public Summary of Training Content\" (Policies) | T3 | \n| Mistral | Large 3 (+42 models in legal center) | Dedicated Large 3 summary + per-model legal docs | T3 | \n| Anthropic | Claude Sonnet 4.5 | 149-page system card, §1.1.1 (two paragraphs: data mix, robots.txt, opt-in users) | T2 | \n| OpenAI | GPT-6.1 Sol | System card addendum (2026-09-29), §2: one sentence — \"same types of data and training as GPT-6 Astra\" | T2 | \n|  | Gemini 3 Pro | Model card PDF (one paragraph on training data) | T2 | \n| Meta | Llama 4 | MODEL_CARD.md (two lines + ~40T / ~22T figures) | T2 | \n| xAI (SpaceXAI) | Grok 4.6 | Model card PDF (2026-08-12), one paragraph | T2 | \n| Cohere | Command A+ | Docs page (three generic lines) | T2 | \n| Amazon | Nova 2 Lite | Responsible-AI card (one paragraph) | T2 | \n| Microsoft | Phi-4 | Model card (4 source categories, 9.8T); MAI series has no card | T2+ | \n| IBM | Granite 4.0 | Card (SFT sources) + tech report (pretrain mixture) | T2+ | \n| Alibaba (Qwen) | Qwen3 | arXiv 2505.09388 (36T + composition) | T2+ | \n| Baidu | ERNIE 5.0 | arXiv 2602.04705 | T2+ | \n| NVIDIA | Nemotron 3 Ultra | Technical report PDF | T2+ | \n| Moonshot (Kimi) | K2 | Tech report (moonshotai.github.io, 15.5T) | T2+ | \n| Cerebras | Cerebras-GPT | Paper + model card | T2+ | \n| Aleph Alpha | Kolibri-1 | Tech report (20T, language split; CoP signatory) — the strongest model card we saw | T2+ | \n| Stability AI | SD 3.5 | Card inside gated Hugging Face repo | T2 (gated) | \n| ByteDance | Doubao | Seed research pages only (API-first) | T1–2 | \n| Zhipu | GLM-4.6 | Card + GLM-4.5 report | T1–2 | \n| Tencent | Hunyuan | Unverified (blog-level only) | ? | \n\n**T3** = dedicated, template-conformant summary · **T2+** = detailed card/report, no dedicated summary · **T2** = brief section · **T1–2** = thin/indirect · **?** = not verifiable from public sources as of 10-08.\n\nA watchdog post from December 2025 (Open Future, \"5+3+3=0 transparency\") quoted the US majors' model-card text and found it \"does not remotely resemble\" the template. We re-checked against primary documents on 10-08:\n\nThat is the pattern: training-data description is a sidebar in model cards everywhere except at the two providers that treated the template as the deliverable.\n\nBecause it changes. Mistral's summary was a section inside technical docs in December 2025 and a dedicated document by now. OpenAI's flagship card moved from GPT-5 to GPT-6.1 in the same period. A one-off article rots; a dated, versioned census with per-provider evidence links doesn't. That is the whole product hypothesis behind this: the value is not \"who complies\" (we can re-run it) but the **diff** — who published, who updated, who moved, against which template version, when.\n\nIssue #2 will add machine-readable per-provider rows (version, date, URL, template section coverage) and a first diff check against this table. Corrections welcome — point us at a document we missed and we'll re-grade the row in the next issue.\n\n*Pennyforge (one-person studio) · full evidence links on request · AI-assisted research, studio-owned*", "url": "https://wpnews.pro/news/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for", "canonical_source": "https://dev.to/pennyforgehq/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for-414h", "published_at": "2026-10-08 19:09:27+00:00", "updated_at": "2026-10-08 19:20:13.230146+00:00", "lang": "en", "topics": ["ai-policy", "large-language-models", "artificial-intelligence", "ai-safety"], "entities": ["European Commission", "EU AI Act", "DeepSeek", "Mistral", "Anthropic", "OpenAI", "Google", "Meta"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for", "markdown": "https://wpnews.pro/news/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for.md", "text": "https://wpnews.pro/news/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for.txt", "jsonld": "https://wpnews.pro/news/only-2-of-21-ai-providers-publish-the-document-the-eu-ai-act-asks-for.jsonld"}}