{"slug": "apertus-misrepresents-open-status-by-conflating-technical-accessibility-with", "title": "Apertus Misrepresents \"Open\" Status by Conflating Technical Accessibility With Legal Freedom", "summary": "Apertus, a Swiss AI initiative, misrepresents its 'open' status by conflating technical accessibility with legal freedom, according to a technical audit. The audit found that while Apertus's model weights and code are open-source, its training data is largely copyrighted web text scraped without explicit licenses, making its 'open data' claim misleading and legally precarious under Swiss and EU copyright law. The report criticizes Apertus for using robots.txt compliance as a proxy for legal permission and for highlighting GDPR compliance while ignoring copyright norms, undermining its 'European Values' positioning.", "body_md": "[Apertus: \"Accessible\" ≠ \"Open\"](#Apertus:--22-Accessible-22---e2--89--a0---22-Open-22-)\n\n# Apertus: “Accessible” ≠ “Open”\n\n## A Technical Audit of the Swiss AI Initiative’s Data Model\n\nExecutive Summary:Apertus markets itself as a “fully open” model built on “European Values.” However, its definition of “open data” relies ontechnical accessibility(robots.txt compliance) rather thanlegal permissiveness(copyright clearance). This creates a fundamental contradiction: a model claimed to be “open” and “transparent” is built on a corpus of copyrighted works whose specific origins are opaque, making true “studyability” and “redistribution” legally ambiguous.\n\n## 1. The Distinction: Open Source Weights vs. Open Data\n\nApertus correctly states that its **architecture** and **model\nweights** are open-source. However, it conflates this with the\nopenness of its **training corpus**.\n\n| Feature | Claimed Status | Technical Reality |\n|---|---|---|\nWeights |\nOpen Source | ✅ True. Licensed permissively. |\nCode |\nOpen Source | ✅ True. Training code is available. |\nTraining Data |\n“Publicly Available” | ⚠️ Misleading. “Publicly Available” means “visible to crawlers,” not “free of rights.” |\nCopyright Status |\nImplicitly Clear | ❌ Opaque. Most web data is copyrighted. Apertus uses it without explicit license. |\n\n### Why “Publicly Available” is Not “Open”\n\nIn legal terms, **Publicly Available** means *anyone can see\nit*.\n\n**Open** (in the context of intellectual property) usually implies\n*anyone can use it* (Public Domain or Creative Commons) or “free as in\nfreedom”.\n\nApertus’s training corpus consists largely of copyrighted web text. By\nlabeling this “open data,” they create a false equivalence. Users\nassume that because the data was “on the web,” it is free to use. It\nis not. It is **scraped**.\n\n## 2. The “European Values” Contradiction\n\nApertus positions itself as a sovereign, European alternative. This claim rests on two pillars: **Data Sovereignty** and **Legal Compliance**.\n\n### 🛑 GDPR vs. Copyright: A Category Error\n\nApertus highlights its **GDPR compliance** extensively. This is true but incomplete.\n\n**GDPR** protects**Personal Data**(identifiable human beings).** Copyright**protects** Creative Works**(articles, code, literature).\n\nApertus filters **PII** (Personally Identifiable Information) from its\ntraining data, which satisfies GDPR.\n\nHowever, it does **not** filter **Copyrighted Material**. Most of the\ninternet is copyrighted. By scraping it, Apertus is operating in a\nlegal gray area common to all LLMs, but they frame this as “ethical\nstandards.”\n\n**The Critique:** To claim “European Values” while relying on a data\nset that largely ignores copyright norms (the default state of the\nweb) is a selective reading of European IP law. The EU AI Act\naddresses *risk*, not *copyright ownership*. Apertus conflates the\ntwo.\n\n### 🇨🇭 Swiss Copyright Law\n\nSwiss copyright law protects literary and artistic works. It does not automatically grant permission for machine learning ingestion. By using “publicly available” data, Apertus is relying on an implicit, untested “fair use” or “text and data mining” exception argument, rather than explicit permission. This makes the “Open” label legally precarious.\n\n## 3. The `robots.txt`\n\nLoophole: Technical vs. Legal\n\nApertus defines “openly available data” as:\n\n“Data which is publicly available… filtered to respect machine-readable opt-out requests… even retroactively.”\n\n### The Technical Definition\n\n**“Open” = “Not Blocked by Robots.txt”**\n\nIf a website does not explicitly block bots via `robots.txt`\n\n, Apertus\nconsiders the data “open.”\n\n### The Legal Flaw\n\nThis equates **Technical Access** with **Legal Permission**.\n\n**Default State:** The default legal status of web content is**Copyrighted**, not Public Domain.** The “Opt-Out” Burden:**Apertus shifts the burden of permission to the creator. You must*actively block*them to retain your rights. In Free Software terms, this is akin to saying software is “open” unless you write a license file.**Retroactivity:** Apertus claims to respect opt-outs “retroactively.” This is a technical feat of filtering, but it does not change the fact that the*initial*ingestion was based on a technical heuristic, not a legal license.\n\n### Why This Breaks “Traceability”\n\nApertus claims the model is “traceable.” But traceability requires a **manifest**.\n\n- If I want to trace the source of a specific output, I need to know which URL generated it.\n- Apertus provides “recipes” (datasets used), not a precise\n**URL-to-Weight** mapping for the entire web scrape. - Therefore, “traceability” is limited to dataset-level attribution, not content-level provenance. This makes it difficult for users to verify if a specific copyrighted work was included or how.\n\n## 4. Request for Apertus v2: Transparency, Not Just Openness\n\nWe do not ask for perfection. We ask for **precision** in their claims.\n\n### ✅ What Apertus Should Publish in v2:\n\n**Data Composition Manifest:**- A breakdown of the training corpus: % Public Domain, % CC-Licensed, % Copyrighted/Scraped.\n- A list of the top 100 largest sources by token count, not just “the web.”\n\n**Clear Copyright Stance:**- Acknowledge that “Openly Available” =\n**Technically Accessible**. - Acknowledge that the model is built on\n**Copyrighted Data** without explicit license, distinguishing it from models trained on fully licensed datasets (like Common Crawl with verified licenses).\n\n- Acknowledge that “Openly Available” =\n**Source Code for the Filter:**- Publish the exact code that parses\n`robots.txt`\n\nto prove the “retroactive” filtering logic is robust and not just a heuristic.\n\n- Publish the exact code that parses\n\n### 🚩 Conclusion\n\nApertus is a **powerful, open-weight model** built on **accessible data**. It is not a “free” model in the legal sense because the underlying data rights are unresolved. It is a **scraped model** wearing an “open” label.\n\nOpen Source Weights, Closed Source Truth.\n\n[Download Apertus Weights](#)\n\n*Technically open. Legally ambiguous.*\n\n**If the download link doesn’t work, that is because we do not recommend non-free LLMs.**", "url": "https://wpnews.pro/news/apertus-misrepresents-open-status-by-conflating-technical-accessibility-with", "canonical_source": "https://gnu.support/large-language-models-llm/Apertus-Misrepresents-Open-Status-by-Conflating-Technical-Accessibility-With-Legal-Freedom-128528.html", "published_at": "2026-08-14 08:32:10+00:00", "updated_at": "2026-08-21 01:12:46.606200+00:00", "lang": "en", "topics": ["ai-ethics", "ai-policy", "artificial-intelligence"], "entities": ["Apertus", "Swiss AI Initiative", "GDPR", "EU AI Act"], "alternates": {"html": "https://wpnews.pro/news/apertus-misrepresents-open-status-by-conflating-technical-accessibility-with", "markdown": "https://wpnews.pro/news/apertus-misrepresents-open-status-by-conflating-technical-accessibility-with.md", "text": "https://wpnews.pro/news/apertus-misrepresents-open-status-by-conflating-technical-accessibility-with.txt", "jsonld": "https://wpnews.pro/news/apertus-misrepresents-open-status-by-conflating-technical-accessibility-with.jsonld"}}