# A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on the Hub, showing that 65% declare no license — so teams can avoid licensing traps (and missing-tag repos) before…

> Source: <https://discuss.huggingface.co/t/a-free-monthly-refreshed-compliance-matrix-of-all-4-893-indic-language-datasets-on-the-hub-showing-that-65-declare-no-license-so-teams-can-avoid-licensing-traps-and-missing-tag-repos-before-they-train-on-them/178373#post_1>
> Published: 2026-08-01 17:15:35+00:00

I scanned every Indic-language dataset on the Hub (hindi, tamil, bengali, telugu, marathi, malayalam, kannada, urdu, +9 more) and published the full results.

Full matrix with risk buckets + guidance for every repo (CSV + Parquet, CC0):

[hardik90/indic-dataset-license-matrix](https://huggingface.co/datasets/hardik90/indic-dataset-license-matrix)

Also in that repo: a [sample audit report](https://huggingface.co/datasets/hardik90/indic-dataset-license-matrix/blob/main/sample-audit-report.md) with a deep-dive on 3 flagship corpora, plus an FAQ covering commercial use, CC-BY-NC, and how to verify any dataset’s license.

Under the EU AI Act GPAI documentation duties and standard enterprise procurement policies, a missing license tag effectively disqualifies a dataset from commercial pipelines — regardless of the author’s original intent.

**If you train or ship on Indic data**, filter the matrix by risk_bucket and build your watchlist before your next fine-tune.

**Maintainers**: I’d love to see your license tags + provenance notes — the matrix is refreshed monthly, and I’m happy to help draft the card text.

*The matrix is factual, API-derived metadata released under CC0. Not legal advice — verify anything you rely on.*
