I scanned every Indic-language dataset on the Hub (hindi, tamil, bengali, telugu, marathi, malayalam, kannada, urdu, +9 more) and published the full results.
Full matrix with risk buckets + guidance for every repo (CSV + Parquet, CC0):
hardik90/indic-dataset-license-matrix Also in that repo: a sample audit report with a deep-dive on 3 flagship corpora, plus an FAQ covering commercial use, CC-BY-NC, and how to verify any dataset’s license.
Under the EU AI Act GPAI documentation duties and standard enterprise procurement policies, a missing license tag effectively disqualifies a dataset from commercial pipelines — regardless of the author’s original intent.
If you train or ship on Indic data, filter the matrix by risk_bucket and build your watchlist before your next fine-tune.
Maintainers: I’d love to see your license tags + provenance notes — the matrix is refreshed monthly, and I’m happy to help draft the card text.
The matrix is factual, API-derived metadata released under CC0. Not legal advice — verify anything you rely on.