A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on the Hub, showing that 65% declare no license — so teams can avoid licensing traps (and missing-tag repos) before… A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on Hugging Face shows that 65% declare no license, potentially disqualifying them from commercial pipelines under the EU AI Act and enterprise procurement policies. The matrix, published by hardik90, includes risk buckets and guidance for every repository, available as CSV and Parquet under CC0. I scanned every Indic-language dataset on the Hub hindi, tamil, bengali, telugu, marathi, malayalam, kannada, urdu, +9 more and published the full results. Full matrix with risk buckets + guidance for every repo CSV + Parquet, CC0 : hardik90/indic-dataset-license-matrix https://huggingface.co/datasets/hardik90/indic-dataset-license-matrix Also in that repo: a sample audit report https://huggingface.co/datasets/hardik90/indic-dataset-license-matrix/blob/main/sample-audit-report.md with a deep-dive on 3 flagship corpora, plus an FAQ covering commercial use, CC-BY-NC, and how to verify any dataset’s license. Under the EU AI Act GPAI documentation duties and standard enterprise procurement policies, a missing license tag effectively disqualifies a dataset from commercial pipelines — regardless of the author’s original intent. If you train or ship on Indic data , filter the matrix by risk bucket and build your watchlist before your next fine-tune. Maintainers : I’d love to see your license tags + provenance notes — the matrix is refreshed monthly, and I’m happy to help draft the card text. The matrix is factual, API-derived metadata released under CC0. Not legal advice — verify anything you rely on.