{"slug": "download-a-fraction-of-data-from-huggingface-datasets", "title": "Download a fraction of data from HuggingFace Datasets", "summary": "A Hugging Face user asked how to download roughly 15GB from each language subsection of the bigcode/the-stack dataset instead of pulling entire subdirectories via load_dataset, and community respondents pointed to the allow_patterns filter in snapshot_download plus the datasets library's data_dir and data_files options. The original poster confirmed they resolved the issue by using HfApi to list repo files and their sizes, though respondents noted that approach does not let you specify a target size or file count directly. A later commenter, Deniz, asked whether the same method works for 20-30GB zipped files that they do not want to download whole.", "body_md": "hey all, I want to download about 15GBs of data from each subsection/language of the [stack dataset](https://huggingface.co/datasets/bigcode/the-stack). Is there any way to do this?\n\nI can download the entire subsection using following code snippet, but i want a fraction of the data(decided by a percentage or number of files).\n\n``` python\nfrom  datasets  import  load_dataset\nds = load_dataset(\"bigcode/the-stack\", data_dir=\"data/c\", split=\"train\")\n```\n\nThanks in advance\n\n \n \nI think it should be possible if you specify the directory name in the **allow_patterns** filter for **snapshot_download**.\n\nI’m not very familiar with it, but it seems that the **datasets** library itself also has functions to extract a part of a dataset.\n\n \n \nyeah, it has data_dir and data_files, but what I was looking for was be able to download 10 GBs or 10 files of data , what the above will do is download the complete sub directory.\n\n \n \nI think this is close, but technically speaking, it’s not the number of files, and you can’t specify the size.\n\nYou can use HfApi to get a list of the files in a repo and see their sizes, but that’s almost like doing it manually…\n\n \n \nyes, thats what I did, thanks for the help btw.\n\n \n \nHello,\n\nIs the method applicable for zipped filestoo?\n\nI have some zippred files aroudn 20 30 GB but I dont want to download them as a whole\n\nBest,\n\nDeniz", "url": "https://wpnews.pro/news/download-a-fraction-of-data-from-huggingface-datasets", "canonical_source": "https://discuss.huggingface.co/t/download-a-fraction-of-data-from-huggingface-datasets/126236#post_7", "published_at": "2026-10-03 22:33:01+00:00", "updated_at": "2026-10-03 22:38:30.382931+00:00", "lang": "en", "topics": ["ai-tools", "developer-tools"], "entities": ["Hugging Face", "bigcode/the-stack", "datasets library", "HfApi", "snapshot_download", "Deniz"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/download-a-fraction-of-data-from-huggingface-datasets", "markdown": "https://wpnews.pro/news/download-a-fraction-of-data-from-huggingface-datasets.md", "text": "https://wpnews.pro/news/download-a-fraction-of-data-from-huggingface-datasets.txt", "jsonld": "https://wpnews.pro/news/download-a-fraction-of-data-from-huggingface-datasets.jsonld"}}