Download a fraction of data from HuggingFace Datasets A Hugging Face user asked how to download roughly 15GB from each language subsection of the bigcode/the-stack dataset instead of pulling entire subdirectories via load_dataset, and community respondents pointed to the allow_patterns filter in snapshot_download plus the datasets library's data_dir and data_files options. The original poster confirmed they resolved the issue by using HfApi to list repo files and their sizes, though respondents noted that approach does not let you specify a target size or file count directly. A later commenter, Deniz, asked whether the same method works for 20-30GB zipped files that they do not want to download whole. hey all, I want to download about 15GBs of data from each subsection/language of the stack dataset https://huggingface.co/datasets/bigcode/the-stack . Is there any way to do this? I can download the entire subsection using following code snippet, but i want a fraction of the data decided by a percentage or number of files . python from datasets import load dataset ds = load dataset "bigcode/the-stack", data dir="data/c", split="train" Thanks in advance I think it should be possible if you specify the directory name in the allow patterns filter for snapshot download . I’m not very familiar with it, but it seems that the datasets library itself also has functions to extract a part of a dataset. yeah, it has data dir and data files, but what I was looking for was be able to download 10 GBs or 10 files of data , what the above will do is download the complete sub directory. I think this is close, but technically speaking, it’s not the number of files, and you can’t specify the size. You can use HfApi to get a list of the files in a repo and see their sizes, but that’s almost like doing it manually… yes, thats what I did, thanks for the help btw. Hello, Is the method applicable for zipped filestoo? I have some zippred files aroudn 20 30 GB but I dont want to download them as a whole Best, Deniz