Stop letting your GPU idle while your CPU struggles to feed it PyTorch developers can eliminate GPU idle time caused by CPU bottlenecks by using multi-process data loading with `num_workers` set to the number of CPU cores, `pin_memory=True`, and binary formats like TFRecord, Apache Parquet, or WebDataset instead of raw CSVs or JSON files. The article provides a custom dataset class example and recommends memory mapping and sharding for datasets exceeding RAM and for distributed training across multiple GPUs. Stop letting your GPU idle while your CPU struggles to feed it Handling the Bottleneck with Prefetching and Parallelism The most common mistake is loading data synchronously. When the model finishes a batch, the GPU sits idle while the CPU fetches the next chunk from the disk. You can kill this latency by using multi-process loading. In PyTorch, this is handled via the num workers parameter in the DataLoader . Setting num workers to the number of CPU cores usually helps, but be careful with memory overhead. If you're using a massive dataset, you should combine this with pin memory=True , which speeds up the transfer from CPU RAM to GPU VRAM by using page-locked memory. Optimized Formats for Large Scale Training Stop using raw CSVs or thousands of tiny JSON files. Opening and closing files creates massive overhead. For a real-world deployment, you need binary formats that support sequential reads and memory mapping. TFRecord: The gold standard for TensorFlow, storing data as a sequence of binary records. Apache Parquet: Incredible for tabular data due to columnar storage, which means you only load the features you actually need. WebDataset: Essential for vision tasks; it wraps data into POSIX tar files, allowing you to stream datasets over a network without needing to download the whole thing to a local SSD first. A Practical Tutorial for Custom Data Pipelines If you're building a custom LLM agent or a fine-tuning script, you'll likely need a custom dataset class. Here is a basic structure to ensure your data is preprocessed on the fly without blocking the training loop. python import torch from torch.utils.data import Dataset, DataLoader class EfficientDataset Dataset : def init self, data path : Load metadata or index files here, not the full dataset self.data = self. load index data path def len self : return len self.data def getitem self, idx : Perform heavy transformations here sample = self.data idx processed sample = self.transform sample return torch.tensor processed sample def transform self, x : Example: Normalization or tokenization return x / 255.0 Deployment configuration for maximum throughput loader = DataLoader dataset=EfficientDataset "data/train" , batch size=64, shuffle=True, num workers=8, pin memory=True, prefetch factor=2 Memory Mapping and Sharding When your dataset exceeds your system RAM, memory mapping mmap is your best friend. It allows the OS to map a file directly into the virtual address space, loading pages only when they are accessed. For distributed training across multiple GPUs, you must implement sharding. This ensures that each GPU sees a unique subset of the data per epoch, preventing redundant computation and ensuring the gradient updates are based on a diverse sample of the global dataset. This is the only way to scale a deep dive project from a single local machine to a cluster. Stop expecting LLMs to be databases because they are 1d ago /en/news/6543/ Since the provided content was only a title 2d ago /en/news/6403/ Sign language AI finally works on a mobile device 4d ago /en/news/6150/ Jeff Dean is chasing a 10 billion dollar valuation for his new 4d ago /en/news/6135/ Medical AI is still hallucinating stereotypes into patient care 7d ago /en/news/5751/ DeepMind WeatherNext actually predicts cyclones with scary 9d ago /en/news/5584/ Next Will Gen Z actually survive the AI takeover of entry-level roles? → /en/news/6721/ these AI tool field notes https://tanyan888.com/ , with plenty of directly applicable cases.