{"slug": "how-to-improve-my-tokens-per-second", "title": "How to improve my tokens per second?", "summary": "An optimized training stack can reach roughly 51% model FLOPs utilization (MFU) on consumer GPUs, according to the LLMQ paper, leaving headroom for a user currently getting about 65 TFLOPS of useful compute at roughly 6 times params FLOPs per token plus gradient-checkpointing recompute. The recommended fixes, in order of expected payoff, are applying torch.compile to fuse RMSNorm, RoPE and SwiGLU into fewer kernels — worth about 1.5x in published single-GPU training benchmarks — moving data off /mnt/c, pre-tokenizing and packing sequences to eliminate padding, and using a fused optimizer. The guidance notes that the large matmuls are compute-bound while norms, RoPE, SwiGLU, softmax and logits are memory-bandwidth-bound, which on a 16 GB card at about 960 GB/s explains why kernel fusion and FP8 deliver large gains, and recommends running the PyTorch profiler for a few steps first to identify whether memory-bound ops, recompute or data loading is the bottleneck.", "body_md": "You’re at roughly 65 TFLOPS of useful compute (about 6 × params FLOPs per token, plus the recompute from gradient checkpointing), which leaves headroom: the LLMQ paper reports around 51% MFU on consumer GPUs with an optimized stack. In order of expected payoff:\n\n`torch.compile` on the model. It fuses RMSNorm, RoPE and SwiGLU into fewer kernels and was worth about 1.5x in published single-GPU training benchmarks.`/mnt/c`, pre-tokenize and pack sequences so there’s no padding, and use a fused optimizer.\nRun the PyTorch profiler for a few steps first to see whether you’re bound by memory-bound ops, recompute or data loading, so you apply the right fix.\n\nAlso worth knowing: the big matmuls are compute-bound, but everything between them (norms, RoPE, SwiGLU, softmax, logits) is memory-bandwidth-bound, and on a 16 GB card with ~960 GB/s that’s where a lot of your step time goes. That’s why kernel fusion (compile, Liger) gives such large gains, and why FP8 helps beyond just faster matmuls.", "url": "https://wpnews.pro/news/how-to-improve-my-tokens-per-second", "canonical_source": "https://discuss.huggingface.co/t/how-to-improve-my-tokens-per-second/190270#post_2", "published_at": "2026-10-10 11:57:10+00:00", "updated_at": "2026-10-10 12:44:43.907017+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "mlops", "ai-infrastructure", "large-language-models"], "entities": ["PyTorch", "torch.compile", "LLMQ", "Liger", "RMSNorm", "RoPE", "SwiGLU", "FP8"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-improve-my-tokens-per-second", "markdown": "https://wpnews.pro/news/how-to-improve-my-tokens-per-second.md", "text": "https://wpnews.pro/news/how-to-improve-my-tokens-per-second.txt", "jsonld": "https://wpnews.pro/news/how-to-improve-my-tokens-per-second.jsonld"}}