cd /news/machine-learning/how-to-improve-my-tokens-per-second · home › topics › machine-learning › article
[ARTICLE · art-148728] src=discuss.huggingface.co ↗ pub= topic=machine-learning verified=true sentiment=· neutral

How to improve my tokens per second?

An optimized training stack can reach roughly 51% model FLOPs utilization (MFU) on consumer GPUs, according to the LLMQ paper, leaving headroom for a user currently getting about 65 TFLOPS of useful compute at roughly 6 times params FLOPs per token plus gradient-checkpointing recompute. The recommended fixes, in order of expected payoff, are applying torch.compile to fuse RMSNorm, RoPE and SwiGLU into fewer kernels — worth about 1.5x in published single-GPU training benchmarks — moving data off /mnt/c, pre-tokenizing and packing sequences to eliminate padding, and using a fused optimizer. The guidance notes that the large matmuls are compute-bound while norms, RoPE, SwiGLU, softmax and logits are memory-bandwidth-bound, which on a 16 GB card at about 960 GB/s explains why kernel fusion and FP8 deliver large gains, and recommends running the PyTorch profiler for a few steps first to identify whether memory-bound ops, recompute or data loading is the bottleneck.

read1 min views2 publishedOct 10, 2026

You’re at roughly 65 TFLOPS of useful compute (about 6 × params FLOPs per token, plus the recompute from gradient checkpointing), which leaves headroom: the LLMQ paper reports around 51% MFU on consumer GPUs with an optimized stack. In order of expected payoff:

torch.compile on the model. It fuses RMSNorm, RoPE and SwiGLU into fewer kernels and was worth about 1.5x in published single-GPU training benchmarks./mnt/c, pre-tokenize and pack sequences so there’s no padding, and use a fused optimizer. Run the PyTorch profiler for a few steps first to see whether you’re bound by memory-bound ops, recompute or data , so you apply the right fix.

Also worth knowing: the big matmuls are compute-bound, but everything between them (norms, RoPE, SwiGLU, softmax, logits) is memory-bandwidth-bound, and on a 16 GB card with ~960 GB/s that’s where a lot of your step time goes. That’s why kernel fusion (compile, Liger) gives such large gains, and why FP8 helps beyond just faster matmuls.

── more in #machine-learning 4 stories · sorted by recency
── more on @pytorch 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-improve-my-to…] indexed:0 read:1min 2026-10-10 · —