12:05
2026-10-07
dev.to
ai-infrastructure
The KV Cache Is the Bottleneck Now — 1-Bit Quantization, MoE Stragglers, and Per-Second GPU Billing
A weekly LLM inference digest highlights new research aimed at KV cache and MoE serving bottlenecks, including TaSQ, a 1-bit KV cache quantization method that its SGLang implementation on a single RTX…