17:36
2026-09-09
aiflash.com
ai-infrastructure
The Architecture of Modern Inference: Engineering the Disaggregated LLM Lifecycle
Baseten engineers Philip Kiely and Ali Taha detailed the shift to disaggregated prefill/decode serving, predictive layer-wise quantization error cancellation, hardware interconnect bottlenecks, and stβ¦