The Architecture of Modern Inference: Engineering the Disaggregated LLM Lifecycle
Baseten engineers Philip Kiely and Ali Taha detailed the shift to disaggregated prefill/decode serving, predictive layer-wise quantization error cancellation, hardware interconnect bottlenecks, and st…