09:00
2026-08-11
blog.doubleword.ai
large-language-models
The case for disaggregated LLM serving
Disaggregated LLM serving, which runs prefill and decode on separate GPU pools and transfers KV caches over the network, should always be used in practice under sufficient load, according to a technicβ¦