cd/entity/DistServe· home entities DistServe
grep -l @distserve /news/*.json | wc -l → 5

DistServe

mentions 5 type Organization feed RSS

// recent coverage 5 mentions

14:30
2026-08-19
hiraditya.github.io
artificial-intelligence

Two Schedulers, One SLO

A vLLM RFC from the llm-d team warns that disaggregated inference deployments, where prefill and decode run on separate schedulers, can trigger recomputation-based preemption inside the decode instanc…

09:00
2026-08-11
blog.doubleword.ai
large-language-models

The case for disaggregated LLM serving

Disaggregated LLM serving, which runs prefill and decode on separate GPU pools and transfers KV caches over the network, should always be used in practice under sufficient load, according to a technic…

04:00
2026-08-03
arxiv.org
artificial-intelligence

Topology-Aware Data Movement for Disaggregated GPU Inference

A new arXiv paper (arXiv:2607.28633v1) proposes a topology-aware transfer orchestrator for disaggregated GPU inference, claiming existing systems like DistServe, Splitwise, and Mooncake ignore that ba…

// co-occurs with top 8 entities
// topics top 6 topics