12:45
2026-10-09
github.com
ai-infrastructure
Show HN: Narwhal ā LLM serving that moves GPUs between prefill/decode in seconds
Narwhal, an open-source adaptive disaggregated LLM inference framework, hot-swaps prefill and decode roles across a fixed GPU fleet in seconds without reloading model weights, scaling from a single GPā¦