{"slug": "show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds", "title": "Show HN: Narwhal – LLM serving that moves GPUs between prefill/decode in seconds", "summary": "Narwhal, an open-source adaptive disaggregated LLM inference framework, hot-swaps prefill and decode roles across a fixed GPU fleet in seconds without reloading model weights, scaling from a single GPU to multi-node deployments. The role controller scores the current split and each adjacent split one engine move away, projecting each from measured engine profiles, offered demand, and resident work, and scores a split by its worst projected SLO ratio across TTFT, TPOT, and decode queueing. It installs via `pip install narwhal-inference` on Linux with Python 3.11 or newer, and its dev mode runs two to eight engines on one NVIDIA CUDA GPU under Ubuntu or WSL2, with the installed template starting two engines on a GPU with 8 GB of VRAM or less and the RTX 5090 reference template starting four engines on an RTX 5090.", "body_md": "Narwhal is an adaptive, disaggregated inference framework that automatically hot-swaps prefill and decode roles as demand changes, and without having to reload model weights. It can scale from a [single GPU](https://athrael-soju.github.io/Narwhal/Dev-Runtime/) to [multi-node deployments](https://athrael-soju.github.io/Narwhal/Deploy/).\n\n| Capability | Behavior | Guide | \n|---|---|---|\n| Role hot-swap | Reassigns prefill and decode roles across a fixed GPU fleet, with NIXL key-value (KV) transfer between them |  | \n| Serving | Serves completion and chat requests, streamed or buffered, with latency-aware admission |  | \n| Fault tolerance | Fails over to a warm-standby router and readmits engines against their live process generation |  | \n| Measurement | Profiles engines, runs ordered benchmark points, and keeps the evidence from each point |  | \n| Observability | Exports router and engine metrics to Prometheus and a provisioned Grafana dashboard |  | \n| Operator tooling | Validates fleet files offline and collects private diagnostic bundles |  | \n| Development mode | Runs two to eight engines on one NVIDIA CUDA GPU under Ubuntu or WSL2 |  | \n\nThe role controller scores the current role split and each adjacent split, one engine move away. It projects each split from measured engine profiles, offered demand, and resident work. A split's score is its worst projected service-level objective (SLO) ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing.\n\nProjections use measured window demand. A decode-to-prefill candidate takes its decode demand from the larger of the short- and long-horizon estimates.\n\nThe controller moves to an adjacent split that improves the score by at least the configured margin. A decode-to-prefill move also needs stable decode demand and a closed arrival-evidence window. The window closes after `controller.reactive.evidence_span_s` with the minimum number of arrivals, or after `controller.reactive.evidence_max_span_s` under sparse traffic. A prefill-to-decode move with prefill load at or below `controller.thresholds.shrink` can proceed while the window is open.\n\nWhen demand over the confirmation span shifts after a settled run, the controller moves one engine. The settled run is `controller.reactive.evidence_span_s`, or one confirmation span shorter when the shift reverses the controller's recent moves. The controller keeps moving engines in that direction on confirmation-span demand while the shift lasts: after its first move for a reversing shift, and after `controller.reactive.evidence_span_s` for any other shift. Under steady demand, the score chooses between adjacent splits once the arrival-evidence window has closed.\n\nEvery move passes guards for pinned engines, role floors, cooldown, dwell time, the resident-stream cap on decode donors, and engine lifecycle holds. While a role is below its configured floor, floor repair moves one engine per monitor pass.\n\nNew requests follow the revised split, and resident requests finish on their assigned engines.\n\nInstall on Linux with Python 3.11 or newer:\n\n```\npython3 -m venv .venv\nsource .venv/bin/activate\npython -m pip install narwhal-inference\nnarwhal-serve --version\nnarwhal --help\n```\n\nThe wheel installs these commands:\n\nNarwhal dev runs a local NVIDIA CUDA fleet on Ubuntu, either directly or under WSL2.\n\n```\nnarwhal dev init\nnarwhal dev up\nnarwhal dev verify\nnarwhal dev status\nnarwhal dev down\n```\n\nThe installed template starts two engines on an NVIDIA GPU with 8 GB of VRAM or less. The RTX 5090 reference template starts four engines on an RTX 5090.\n\nRun these gates in order from a management workstation:\n\n1. Freeze inputs and discover the deployment in [Gate A](https://athrael-soju.github.io/Narwhal/deploy/01-Discover/) .\n2. Package and install the approved revision in [Gate B](https://athrael-soju.github.io/Narwhal/deploy/02-Install/) .\n3. Validate and start every engine in [Gate C](https://athrael-soju.github.io/Narwhal/deploy/03-Validate-Engines/) .\n4. Qualify the transfer fabric in [Gate D](https://athrael-soju.github.io/Narwhal/deploy/04-Qualify-Fabric/) .\n5. Attest the live engines in [Gate E](https://athrael-soju.github.io/Narwhal/deploy/05-Attest/) .\n6. Profile the engines and run preflight in [Gate F](https://athrael-soju.github.io/Narwhal/deploy/06-Profile-and-Preflight/) .\n7. Start the router and validate capacity through an SSH tunnel in [Gate G](https://athrael-soju.github.io/Narwhal/deploy/07-Serve-and-Measure/) .\n\nNarwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).", "url": "https://wpnews.pro/news/show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds", "canonical_source": "https://github.com/athrael-soju/Narwhal", "published_at": "2026-10-09 12:45:01+00:00", "updated_at": "2026-10-09 12:54:05.856333+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-tools", "mlops", "developer-tools"], "entities": ["Narwhal", "NIXL", "NVIDIA", "CUDA", "Prometheus", "Grafana", "RTX 5090", "WSL2"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds", "markdown": "https://wpnews.pro/news/show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds.md", "text": "https://wpnews.pro/news/show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds.txt", "jsonld": "https://wpnews.pro/news/show-hn-narwhal-llm-serving-that-moves-gpus-between-prefill-decode-in-seconds.jsonld"}}