Part 6 of this series name-dropped pd-ratio-coordinator in passing — a Kubernetes operator I’d written to autonomously rebalance prefill and decode replica counts in an llm-d cluster, reacting to queue velocity and a joint GPU budget instead of the generic queue-depth checks a standard autoscaler uses. I cited it as “the part that connects directly” to the memory-wall argument. What I didn’t say at the time: it had never actually run against a real cluster. ~1,600 lines of Go, unit-tested, never GPU-tested.
This post is what happened when that stopped being true.
What pd-ratio-coordinator Actually Does #
llm-d separates inference into two pools — prefill (compute-intensive first-token generation) and decode (memory-bandwidth-intensive token generation). WVA, the standard autoscaler, scales each pool independently based on current queue depth. Two gaps that leaves:
Spike blindness. WVA reacts after a queue has already saturated. pd-ratio-coordinator detects queue velocity — rate of growth — and scales before saturation hits.
Pool competition. Without a joint constraint, prefill and decode can
each independently request more GPUs than the cluster has, leaving both
pools stuck Pending. pd-ratio-coordinator enforces
prefill.replicas + decode.replicas <= gpuBudget as a hard constraint,
always.
It also handles graceful decode drain (label a pod, wait for in-flight sequences to finish, then scale down — decode pods hold live KV cache, so killing one mid-sequence drops a user’s request) and cooldown/hysteresis to prevent oscillation.
Before Any of That Could Be Tested: Does It Even Compile? #
No. go build ./... failed on the actual committed code (branch KRAG-1).
Two real bugs, not just missing dependencies:
- A DeepCopy type bug in
register.go—PDRatioPolicyStatus’s hand-writtenDeepCopyIntotried to append[]interface{}into a[]metav1.Conditionfield. Fixed to a proper typed copy. - Unqualified constants in
scaler/decode.go—BottleneckNone,BottleneckPrefill,BottleneckDecode,BottleneckBothwere referenced as if they were local, but they’re defined inapi/v1alpha1and were never imported. Qualified all four.
go.sum had also never been committed, so the module wouldn’t build for
anyone who cloned it fresh. And config/crd/ — the directory the README
instructs users to kubectl apply -f — didn’t exist. controller-gen
needed a +groupName marker that was never added (the group was only set
as a Go runtime variable, which static marker-scanning can’t see).
All four fixed, pushed to KRAG-1, before any GPU was rented. If you’re
going to validate a tool, “does it compile” is not optional pre-work.
The Scope Decision: Skip Real P/D Disaggregation on Purpose #
The repo’s own docs/testing.md assumes real P/D disaggregation is
already running — NIXL over NVLink, prefill and decode on separate GPUs
with working KV transfer. Two problems with that assumption. First, Part
5
already found that a single time-sliced node has no RDMA path — NIXL
fails, the architecture collapses into aggregated serving. Second, the
“1x GH200 + H100” combo instance the docs assume may not even be a real
Lambda SKU — the console only ever showed separate GH200, H100, A10, and
A100 listings.
Here’s the way out: pd-ratio-coordinator doesn’t actually care whether
KV transfer between prefill and decode is real. Its entire job is
reading Prometheus metrics and calling the Kubernetes API to scale
Deployments. It never inspects NIXL, never touches KV transfer
correctness. So the tool’s whole value proposition — velocity detection,
joint GPU budget, graceful drain, cooldown — can be validated with two
plain vLLM Deployments labeled prefill and decode, real Locust load,
real Prometheus scraping, real kubectl scale calls, and zero dependency
on cross-GPU KV transfer working at all.
That means one GPU is enough. No NVLink, no combo SKU, no repeating Part 5’s unresolved problem for a reason that has nothing to do with what this tool actually does.
Setup #
Single-node k3s on a rented A100 (40GB). Two plain vLLM Deployments,
Qwen3-0.6B, labeled prefill and decode, both sharing the one physical
GPU via --gpu-memory-utilization=0.2 each rather than requesting
isolated nvidia.com/gpu units — deliberately, so multiple replicas could
coexist on one card. Lightweight annotation-based Prometheus (no Helm
chart). The controller itself run via make run directly on the host,
not as an in-cluster pod — simpler to iterate on, closer to how I was
actually debugging it.
What broke getting the cluster up (kept in full, same as every other post in this series)
- No GPU access at all, at first. Skipping
nvidia.com/gpuresource requests also meant containerd never injected GPU device access into the containers —RuntimeError: Failed to infer device type. Fixed with a KubernetesRuntimeClassobject (handler: nvidia) andruntimeClassName: nvidiaon the pod specs, rather than forcing nvidia as the cluster-wide containerd default (which would also apply to Prometheus, which needs no GPU at all). - **``` containerd: failed to unmarshal TOML: toml: table containerd already exists
, then`table nvidia already exists` on the next attempt. Turns
out k3s on this GPU-ready box already auto-detects`nvidia-container-runtime` and configures a working`nvidia` runtime
handler on its own. My custom containerd template wasn’t needed at all
— removed entirely once confirmed.
3. **Stale containerd-shim processes blocked a clean k3s restart.** Same
“child process survives killing the parent” shape as the zombie
EngineCore issue from the earlier vLLM post — containerd shims are*designed* to survive their parent’s death, which is exactly what makes
them annoying here. Fixed with k3s’s own`k3s-killall.sh` before
restarting.
4. **In-cluster service DNS doesn’t resolve from the host.** Running the
controller outside the cluster meant`http://prometheus.pd-validation:9090` couldn’t resolve. Fixed with a persistent port-forward and pointing`prometheusURL` at`localhost` .
5. **`go run` spawns a child binary that survives killing the parent.** The zombie-child pattern a third time, now in Go’s own tooling —`pkill -f 'go run main.go'` doesn’t touch the actual compiled binary at`/tmp/go-build.../exe/main` . Had to find and kill it by PID directly.
6. **Two real metric-name drifts in vLLM v0.30.0** , found by diffing the
operator’s hardcoded PromQL against the server’s actual`/metrics` output:
- `vllm:gpu_cache_usage_perc` → renamed to`vllm:kv_cache_usage_perc` —
the*same* rename[Part 2](https://kraghavan.ca/llm-infrastructure/inference/2026/04/16/vllm-ollama-apple-silicon-experiment2.html) already documented as a gotcha, on a different vLLM version. This
metric name has been drifting for a while, not a one-off.
- `vllm:time_per_output_token_seconds` → renamed to`vllm:request_time_per_output_token_seconds` .
Both fixed in `internal/metrics/prometheus.go` , rebuilt, redeployed.
## Results
Four tests, adapted from the repo’s own `docs/testing.md`, de-risked as
described above.
### Test 1 — TPOT SLO breach → scale decode: passed
30 concurrent users, 60s, against decode. **1,970 requests, 0 failures**,
p50 530ms, p95 560ms.
Real finding before the real pass: measured `decode.tpot_p95_ms` sat at
**9.5ms** under this load — far below the first “aggressive-sounding” 20ms
SLO I set, which never triggered anything. Qwen3-0.6B decodes faster than
that guess, even sharing a GPU at 20% utilization. Had to lower the SLO to
5ms — below the measured baseline — to get a clean, deterministic breach.
Stated plainly: that’s a hand-tuned threshold to exercise the code path,
not a realistic production number for this model.
Once it triggered: `tpot_slo_breach(10ms>5ms)` in the log, cooldown
respected, then a real scale action — `"scaled decode", "from": 1, "to": 2`
— and a real second decode pod came up.
### Test 2 — queue velocity spike → scale prefill: inconclusive, and that’s the honest result
Escalated three times — 20 users, then 150, then 300. Final run:
**12,517 requests, 0 failures**, 520.7 req/s sustained, p50 470ms, p95
560ms, p99 680ms.
`prefill.queue` stayed at **zero for the entire duration of all three
runs** — confirmed both by the controller’s own 10-second reconcile
snapshots and by polling Prometheus directly every 2 seconds during the
final burst. Even at 300 concurrent requests and 500+ req/s, nothing ever
backed up into a waiting state.
Most likely reason: vLLM’s continuous batching on a 0.6B model admits
requests into the running batch fast enough that nothing queues, at any
concurrency this setup could throw at it. That’s a property of this
model/hardware pairing, not a bug in the detection code — `AnalysePrefill`
already has passing unit tests against synthetic queue data. What this
round couldn’t do is produce a real-world backlog to exercise that path
operationally. A v2 validation would need either a much larger model or a
deliberately starved `--max-num-seqs` to actually generate queue depth.
### Test 3 — GPU budget enforcement: passed
Budget set equal to the current total (prefill=1, decode=2, budget=3). 30
users, 40s, against decode. **1,376 requests, 0 failures**, p50 530ms,
p95 550ms.
Real pressure was present — `decode.under_pressure: true`,
`tpot_slo_breach(10ms>5ms)` — and decode replicas stayed at 2 the entire
time. The controller correctly refused to scale past budget even with
genuine, sustained pressure asking it to. `prefill + decode <= gpuBudget`
held throughout. Clean, unambiguous result.
### Test 4 — drain before scale-down: nuanced, and worth explaining rather than rounding up
8 users, 70s, against decode. **696 requests, 0 failures**, p50 470ms,
p95 480ms. Budget cut from 3 to 2 mid-flight to force a scale-down.
The `llmd.io/draining=true` label was applied correctly to the targeted
pod. But the drain **timed out after its configured 20 seconds** —
`"drain timeout, force scaling down"` — and the controller force-scaled
down anyway.
The more important finding is underneath that: this setup uses a plain
Kubernetes `Service` (kube-proxy round-robin), not llm-d’s actual EPP
gateway. The draining label is designed to be read by that EPP to stop
routing new requests to the pod — a plain Service has no concept of the
label at all. So the drain mechanism’s *write path* — apply label, poll
in-flight request count, force-scale on timeout — is confirmed to execute
correctly end to end. Its *actual protective effect* was not validated
here, because half the mechanism (routing exclusion) isn’t present in
this de-risked setup. Zero failures is good news, but it may simply mean
nothing happened to be in-flight on that specific pod at the moment it
was killed, not that the drain protected it. Real validation of this
guarantee needs llm-d’s EPP in front of the pods — explicitly outside
this round’s scope, not a result I’m claiming.
### Summary
| Test | Result | Requests | Failures | p50 | p95 |
|---|---|---|---|---|---|
| TPOT breach → scale decode | **Passed** | 1,970 | 0 | 530ms | 560ms |
| Velocity spike → scale prefill | **Inconclusive** (real limitation) | 12,517 | 0 | 470ms | 560ms |
| GPU budget enforcement | **Passed** | 1,376 | 0 | 530ms | 550ms |
| Drain before scale-down | **Nuanced** (write path confirmed, protection unverified) | 696 | 0 | 470ms | 480ms |
**16,559 total requests across every test. Zero failures, throughout.**
## What This Means
Two of four mechanisms are cleanly proven: the velocity/SLO detection logic reacts to real signals and makes real scaling decisions, and the joint GPU budget constraint holds even under genuine competing pressure. One is a real, named limitation — this model/hardware pairing couldn’t produce the queue backlog needed to exercise spike detection, through no fault of the detection code itself. One is a genuinely nuanced result that I’d rather report honestly than round up to a pass: the drain mechanism’s bookkeeping works, but proving it protects real traffic needs a piece of infrastructure (llm-d’s EPP) this round deliberately scoped out.
That’s four real answers, not four passes. I think that’s more useful to anyone deciding whether to actually run this operator than a clean scorecard would have been.
**Open for v2:** a workload that can genuinely saturate prefill admission
(bigger model, or a resource-starved config), and a real llm-d EPP in
front of the pods to validate the drain label’s actual routing-exclusion
effect, not just its bookkeeping.
## Is This Novel? No — And Here’s Who Already Did It
Before publishing this, I asked for an honest review of whether any of this is actually new. Short answer: no. This is a small, llm-d-shaped implementation of a pattern that already exists in production elsewhere, and the post should say that plainly instead of letting “I built an operator” imply more than it does.
Specifically, checked against the primary sources, not just taken on faith:
- **[NVIDIA Dynamo’s Planner](https://docs.nvidia.com/dynamo/v1.2.1/components/planner/planner-guide)** already has a`load` scaling mode that reacts to prefill-queue-token
and decode-KV-utilization thresholds — structurally the same idea as`AnalysePrefill` /`AnalyseDecode` here. It also already enforces a joint
GPU budget across prefill and decode (`max_gpu_budget` , a hard cap on
combined replicas) — the exact mechanism Test 3 validated above. Not
novel. Confirmed by reading the config reference directly.
- **[HeteroScale](https://arxiv.org/abs/2508.19559)** (ByteDance,
published August 2025) runs a single joint autoscaling metric across
prefill/decode pools in production on tens of thousands of GPUs. The
entire “pools compete for GPUs, scale them together” framing of this
post is their result at a scale this validation can’t touch. Not
novel.
- llm-d’s own
[autoscaling roadmap](https://github.com/llm-d/llm-d-autoscaling/issues/1079) already lists a “rate-based (velocity) scaling signal” as a planned
item. The queue-velocity idea in`AnalysePrefill` isn’t an invention,
it’s llm-d’s own stated direction, arrived at independently and about
a release early.
- One genuinely useful finding did fall out of checking Dynamo, though:
its docs state the `load` planner mode is**non-functional** on vLLM,
SGLang, and TRT-LLM today, because none of those backends expose real
prefill queue metrics. That’s the same wall Test 2 hit. It means the
inconclusive result above isn’t a quirk of this 0.6B model on one
A100 — it’s a structural gap in what current inference engines expose,
seen independently by a much bigger team. That’s worth more than a
clean pass would have been.
What’s left standing after all that: the drain-before-scale-down mechanism (Test 4) is the one piece I could not find already shipped anywhere. I checked llm-d’s own autoscaling roadmap and Dynamo’s planner guide directly for anything about draining a pod or excluding it from routing before scale-down — neither mentions it. That doesn’t make it novel research; it’s a small, missing piece of plumbing for one specific ecosystem (llm-d), not a new idea. And it’s also the one Test 4 couldn’t actually validate end-to-end, since that requires a real llm-d EPP in front of the pods to prove the exclusion works, not just that the label gets set correctly.
So: a tiny, real dent, in a very large and already well-populated field. If you’re evaluating this for your own cluster, evaluate it as “a lightweight llm-d-native version of what Dynamo and ByteDance already ship at scale,” not as a new approach to the problem.
## The Scripts
Everything here — the k8s manifests, the Locust load generator, the setup
and test-runner scripts, the Go fixes — is in
[gpu-labs](https://github.com/kraghavan/gpu-labs/tree/main/pd-ratio-coordinator-validation),
alongside the fixed-up [pd-ratio-coordinator](https://github.com/kraghavan/pd-ratio-coordinator/tree/KRAG-1)
repo itself. If you find a way to get real queue backlog out of a small
model, or want to wire up a real EPP for the drain test, I’d genuinely
like to know.
*Experiments run on Lambda Cloud, 1x A100 (40GB SXM4), k3s, vLLM v0.30.0,
Qwen3-0.6B, pd-ratio-coordinator branch KRAG-1. Platform engineer with
11+ years in distributed systems going deep on LLM serving
infrastructure.*