{"slug": "pd-ratio-coordinator-does-my-own-tool-actually-work", "title": "pd-ratio-coordinator — Does My Own Tool Actually Work?", "summary": "A developer's Kubernetes operator, pd-ratio-coordinator, failed to compile on its committed KRAG-1 branch, with two real bugs — a DeepCopy type mismatch in register.go and four unqualified Bottleneck constants in scaler/decode.go — plus a missing go.sum and a nonexistent config/crd/ directory, before any GPU testing could begin. The ~1,600-line Go tool autonomously rebalances prefill and decode replica counts in an llm-d cluster using queue velocity and a hard prefill.replicas + decode.replicas <= gpuBudget constraint, and the author validated its velocity detection, joint GPU budget, graceful drain and cooldown behavior against two plain vLLM Deployments with Locust load and Prometheus scraping, deliberately skipping real P/D disaggregation because a single time-sliced node has no RDMA path for NIXL.", "body_md": "Part 6 of this series name-dropped [pd-ratio-coordinator](https://github.com/kraghavan/pd-ratio-coordinator)\nin passing — a Kubernetes operator I’d written to autonomously rebalance\nprefill and decode replica counts in an llm-d cluster, reacting to queue\n*velocity* and a joint GPU budget instead of the generic queue-depth\nchecks a standard autoscaler uses. I cited it as “the part that connects\ndirectly” to the memory-wall argument. What I didn’t say at the time: it\nhad never actually run against a real cluster. ~1,600 lines of Go,\nunit-tested, never GPU-tested.\n\nThis post is what happened when that stopped being true.\n\n## What pd-ratio-coordinator Actually Does\n\nllm-d separates inference into two pools — prefill (compute-intensive first-token generation) and decode (memory-bandwidth-intensive token generation). WVA, the standard autoscaler, scales each pool independently based on current queue depth. Two gaps that leaves:\n\n**Spike blindness.** WVA reacts after a queue has already saturated.\npd-ratio-coordinator detects queue *velocity* — rate of growth — and\nscales before saturation hits.\n\n**Pool competition.** Without a joint constraint, prefill and decode can\neach independently request more GPUs than the cluster has, leaving both\npools stuck Pending. pd-ratio-coordinator enforces\n`prefill.replicas + decode.replicas <= gpuBudget` as a hard constraint,\nalways.\n\nIt also handles graceful decode drain (label a pod, wait for in-flight\nsequences to finish, *then* scale down — decode pods hold live KV cache,\nso killing one mid-sequence drops a user’s request) and cooldown/hysteresis\nto prevent oscillation.\n\n## Before Any of That Could Be Tested: Does It Even Compile?\n\nNo. `go build ./...` failed on the actual committed code (branch `KRAG-1`).\nTwo real bugs, not just missing dependencies:\n\n1. **A DeepCopy type bug in `register.go`** —` PDRatioPolicyStatus` ’s\nhand-written`DeepCopyInto` tried to append`[]interface{}` into a`[]metav1.Condition` field. Fixed to a proper typed copy.\n2. **Unqualified constants in `scaler/decode.go`** —` BottleneckNone` ,`BottleneckPrefill` ,`BottleneckDecode` ,`BottleneckBoth` were referenced\nas if they were local, but they’re defined in`api/v1alpha1` and were\nnever imported. Qualified all four.\n\n`go.sum` had also never been committed, so the module wouldn’t build for\nanyone who cloned it fresh. And `config/crd/` — the directory the README\ninstructs users to `kubectl apply -f` — didn’t exist. `controller-gen`\nneeded a `+groupName` marker that was never added (the group was only set\nas a Go runtime variable, which static marker-scanning can’t see).\n\nAll four fixed, pushed to `KRAG-1`, before any GPU was rented. If you’re\ngoing to validate a tool, “does it compile” is not optional pre-work.\n\n## The Scope Decision: Skip Real P/D Disaggregation on Purpose\n\nThe repo’s own `docs/testing.md` assumes real P/D disaggregation is\nalready running — NIXL over NVLink, prefill and decode on separate GPUs\nwith working KV transfer. Two problems with that assumption. First, [Part\n5](https://kraghavan.ca/llm-infrastructure/inference/2026/04/21/llm-d-pd-disaggregation.html)\nalready found that a single time-sliced node has no RDMA path — NIXL\nfails, the architecture collapses into aggregated serving. Second, the\n“1x GH200 + H100” combo instance the docs assume may not even be a real\nLambda SKU — the console only ever showed separate GH200, H100, A10, and\nA100 listings.\n\nHere’s the way out: **pd-ratio-coordinator doesn’t actually care whether\nKV transfer between prefill and decode is real.** Its entire job is\nreading Prometheus metrics and calling the Kubernetes API to scale\nDeployments. It never inspects NIXL, never touches KV transfer\ncorrectness. So the tool’s whole value proposition — velocity detection,\njoint GPU budget, graceful drain, cooldown — can be validated with two\nplain vLLM Deployments labeled `prefill` and `decode`, real Locust load,\nreal Prometheus scraping, real `kubectl scale` calls, and zero dependency\non cross-GPU KV transfer working at all.\n\nThat means **one GPU is enough.** No NVLink, no combo SKU, no repeating\nPart 5’s unresolved problem for a reason that has nothing to do with what\nthis tool actually does.\n\n## Setup\n\nSingle-node k3s on a rented A100 (40GB). Two plain vLLM Deployments,\nQwen3-0.6B, labeled `prefill` and `decode`, both sharing the one physical\nGPU via `--gpu-memory-utilization=0.2` each rather than requesting\nisolated `nvidia.com/gpu` units — deliberately, so multiple replicas could\ncoexist on one card. Lightweight annotation-based Prometheus (no Helm\nchart). The controller itself run via `make run` directly on the host,\nnot as an in-cluster pod — simpler to iterate on, closer to how I was\nactually debugging it.\n\n### What broke getting the cluster up (kept in full, same as every other post in this series)\n\n1. **No GPU access at all, at first.** Skipping`nvidia.com/gpu` resource\nrequests also meant containerd never injected GPU device access into\nthe containers —`RuntimeError: Failed to infer device type` . Fixed\nwith a Kubernetes`RuntimeClass` object (`handler: nvidia` ) and`runtimeClassName: nvidia` on the pod specs, rather than forcing nvidia\nas the cluster-wide containerd default (which would also apply to\nPrometheus, which needs no GPU at all).\n2. **```\ncontainerd: failed to unmarshal TOML: toml: table containerd already\nexists\n```**\n, then`table nvidia already exists` on the next attempt. Turns\nout k3s on this GPU-ready box already auto-detects`nvidia-container-runtime` and configures a working`nvidia` runtime\nhandler on its own. My custom containerd template wasn’t needed at all\n— removed entirely once confirmed.\n3. **Stale containerd-shim processes blocked a clean k3s restart.** Same\n“child process survives killing the parent” shape as the zombie\nEngineCore issue from the earlier vLLM post — containerd shims are*designed* to survive their parent’s death, which is exactly what makes\nthem annoying here. Fixed with k3s’s own`k3s-killall.sh` before\nrestarting.\n4. **In-cluster service DNS doesn’t resolve from the host.** Running the\ncontroller outside the cluster meant`http://prometheus.pd-validation:9090` couldn’t resolve. Fixed with a persistent port-forward and pointing`prometheusURL` at`localhost` .\n5. **`go run` spawns a child binary that survives killing the parent.** The zombie-child pattern a third time, now in Go’s own tooling —`pkill -f 'go run main.go'` doesn’t touch the actual compiled binary at`/tmp/go-build.../exe/main` . Had to find and kill it by PID directly.\n6. **Two real metric-name drifts in vLLM v0.30.0** , found by diffing the\noperator’s hardcoded PromQL against the server’s actual`/metrics` output:\n  - `vllm:gpu_cache_usage_perc` → renamed to`vllm:kv_cache_usage_perc` —\nthe*same* rename[Part 2](https://kraghavan.ca/llm-infrastructure/inference/2026/04/16/vllm-ollama-apple-silicon-experiment2.html) already documented as a gotcha, on a different vLLM version. This\nmetric name has been drifting for a while, not a one-off.\n  - `vllm:time_per_output_token_seconds` → renamed to`vllm:request_time_per_output_token_seconds` .\n Both fixed in `internal/metrics/prometheus.go` , rebuilt, redeployed.\n\n## Results\n\nFour tests, adapted from the repo’s own `docs/testing.md`, de-risked as\ndescribed above.\n\n### Test 1 — TPOT SLO breach → scale decode: passed\n\n30 concurrent users, 60s, against decode. **1,970 requests, 0 failures**,\np50 530ms, p95 560ms.\n\nReal finding before the real pass: measured `decode.tpot_p95_ms` sat at\n**9.5ms** under this load — far below the first “aggressive-sounding” 20ms\nSLO I set, which never triggered anything. Qwen3-0.6B decodes faster than\nthat guess, even sharing a GPU at 20% utilization. Had to lower the SLO to\n5ms — below the measured baseline — to get a clean, deterministic breach.\nStated plainly: that’s a hand-tuned threshold to exercise the code path,\nnot a realistic production number for this model.\n\nOnce it triggered: `tpot_slo_breach(10ms>5ms)` in the log, cooldown\nrespected, then a real scale action — `\"scaled decode\", \"from\": 1, \"to\": 2`\n— and a real second decode pod came up.\n\n### Test 2 — queue velocity spike → scale prefill: inconclusive, and that’s the honest result\n\nEscalated three times — 20 users, then 150, then 300. Final run:\n**12,517 requests, 0 failures**, 520.7 req/s sustained, p50 470ms, p95\n560ms, p99 680ms.\n\n`prefill.queue` stayed at **zero for the entire duration of all three\nruns** — confirmed both by the controller’s own 10-second reconcile\nsnapshots and by polling Prometheus directly every 2 seconds during the\nfinal burst. Even at 300 concurrent requests and 500+ req/s, nothing ever\nbacked up into a waiting state.\n\nMost likely reason: vLLM’s continuous batching on a 0.6B model admits\nrequests into the running batch fast enough that nothing queues, at any\nconcurrency this setup could throw at it. That’s a property of this\nmodel/hardware pairing, not a bug in the detection code — `AnalysePrefill`\nalready has passing unit tests against synthetic queue data. What this\nround couldn’t do is produce a real-world backlog to exercise that path\noperationally. A v2 validation would need either a much larger model or a\ndeliberately starved `--max-num-seqs` to actually generate queue depth.\n\n### Test 3 — GPU budget enforcement: passed\n\nBudget set equal to the current total (prefill=1, decode=2, budget=3). 30\nusers, 40s, against decode. **1,376 requests, 0 failures**, p50 530ms,\np95 550ms.\n\nReal pressure was present — `decode.under_pressure: true`,\n`tpot_slo_breach(10ms>5ms)` — and decode replicas stayed at 2 the entire\ntime. The controller correctly refused to scale past budget even with\ngenuine, sustained pressure asking it to. `prefill + decode <= gpuBudget`\nheld throughout. Clean, unambiguous result.\n\n### Test 4 — drain before scale-down: nuanced, and worth explaining rather than rounding up\n\n8 users, 70s, against decode. **696 requests, 0 failures**, p50 470ms,\np95 480ms. Budget cut from 3 to 2 mid-flight to force a scale-down.\n\nThe `llmd.io/draining=true` label was applied correctly to the targeted\npod. But the drain **timed out after its configured 20 seconds** —\n`\"drain timeout, force scaling down\"` — and the controller force-scaled\ndown anyway.\n\nThe more important finding is underneath that: this setup uses a plain\nKubernetes `Service` (kube-proxy round-robin), not llm-d’s actual EPP\ngateway. The draining label is designed to be read by that EPP to stop\nrouting new requests to the pod — a plain Service has no concept of the\nlabel at all. So the drain mechanism’s *write path* — apply label, poll\nin-flight request count, force-scale on timeout — is confirmed to execute\ncorrectly end to end. Its *actual protective effect* was not validated\nhere, because half the mechanism (routing exclusion) isn’t present in\nthis de-risked setup. Zero failures is good news, but it may simply mean\nnothing happened to be in-flight on that specific pod at the moment it\nwas killed, not that the drain protected it. Real validation of this\nguarantee needs llm-d’s EPP in front of the pods — explicitly outside\nthis round’s scope, not a result I’m claiming.\n\n### Summary\n\n| Test | Result | Requests | Failures | p50 | p95 | \n|---|---|---|---|---|---|\n| TPOT breach → scale decode | **Passed** | 1,970 | 0 | 530ms | 560ms | \n| Velocity spike → scale prefill | **Inconclusive** (real limitation) | 12,517 | 0 | 470ms | 560ms | \n| GPU budget enforcement | **Passed** | 1,376 | 0 | 530ms | 550ms | \n| Drain before scale-down | **Nuanced** (write path confirmed, protection unverified) | 696 | 0 | 470ms | 480ms | \n\n**16,559 total requests across every test. Zero failures, throughout.**\n\n## What This Means\n\nTwo of four mechanisms are cleanly proven: the velocity/SLO detection logic reacts to real signals and makes real scaling decisions, and the joint GPU budget constraint holds even under genuine competing pressure. One is a real, named limitation — this model/hardware pairing couldn’t produce the queue backlog needed to exercise spike detection, through no fault of the detection code itself. One is a genuinely nuanced result that I’d rather report honestly than round up to a pass: the drain mechanism’s bookkeeping works, but proving it protects real traffic needs a piece of infrastructure (llm-d’s EPP) this round deliberately scoped out.\n\nThat’s four real answers, not four passes. I think that’s more useful to anyone deciding whether to actually run this operator than a clean scorecard would have been.\n\n**Open for v2:** a workload that can genuinely saturate prefill admission\n(bigger model, or a resource-starved config), and a real llm-d EPP in\nfront of the pods to validate the drain label’s actual routing-exclusion\neffect, not just its bookkeeping.\n\n## Is This Novel? No — And Here’s Who Already Did It\n\nBefore publishing this, I asked for an honest review of whether any of this is actually new. Short answer: no. This is a small, llm-d-shaped implementation of a pattern that already exists in production elsewhere, and the post should say that plainly instead of letting “I built an operator” imply more than it does.\n\nSpecifically, checked against the primary sources, not just taken on faith:\n\n- **[NVIDIA Dynamo’s Planner](https://docs.nvidia.com/dynamo/v1.2.1/components/planner/planner-guide)** already has a`load` scaling mode that reacts to prefill-queue-token\nand decode-KV-utilization thresholds — structurally the same idea as`AnalysePrefill` /`AnalyseDecode` here. It also already enforces a joint\nGPU budget across prefill and decode (`max_gpu_budget` , a hard cap on\ncombined replicas) — the exact mechanism Test 3 validated above. Not\nnovel. Confirmed by reading the config reference directly.\n- **[HeteroScale](https://arxiv.org/abs/2508.19559)** (ByteDance,\npublished August 2025) runs a single joint autoscaling metric across\nprefill/decode pools in production on tens of thousands of GPUs. The\nentire “pools compete for GPUs, scale them together” framing of this\npost is their result at a scale this validation can’t touch. Not\nnovel.\n- llm-d’s own\n[autoscaling roadmap](https://github.com/llm-d/llm-d-autoscaling/issues/1079) already lists a “rate-based (velocity) scaling signal” as a planned\nitem. The queue-velocity idea in`AnalysePrefill` isn’t an invention,\nit’s llm-d’s own stated direction, arrived at independently and about\na release early.\n- One genuinely useful finding did fall out of checking Dynamo, though:\nits docs state the `load` planner mode is**non-functional** on vLLM,\nSGLang, and TRT-LLM today, because none of those backends expose real\nprefill queue metrics. That’s the same wall Test 2 hit. It means the\ninconclusive result above isn’t a quirk of this 0.6B model on one\nA100 — it’s a structural gap in what current inference engines expose,\nseen independently by a much bigger team. That’s worth more than a\nclean pass would have been.\n\nWhat’s left standing after all that: the drain-before-scale-down mechanism (Test 4) is the one piece I could not find already shipped anywhere. I checked llm-d’s own autoscaling roadmap and Dynamo’s planner guide directly for anything about draining a pod or excluding it from routing before scale-down — neither mentions it. That doesn’t make it novel research; it’s a small, missing piece of plumbing for one specific ecosystem (llm-d), not a new idea. And it’s also the one Test 4 couldn’t actually validate end-to-end, since that requires a real llm-d EPP in front of the pods to prove the exclusion works, not just that the label gets set correctly.\n\nSo: a tiny, real dent, in a very large and already well-populated field. If you’re evaluating this for your own cluster, evaluate it as “a lightweight llm-d-native version of what Dynamo and ByteDance already ship at scale,” not as a new approach to the problem.\n\n## The Scripts\n\nEverything here — the k8s manifests, the Locust load generator, the setup\nand test-runner scripts, the Go fixes — is in\n[gpu-labs](https://github.com/kraghavan/gpu-labs/tree/main/pd-ratio-coordinator-validation),\nalongside the fixed-up [pd-ratio-coordinator](https://github.com/kraghavan/pd-ratio-coordinator/tree/KRAG-1)\nrepo itself. If you find a way to get real queue backlog out of a small\nmodel, or want to wire up a real EPP for the drain test, I’d genuinely\nlike to know.\n\n*Experiments run on Lambda Cloud, 1x A100 (40GB SXM4), k3s, vLLM v0.30.0,\nQwen3-0.6B, pd-ratio-coordinator branch KRAG-1. Platform engineer with\n11+ years in distributed systems going deep on LLM serving\ninfrastructure.*", "url": "https://wpnews.pro/news/pd-ratio-coordinator-does-my-own-tool-actually-work", "canonical_source": "https://kraghavan.ca/llm-infrastructure/inference/2026/10/01/pd-ratio-coordinator-validation.html", "published_at": "2026-10-01 00:00:00+00:00", "updated_at": "2026-10-01 04:47:10.356298+00:00", "lang": "en", "topics": ["ai-infrastructure", "mlops", "large-language-models", "ai-tools"], "entities": ["pd-ratio-coordinator", "llm-d", "Kubernetes", "WVA", "vLLM", "Prometheus", "NIXL", "Locust"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/pd-ratio-coordinator-does-my-own-tool-actually-work", "markdown": "https://wpnews.pro/news/pd-ratio-coordinator-does-my-own-tool-actually-work.md", "text": "https://wpnews.pro/news/pd-ratio-coordinator-does-my-own-tool-actually-work.txt", "jsonld": "https://wpnews.pro/news/pd-ratio-coordinator-does-my-own-tool-actually-work.jsonld"}}