What Breaking and Rebuilding an LLM Serving Platform Taught Me About Reliability A developer rebuilt and stress-tested InferOps, an LLM serving platform, and found that Kubernetes pod-level readiness can report healthy while the serving path has zero ready endpoints. In one experiment, deleting the only serving-runtime pod produced a caller-visible outage of 31.960 seconds, while the first successful completion arrived 63.724 seconds after deletion — a figure that includes the request's own 30.287-second duration. The author concludes operators should ask not just whether a Ready pod exists but whether the serving path currently has a ready backend. After I deleted InferOps’s only serving-runtime pod, Kubernetes produced an observation that initially looked reassuring: a runtime pod was still reporting Ready: True. But that pod was the one being terminated. During the replacement process, the runtime Service had zero ready endpoints. That disagreement became a useful starting point for a broader question: when different layers of an LLM-serving platform describe different realities, which signal should an operator trust? I explored that question through four InferOps V1 investigations: The experiments were different, but they kept exposing the same pattern. A pod can be Ready without being useful to the serving path. A process can be alive while the model is unavailable. More concurrent callers can create more waiting without materially more completed work. And a clean Git checkout is not the same thing as a clean machine. The interesting engineering was in understanding why those signals disagreed. These are bounded observations, not general benchmarks. The four investigations ran on one prepared Windows workstation with CPU inference and one reference model. Their Kubernetes experiments used Docker Desktop Kubernetes, one API replica, and one runtime replica; the fresh-clone journey also included local container inference. Recovery and load requests reached the API pod through a loopback port-forward, so their latency measurements include that path rather than an external ingress or the API Service. The numbers below should not be read as Kubernetes benchmarks, production SLOs, portable capacity figures, or availability guarantees. The goal was narrower: expose behavior, trace it across layers, and make only the claim the executed evidence supports. The first experiment deleted the only serving-runtime pod while the committed load profile was running. Kubernetes replaced it automatically. But “when did the service recover?” turned out to have several possible answers. Figure 1. Recovery measurements use different start and end events. Event boxes are not spaced to scale, and readiness observations are samples rather than a continuous trace. The defined caller-visible outage ran from completion of the first unsuccessful request dispatched after deletion to completion of the first subsequent successful request. That interval was 31.960 seconds. The first successful completion arrived 63.724 seconds after deletion, but that request itself took 30.287 seconds. It was dispatched before the replacement was observed Ready and completed afterward. The measurement does not isolate how much of its duration was waiting and how much was inference. So 63.724 seconds is not a clean measurement of platform recovery time either. It includes the successful request’s end-to-end duration, including waiting and inference. The experiment produced several valid measurements answering different questions. The readiness signals also disagreed During the replacement process, seven readiness samples were collected. In samples from approximately 1.4 to 28.6 seconds after deletion, Kubernetes listed both the terminating old pod and the replacement pod. One runtime pod continued reporting Ready: True, but the runtime Service had zero ready endpoints. The Ready pod was the old pod inside its termination grace period. At the next recorded sample, around 33.7 seconds, only the replacement remained and the Service again had one ready endpoint. This does not establish that pod readiness and endpoint readiness disagreed throughout the entire caller-visible outage. The samples and the caller-outage interval have different boundaries. What it establishes is more specific: during part of the replacement process, pod-level readiness looked healthy while the Service had no endpoint in its ready set. That is enough to change the operational question. Alongside “Is there a Ready pod?” I also want to know, “Does the serving path currently have a ready backend?” Those are related signals, but they are not identical. Evidence: Inference pod recovery — evidence archived in InferOps v1.0.0. Lesson Words such as recovered need an explicit end condition. Recovery might mean: A recovery number without that definition can be numerically correct and operationally misleading. The second experiment tested a different failure state. The runtime container was deliberately constrained to 10 millicores of CPU. llama-server started, bound its port, and began loading the model. It did not finish loading during the registered observation window. The condition was held for 179.755 seconds. The runtime process remained alive and recorded zero restarts in the sampled observations. Its own health endpoint returned: 503 Loading model The API remained live, but its readiness endpoint returned 503. The platform distinguished an alive process from a runtime capable of serving inference. That was expected. The more interesting finding appeared in the error semantics. The runtime knew more than the API adapter could observe The runtime itself could say “Loading model.” But completion requests sent directly to the API pod returned: 503 capability-unavailable condition: runtime-unreachable retryable: true No completion probe received model-not-ready, because runtime readiness had already affected routing. During the unready window, the sampled state was: The API adapter communicates with the runtime through the runtime Service. Because that Service had no endpoint in its ready set, the adapter did not get the opportunity to inspect the runtime’s more precise 503 Loading model response. The readiness gate had answered the question before the adapter could. Error semantics are not determined only by application code. Routing topology can decide which component gets the opportunity to describe the failure. Another boundary: how the experiment reached the API This experiment did not send its requests through the API Service. Because both Services had no ready endpoints, the workflow used loopback port-forwards directly to pods. Figure 2. The probes reached the API and runtime pods directly. The response through the API Service was not measured. The result establishes that when the API pod itself was reached during this condition, its completion endpoint returned capability-unavailable / runtime-unreachable. It does not establish that a caller entering through the API Service would receive that same JSON response. In the observed state, the API Service itself had no ready endpoint. A caller using that Service would have encountered that routing boundary first; the experiment states that condition but did not measure the resulting caller response through the Service. What a caller sees therefore depends partly on where the caller enters the topology. Evidence: Unready-model recovery — evidence archived in InferOps v1.0.0. https://github.com/asadhanif3188/InferOps/blob/718ad2e0fac8ae70c6d053a2decb00e61ff3de39/docs/proof/serving/v1-s4-007-pr1-unready-model-recovery.md Testing an error mapper in isolation is not enough. Readiness probes, Services, proxies, gateways, retries, timeouts, and load balancers can determine whether that mapper ever observes the underlying error. For platform APIs, failure semantics should be tested from the actual entry points users depend on. The third experiment investigated capacity behavior. I wanted to know when additional concurrency stops increasing completed work and starts increasing waiting. The runtime was configured with one serving slot: --parallel 1 --threads 6 --ctx-size 4096 --n-predict 128 --temp 0 The workload used one fixed prompt containing 35 input tokens and produced 17 output tokens per request. The generator was closed loop: each worker waited for its response before sending its next request. After a three-request warm-up, each repetition ran: The whole matrix was repeated twice. Across the two measured runs and warm-up, 366 requests returned HTTP 200 . There were no request timeouts, transport failures, or refusals. Latency rose. Successful completion rate barely moved. Figure 3. The ranges span two repetitions. Processing and deferred counts are sampled gauges; the figures describe one repeated-prompt workload with one serving slot. At concurrency 2, median latency was about 2× baseline. At concurrency 4, it was about 4× baseline. Successful completion rate across all measured phases remained between 0.554 and 0.575 requests per second. Runtime telemetry showed the queue Sampled telemetry was consistent with the single-slot configuration: one request processing, with deferred requests reaching zero, one, and three at concurrency one, two, and four respectively. The API held more requests in flight while the runtime continued processing one at a time. In this configuration, extra caller concurrency primarily created waiting. There is another workload detail worth keeping in view. Runtime prompt-token counters were consistent with substantial prompt-prefix reuse for the repeated fixed prompt. That behavior was inferred from aggregate counters; it was not observed independently per request. This matters because the measured service times describe this particular repeated-prompt workload, not arbitrary prompt shapes. CPU still does not give me a complete causal answer The runtime averaged approximately 99.1–99.9% of its 6-CPU limit across every measured phase, including concurrency one. That is an important signal, but it is not enough to conclude that CPU was definitely the bottleneck. CPU throttling was not sampled, and the experiment did not test alternative serving-slot configurations under the same compute budget. The supported conclusion is narrower: with this one-slot configuration and closed-loop workload, increasing concurrency from one to four increased request waiting while successful completion rate remained roughly flat. Evidence: Performance findings — evidence archived in InferOps v1.0.0 https://github.com/asadhanif3188/InferOps/blob/718ad2e0fac8ae70c6d053a2decb00e61ff3de39/docs/proof/serving/v1-s4-004-pr2-performance-findings.md . Concurrency is not capacity. For inference systems, I care more about where additional demand stops increasing completed useful work than about how many requests the API can hold open. The fourth investigation was the documented InferOps journey executed from new checkouts. The first attempt failed. The second attempt failed. Across those attempts, the workflow exposed five repository defects. After fixes, the third fresh clone completed all 18 checklist steps. Two defects capture why this exercise mattered. Ambient environment changed what a test was testing One test was supposed to validate behavior when no provider had been selected. But the surrounding workflow exported its own provider selection, and a test helper copied that environment into its subprocess. The test was no longer exercising the condition its name claimed it was testing. The development environment had leaked into the test. The fix was to strip inherited INFEROPS variables from that test boundary. Cleanup treated eventual state as immediate state Another failure occurred after Helm uninstall. The workflow saw three release-labelled pods and concluded that residue remained. A moment later, they were gone. The problem was the verification assumption. In this run, helm uninstall --wait had returned, but Deployment-owned pods were still being removed asynchronously by Kubernetes garbage collection. The check asked the right question at the wrong time. It was changed to use a bounded retry instead of treating the first observation as final state. That is a distributed-systems problem in a small form: eventual state should not be verified as though it were instantaneous. The completed run still encountered failure The successful journey took 1 hour 30 minutes 32 seconds, from the first step’s start to the last step’s completion, including pauses between three workflow invocations. It covered prerequisites, tests, workload scaffolding, model acquisition, real local inference, Kubernetes deployment, telemetry, load testing, failure testing, cleanup, and cluster verification. The model transfer stopped twice during that completed journey. The workflow preserved the partial transfer, resumed it using HTTP Range requests, and verified the resulting SHA-256 before continuing. A useful platform workflow should be able to fail without blindly discarding useful state or trusting an incomplete artifact. A fresh clone does not mean a fresh machine The checkout was fresh. The host was not. The machine already contained useful state including package and image-build caches, and the runtime image was already present. Recorded host actions were also required, including selecting a compatible kubectl and identifying the storage drive used by Docker Desktop. Figure 4. The completed journey depended on a prepared host. Cleanup removed scoped resources, verified the cluster, and retained documented caches, images, and records. No second engineer had repeated the journey in the evidence recorded here. This experiment therefore does not establish that InferOps is reproducible anywhere. It establishes something narrower: fresh-clone testing exposed hidden assumptions, the defects were corrected, and the documented path subsequently completed from a clean checkout on that prepared host. The stronger next test is another engineer on another machine. Evidence: Clean-clone run — evidence archived in InferOps v1.0.0 https://github.com/asadhanif3188/InferOps/blob/718ad2e0fac8ae70c6d053a2decb00e61ff3de39/docs/proof/environment/v1-s5-001-pr2-clean-clone-run.md . These investigations crossed different parts of the stack, but they exposed the same engineering habit: do not assume two related signals mean the same thing. Surface observation What the experiment forced me to ask Pod reports Ready Is there actually a backend in the Service ready set? Runtime process is alive Is the model ready to serve inference? Runtime knows it is loading Can that information survive routing and reach the caller? API has four requests in flight How many requests is the runtime actually processing? Fresh checkout exists What state is still coming from the host? Each signal is useful. The mistake is interpreting it outside the boundary it actually measures. That gives me a practical troubleshooting model for AI platforms: Infrastructure → runtime → routing → API → caller And, for operational workflows: Repository → tooling → host state → deployment → verification When observations disagree, I want to identify the boundary between them rather than immediately deciding that one metric is wrong. Often that boundary explains the behavior. Each V1 experiment left behind a stronger next question. For pod recovery, I want to know what changes with multiple replicas and which failures remain visible to callers. For model readiness, I want to test through the actual caller entry point, including the API Service, and decide where model-loading state should be represented without routing traffic to a runtime that cannot yet serve it. For concurrency, the next experiment should vary serving parallelism while holding the compute budget explicit, add CPU-throttling visibility, and test whether additional slots increase completed work or only introduce contention. For reproducibility, the next milestone is straightforward: another engineer, another machine, the documented workflow, and no access to the original developer’s unstated knowledge. The pattern I want to preserve is: Observe → form a hypothesis → design an experiment → measure → bound the conclusion → choose the next experiment. That is more useful to me than simply adding another feature to the platform. The most useful InferOps V1 results were the places where apparently related signals did not mean the same thing. Pod readiness was not equivalent to a ready Service backend. Process liveness was not model readiness. The runtime’s explanation of its own state was not necessarily the error the API could observe—or the result a Service-level caller would encounter. More concurrency was not more completed work. And a fresh checkout was not a fresh machine. That leads to the engineering principle I am taking forward: Reliable AI infrastructure requires connecting what the infrastructure reports, what the runtime is doing, what the serving path permits, what the caller experiences, and what assumptions the operating workflow depends on. The disagreements between those perspectives are often where the architecture becomes visible.