vLLM, KServe and llm-d: serving language models on Kubernetes A Kubedex guide outlines how vLLM, KServe and llm-d combine to serve language models on Kubernetes, recommending teams start with a single vLLM deployment and add KServe for a shared model-serving API and llm-d for distributed routing only when they can observe the bottleneck each solves. Kubedex states it has not run a GPU benchmark or production rollout for the guide, and advises pinning Kubernetes, GPU driver, runtime, engine, model-serving controller, gateway and inference-extension versions together with exact image digests, chart versions and custom-resource API versions rather than upgrading every component independently. To put a language model behind a chat or text-generation API, you need software that loads the model, accepts prompts and generates answers. You also need enough GPU memory, a way to send requests to available workers and a process for updating the model without breaking the service. Kubernetes runs and replaces those workers, but it does not supply the whole AI-serving system. vLLM is an inference engine: it runs the model and generates output tokens, the pieces of text that make up an answer. KServe manages model-serving deployments in Kubernetes, giving teams a consistent way to declare and run services. llm-d supplies components for distributed serving, including routing requests across model workers. They can work together; this is not a vLLM-versus-KServe replacement decision. Start with vLLM and a small deployment when you need to serve one model and learn its memory and response-time requirements. Add KServe when several teams or models need a common deployment interface. Evaluate llm-d when multiple workers need more sophisticated routing or you have measured a reason to split model-processing stages across them. Each additional component should solve a problem you can observe. This guide explains those decisions and provides a proposed measurement exercise. Kubedex has not run a GPU benchmark or production rollout for it. The sections below show what to measure before claiming a latency, capacity or cost improvement. What to add around the model server A direct vLLM deployment can be a useful starting point for one model with a clear owner. You still need routing, authentication, health checks, artifact access and capacity management. The vLLM Production Stack https://docs.vllm.ai/projects/production-stack/en/latest/deployment/ offers deployment and routing components around the engine; adopting it is a platform choice with its own configuration and upgrade responsibilities. KServe is useful when teams need a consistent model-serving API and lifecycle. Its administrator guide https://kserve.github.io/website/docs/admin-guide/overview distinguishes InferenceService from LLMInferenceService and separates Standard Kubernetes deployments from Knative-based serving. A standard LLM deployment does not automatically need the most elaborate distributed serving path. llm-d's architecture https://llm-d.ai/docs/architecture groups model-server Pods in an InferencePool and uses a router with an endpoint picker to select a destination. Prefix-cache routing and prefill/decode disaggregation address different constraints. Adopt either only after you can identify the bottleneck it is intended to reduce. Pin a compatible release set Record Kubernetes, GPU driver, runtime, engine, model-serving controller, gateway and inference-extension versions together. Include the exact image digests, chart versions and custom-resource API versions. A manifest from a development documentation branch may not match the installed CRD. Use the installation and dependency set for the selected KServe release or llm-d deployment path. Do not independently upgrade every component to its latest release. For example, current KServe and llm-d documentation describes different autoscaling arrangements; that is a reason to check integration versions, not to combine their examples. Consult the KServe installation overview https://kserve.github.io/website/docs/install/overview and dependency matrix https://kserve.github.io/website/docs/install/dependencies . Optional local model caching, network providers and multi-node controllers each add operating responsibilities. Describe your prompts and traffic before sizing GPUs Create a benchmark record with the following fields: - Artifact: model and tokenizer revisions, quantization, precision, weight source and access terms. - Workload: prompt-length and output-length distributions, concurrency or arrival rate, streaming behavior, cancellations and repeated prefixes. - Hardware: accelerator model and memory, device count, interconnect, CPU, RAM and storage path. - Configuration: context limit, parallelism, engine arguments, cache settings, router policy and replica count. - Acceptance: latency percentiles, error budget, output-quality checks and a cost measurement interval. A short-prompt throughput test cannot establish capacity for long-context interactive traffic. A repeated system prompt may also make a warm cache look unusually effective. Keep representative and deliberately uncached requests as separate test cases. Measure the whole request, including queueing Measure client-observed time to first token and token-delivery latency as well as completed requests, errors and output-token throughput. Connect those observations to engine queue time, running and waiting requests, cache pressure and GPU memory. The vLLM metric reference https://docs.vllm.ai/en/latest/usage/metrics/ includes separate queue, prefill, inference and first-token measurements; confirm the names exposed by your pinned release. Use the matching vLLM serving benchmark https://docs.vllm.ai/en/latest/cli/bench/serve/ or a controlled client harness. Record arrival patterns, warm-up, run duration and failed requests alongside percentiles. Compare two configurations with the same request population and artifact identity. Count successful work meeting the chosen latency objective, not just all generated tokens. Keep three latency measurements separate. Time to first token TTFT measures the wait for the first output. Time per output token TPOT averages the remaining generation time within each request. Inter-token latency ITL measures gaps between successive streamed outputs; an output can contain several tokens. Their percentile distributions therefore need not match. Also distinguish client measurements from server histograms: the client sees the network and gateway path. The benchmark definitions https://docs.vllm.ai/en/stable/benchmarking/cli/ and metrics design https://docs.vllm.ai/en/stable/design/metrics/ describe the differences. An acceptance record should name the metric and measurement point: model revision, hardware, request set, client p95 TTFT, per-request p95 TPOT, errors and completed tokens over the declared interval. Fill it with observed values. There is no hardware-independent latency number that this article can supply. Start with a bounded diagnostic run This unexecuted command template assumes an isolated, already-running text-completions endpoint available on localhost, a compatible installed vLLM benchmark client, and a local tokenizer matching the served model revision. Replace both placeholders and check vllm bench serve --help for your pinned client. It exercises a small synthetic measurement batch; it is not a model-quality evaluation or a capacity test. served model=your-served-model-name tokenizer path=/absolute/path/to/pinned-tokenizer vllm bench serve \ --backend openai \ --base-url http://127.0.0.1:8000 \ --endpoint /v1/completions \ --model "$served model" \ --tokenizer "$tokenizer path" \ --dataset-name random \ --random-input-len 128 \ --random-output-len 64 \ --num-prompts 20 \ --request-rate 1 \ --max-concurrency 2 \ --percentile-metrics ttft,tpot,itl,e2el \ --metric-percentiles 50,95,99 \ --save-result \ --result-dir ./benchmark-results Inspect failures and actual token counts in the saved result before increasing load. Early stopping can change output length. The concurrency cap can reduce the achieved arrival rate, so record the measured rate too. Twenty requests can expose a configuration problem; they cannot establish reliable production tail latency. Follow with the representative workload and longer observation period defined above. The CLI reference https://docs.vllm.ai/en/latest/cli/bench/serve/ documents the options and distinguishes text-completions from chat endpoints. Measure how long a new worker takes to start Adding a replica may require node provisioning, image transfer, model download, volume attachment, engine initialization and warm-up before it can serve. Measure each stage. Test with an empty node cache as well as a warm restart; a model cached on one node is not necessarily available on its replacement. Choose storage behavior deliberately. A shared filesystem introduces throughput and availability dependencies. Local caching adds placement, eviction and prewarming decisions. Object storage requires identity, network access and a measured download path. Readiness must mean the selected model can serve a representative request, not simply that a process is listening. Keep warm capacity where an interactive latency target cannot absorb the measured cold path. For batch traffic, a longer wait may be acceptable. KServe's Knative autoscaling guide https://kserve.github.io/website/docs/model-serving/predictive-inference/autoscaling/kpa-autoscaler describes a particular deployment mode; its scale-to-zero behavior is not a promise for every KServe API, mode or LLM configuration. Add intelligent routing when the workload benefits With several identical replicas, ordinary load balancing can send a request to a worker without its useful cached prefix. vLLM prefix caching https://docs.vllm.ai/en/latest/features/automatic prefix caching/ reuses eligible work within the engine. A cache-aware router adds the separate decision of which worker should receive the request. Measure cache reuse and queue delay together: affinity to an overloaded worker can offset saved prefill work. Prefill processes the input prompt; decode generates the answer. Separating those stages across workers is called prefill/decode disaggregation. Include transfer of the model's cached intermediate state the KV cache and network behavior in the experiment. Compare against a simpler colocated baseline on the same workload and resource budget. More specialized workers are justified only if the measured benefit survives transfer overhead, failure handling and operational cost. Keep autoscaling and admission ownership clear Use a load signal that matches the bottleneck, such as outstanding work or measured queue pressure, and verify its behavior during bursts and metrics outages. CPU utilization alone can miss a saturated GPU-serving path. Choose one owner of replica count and check which controller creates the HPA or other scaling resources for the chosen release. Replica demand cannot create unavailable GPUs. Coordinate serving scale-up with node capacity, quotas, device allocation and the measured cold-start delay. Set finite queue and request limits so excess work has a defined outcome. During scale-down, drain in-flight streams and observe cancellations and retries; a drop in replica count is not evidence that users completed their requests. GPU scheduling, DRA and Kueue https://kubedex.com/kubernetes-gpu-scheduling/ cover workload admission and device placement. Kueue is not the token-by-token router for HTTP requests. Batch admission policy and an interactive request queue need separate objectives even when they share a GPU fleet. Roll out model state and protect the endpoint Keep the previous model revision, serving configuration and routing declaration available. Validate output quality with a fixed evaluation set before shifting a small request cohort. Record a rollback condition for both service behavior and model quality. HTTP API compatibility does not establish equivalent answers. Test old and new versions under the actual routing policy, including cache warm-up and any session assumptions. Restoring the router to an old endpoint works only while that endpoint remains healthy and has capacity. Include gateway timeouts and streaming behavior in the check. Apply authentication and tenant limits at the intended public boundary; restrict internal engine, metrics and control endpoints. Model downloads and runtime options can introduce code and credential risks, so use trusted artifacts and the vLLM security guidance https://docs.vllm.ai/en/latest/usage/security/ . Keep prompts, credentials and model-access tokens out of routine debug logs. Define the evidence needed to go live A useful acceptance run includes a representative warm load, a cold scale-up, a rejected or timed-out request, one worker loss and a model-version reversal. Report what was actually exercised, with versions and raw observations. Calculate cost over that same interval, including idle capacity and failed work. This produces a defensible baseline for your application; it does not establish a universal ranking of vLLM, KServe, llm-d or GPU models. Responsibilities in an inference deployment Choose columns | Responsibilities in an inference deployment | | | |---|---|---| | Layer | Primary decision | Evidence to collect | |---|---|---| | vLLM engine | Model execution, memory and batching configuration | Quality, engine latency, queueing and memory for a declared workload | | KServe control plane | Serving API, lifecycle and deployment mode | Reconciliation, readiness, dependency compatibility and recovery | | llm-d request path | Inference-aware routing and optional distributed patterns | Cache reuse, transfer cost, routing latency and failure behavior | | Kubernetes and device components | Capacity, placement and device allocation | Admission wait, scheduling delay, cold starts and device health | 4 rows Sources & further reading Spotted something that needs another look? Help improve this page → https://kubedex.com/contact/