{"slug": "what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows", "title": "What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You", "summary": "A developer's cost model shows that self-hosting an open-weight LLM is dominated by GPU utilization rather than hardware price, with continuous batching and realistic token mixes moving cost per million tokens by more than 10x. The writeup gives a back-of-the-envelope formula, cost_per_1M_tokens = (gpu_hourly_cost / tokens_per_hour) * 1_000_000, and argues that KV-cache memory, not raw compute, is usually the ceiling on effective concurrency. It advises measuring gpu_cache_usage before buying a bigger GPU and points to num_kv_heads and kv_cache_dtype as the levers that raise that ceiling.", "body_md": "Self-hosting an open-weight LLM is rarely expensive because the GPU is expensive. It is expensive because most people run the GPU at single-digit utilization. The headline rental price of an accelerator is fixed per hour, so your true cost per million tokens is set almost entirely by how many tokens you push through that hour. Continuous batching, a realistic input/output token mix, and honest utilization numbers move the price by more than 10x. This post gives you a back-of-the-envelope model you can plug your own numbers into, plus the three server settings that actually change the result.\n\nWhen someone says \"an H100 costs about 2 to 3 dollars per hour,\" they have quoted an input to the problem, not the answer. What you sell or consume is tokens, not GPU-seconds. The conversion factor between them is throughput, and throughput is not a property of the GPU alone. It is a property of the GPU plus the model plus the request pattern plus your server configuration.\n\nThe single formula that matters:\n\n```\ncost_per_1M_tokens = (gpu_hourly_cost / tokens_per_hour) * 1_000_000\ntokens_per_hour    = throughput_tokens_per_sec * 3600 * utilization\n```\n\nTwo deployments renting the identical GPU at the identical hourly price can differ by an order of magnitude in `cost_per_1M_tokens`, purely because one achieves 800 tokens per second at 70 percent utilization and the other achieves 90 tokens per second at 15 percent utilization. The hardware invoice is the same. The unit economics are not.\n\nA naive server processes one request, waits for it to finish generating, then starts the next. The GPU spends most of its time idle between the compute-heavy prefill and the memory-bound decode steps. Continuous batching, popularized by the vLLM project, instead keeps a running set of in-flight sequences and injects new requests into the batch as soon as slots free up, token by token. The vLLM docs describe this as iteration-level scheduling, and it is the main reason a well-tuned open server reaches throughput that feels implausible if you have only ever benchmarked batch size 1.\n\nHere is the intuition in a tiny simulator. It is not a GPU model, it just shows why aggregate throughput climbs with concurrency until something saturates:\n\n``` python\ndef tokens_per_hour(single_stream_tps, max_concurrency, scaling_efficiency):\n    # scaling_efficiency < 1 captures memory-bandwidth and KV-cache limits\n    effective = single_stream_tps * max_concurrency * scaling_efficiency\n    return effective * 3600\n\nfor c in (1, 8, 32, 64):\n    tph = tokens_per_hour(42, c, scaling_efficiency=0.55)\n    cost = (2.5 / tph) * 1_000_000   # $2.5/hr GPU, illustrative\n    print(f\"concurrency={c:>3}  tok/hr={tph/1e6:5.2f}M  $/1M={cost:6.2f}\")\n```\n\nThe numbers above are illustrative, not measured, but the shape is real: moving from one in-flight request to a few dozen can turn a 20-dollar price per million tokens into something close to 1 dollar. The cover chart shows the same relationship. The curve flattens once you hit a bottleneck, which is usually KV-cache memory rather than raw compute.\n\nThe quiet ceiling is the KV cache. Every active sequence stores keys and values for all previous tokens, and that memory grows with context length and batch size. When the cache fills, the scheduler either queues new requests or preempts running ones, and your effective concurrency stops climbing.\n\nA rough KV-cache size estimate for a transformer:\n\n```\nkv_bytes_per_token = 2 * num_layers * num_kv_heads * head_dim * bytes_per_elem\n# factor of 2 is for keys AND values\n```\n\nTwo levers shrink this and therefore raise the concurrency ceiling:\n\n`num_kv_heads`, which is why many recent open models adopt it.` bytes_per_elem`. vLLM exposes this through `kv_cache_dtype`.\nIf your p99 latency is fine but throughput plateaus early, you are almost certainly cache-bound, not compute-bound. Measure `gpu_cache_usage` in the server metrics before you reach for a bigger GPU.\n\nWeight quantization (for example 4-bit GGUF for llama.cpp, or AWQ and GPTQ for GPU serving) does two separate things, and people conflate them.\n\nWhat it does not reliably do is cut cost on a GPU you were already underutilizing. If you serve one user at a time, a 4-bit model on an idle GPU still produces an embarrassing price per million tokens, because the denominator (tokens per hour) is still tiny. Quantization pays off when you combine it with batching so the freed memory becomes extra concurrency.\n\nGGUF on CPU is a related but different trade. llama.cpp can serve quantized models with no GPU at all, which is attractive for low-traffic internal tools where a GPU would sit idle. The crossover point is traffic volume: below some requests-per-minute threshold, a CPU box you already own beats a rented GPU you barely use. Above it, the GPU wins decisively because its tokens-per-hour ceiling is so much higher.\n\nYou need four inputs, three of which you can measure and one you must assume:\n\n```\ngpu_hourly = 2.5          # your actual rental or amortized price\nmeasured_tps = 650        # steady-state tokens/sec under realistic load\nutilization = 0.60        # fraction of the hour the GPU is actually busy\noverhead = 1.15           # networking, retries, idle head start, replicas\n\ntph = measured_tps * 3600 * utilization\ncost_per_1M = (gpu_hourly / tph) * 1_000_000 * overhead\nprint(round(cost_per_1M, 2), \"USD per 1M tokens\")\n```\n\nThe one input people fake is `measured_tps`. Do not copy a vendor benchmark run at batch 256 with 128-token outputs if your workload is batch 4 with 1,000-token outputs. Output-heavy workloads spend far more time in the bandwidth-bound decode phase, so their tokens-per-hour is lower and their cost is higher. Benchmark with your own input/output length distribution or the number is fiction.\n\nAfter tuning many open-model deployments, the settings with the largest effect on cost per token are consistently these:\n\n`max_num_seqs` (or the equivalent max concurrency). Set it too low and you leave throughput on the table. Set it too high and you thrash the KV cache and trigger preemption. Sweep it against your real traffic.`gpu_memory_utilization`. vLLM pre-allocates the KV cache from this fraction. Nudging it from 0.80 to 0.90 can meaningfully raise the concurrency ceiling, as long as you keep headroom for activation spikes.\nEverything else (speculative decoding, chunked prefill, tensor parallel degree) matters, but these three are where the first large wins live.\n\n**Is self-hosting cheaper than a hosted API?**\n\nIt depends entirely on utilization. A hosted API amortizes one GPU across thousands of tenants, so it runs at high utilization you would struggle to match. Self-hosting wins on cost only when your sustained traffic keeps your own GPU busy, or when data residency and latency, not price, are the reason.\n\n**Why is my cost per token so much higher than the blog posts I read?**\n\nAlmost always because those posts report peak throughput at large batch sizes with short outputs, while your workload has long outputs and modest concurrency. Re-run the benchmark with your own token distribution.\n\n**Does a bigger GPU always lower cost per token?**\n\nNo. A bigger GPU only helps if you can fill it. If you are already at low utilization, a larger accelerator raises your hourly cost while the tokens-per-hour barely moves, making the unit cost worse.\n\n**Where do continuous batching and quantization overlap?**\n\nQuantization frees memory, continuous batching turns that freed memory into extra concurrent sequences. Used together they compound. Used alone on an idle GPU, neither fixes the underlying utilization problem.\n\n**How do I know if I am compute-bound or memory-bound?**\n\nWatch KV-cache usage and GPU compute utilization together. If cache usage hits the ceiling while compute sits below 100 percent, you are memory-bound and should attack the cache (GQA models, FP8 cache, shorter contexts) before buying compute.\n\n**What is a realistic utilization target?**\n\nFor bursty interactive traffic, sustained 50 to 70 percent is a reasonable and honest target. Claiming 95 percent usually means you are either batching offline jobs or quoting a benchmark, not a production SLA.\n\nAll throughput and cost figures in this post are labeled illustrative and are meant to show the shape of the relationships. Replace them with your own measured values before making a budget decision.", "url": "https://wpnews.pro/news/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows", "canonical_source": "https://dev.to/wolfsea2357/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows-you-4pmp", "published_at": "2026-10-06 00:43:30+00:00", "updated_at": "2026-10-06 00:47:32.213869+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools"], "entities": ["vLLM", "H100"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows", "markdown": "https://wpnews.pro/news/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows.md", "text": "https://wpnews.pro/news/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows.txt", "jsonld": "https://wpnews.pro/news/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows.jsonld"}}