What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You A developer's cost model shows that self-hosting an open-weight LLM is dominated by GPU utilization rather than hardware price, with continuous batching and realistic token mixes moving cost per million tokens by more than 10x. The writeup gives a back-of-the-envelope formula, cost_per_1M_tokens = (gpu_hourly_cost / tokens_per_hour) * 1_000_000, and argues that KV-cache memory, not raw compute, is usually the ceiling on effective concurrency. It advises measuring gpu_cache_usage before buying a bigger GPU and points to num_kv_heads and kv_cache_dtype as the levers that raise that ceiling. Self-hosting an open-weight LLM is rarely expensive because the GPU is expensive. It is expensive because most people run the GPU at single-digit utilization. The headline rental price of an accelerator is fixed per hour, so your true cost per million tokens is set almost entirely by how many tokens you push through that hour. Continuous batching, a realistic input/output token mix, and honest utilization numbers move the price by more than 10x. This post gives you a back-of-the-envelope model you can plug your own numbers into, plus the three server settings that actually change the result. When someone says "an H100 costs about 2 to 3 dollars per hour," they have quoted an input to the problem, not the answer. What you sell or consume is tokens, not GPU-seconds. The conversion factor between them is throughput, and throughput is not a property of the GPU alone. It is a property of the GPU plus the model plus the request pattern plus your server configuration. The single formula that matters: cost per 1M tokens = gpu hourly cost / tokens per hour 1 000 000 tokens per hour = throughput tokens per sec 3600 utilization Two deployments renting the identical GPU at the identical hourly price can differ by an order of magnitude in cost per 1M tokens , purely because one achieves 800 tokens per second at 70 percent utilization and the other achieves 90 tokens per second at 15 percent utilization. The hardware invoice is the same. The unit economics are not. A naive server processes one request, waits for it to finish generating, then starts the next. The GPU spends most of its time idle between the compute-heavy prefill and the memory-bound decode steps. Continuous batching, popularized by the vLLM project, instead keeps a running set of in-flight sequences and injects new requests into the batch as soon as slots free up, token by token. The vLLM docs describe this as iteration-level scheduling, and it is the main reason a well-tuned open server reaches throughput that feels implausible if you have only ever benchmarked batch size 1. Here is the intuition in a tiny simulator. It is not a GPU model, it just shows why aggregate throughput climbs with concurrency until something saturates: python def tokens per hour single stream tps, max concurrency, scaling efficiency : scaling efficiency < 1 captures memory-bandwidth and KV-cache limits effective = single stream tps max concurrency scaling efficiency return effective 3600 for c in 1, 8, 32, 64 : tph = tokens per hour 42, c, scaling efficiency=0.55 cost = 2.5 / tph 1 000 000 $2.5/hr GPU, illustrative print f"concurrency={c: 3} tok/hr={tph/1e6:5.2f}M $/1M={cost:6.2f}" The numbers above are illustrative, not measured, but the shape is real: moving from one in-flight request to a few dozen can turn a 20-dollar price per million tokens into something close to 1 dollar. The cover chart shows the same relationship. The curve flattens once you hit a bottleneck, which is usually KV-cache memory rather than raw compute. The quiet ceiling is the KV cache. Every active sequence stores keys and values for all previous tokens, and that memory grows with context length and batch size. When the cache fills, the scheduler either queues new requests or preempts running ones, and your effective concurrency stops climbing. A rough KV-cache size estimate for a transformer: kv bytes per token = 2 num layers num kv heads head dim bytes per elem factor of 2 is for keys AND values Two levers shrink this and therefore raise the concurrency ceiling: num kv heads , which is why many recent open models adopt it. bytes per elem . vLLM exposes this through kv cache dtype . If your p99 latency is fine but throughput plateaus early, you are almost certainly cache-bound, not compute-bound. Measure gpu cache usage in the server metrics before you reach for a bigger GPU. Weight quantization for example 4-bit GGUF for llama.cpp, or AWQ and GPTQ for GPU serving does two separate things, and people conflate them. What it does not reliably do is cut cost on a GPU you were already underutilizing. If you serve one user at a time, a 4-bit model on an idle GPU still produces an embarrassing price per million tokens, because the denominator tokens per hour is still tiny. Quantization pays off when you combine it with batching so the freed memory becomes extra concurrency. GGUF on CPU is a related but different trade. llama.cpp can serve quantized models with no GPU at all, which is attractive for low-traffic internal tools where a GPU would sit idle. The crossover point is traffic volume: below some requests-per-minute threshold, a CPU box you already own beats a rented GPU you barely use. Above it, the GPU wins decisively because its tokens-per-hour ceiling is so much higher. You need four inputs, three of which you can measure and one you must assume: gpu hourly = 2.5 your actual rental or amortized price measured tps = 650 steady-state tokens/sec under realistic load utilization = 0.60 fraction of the hour the GPU is actually busy overhead = 1.15 networking, retries, idle head start, replicas tph = measured tps 3600 utilization cost per 1M = gpu hourly / tph 1 000 000 overhead print round cost per 1M, 2 , "USD per 1M tokens" The one input people fake is measured tps . Do not copy a vendor benchmark run at batch 256 with 128-token outputs if your workload is batch 4 with 1,000-token outputs. Output-heavy workloads spend far more time in the bandwidth-bound decode phase, so their tokens-per-hour is lower and their cost is higher. Benchmark with your own input/output length distribution or the number is fiction. After tuning many open-model deployments, the settings with the largest effect on cost per token are consistently these: max num seqs or the equivalent max concurrency . Set it too low and you leave throughput on the table. Set it too high and you thrash the KV cache and trigger preemption. Sweep it against your real traffic. gpu memory utilization . vLLM pre-allocates the KV cache from this fraction. Nudging it from 0.80 to 0.90 can meaningfully raise the concurrency ceiling, as long as you keep headroom for activation spikes. Everything else speculative decoding, chunked prefill, tensor parallel degree matters, but these three are where the first large wins live. Is self-hosting cheaper than a hosted API? It depends entirely on utilization. A hosted API amortizes one GPU across thousands of tenants, so it runs at high utilization you would struggle to match. Self-hosting wins on cost only when your sustained traffic keeps your own GPU busy, or when data residency and latency, not price, are the reason. Why is my cost per token so much higher than the blog posts I read? Almost always because those posts report peak throughput at large batch sizes with short outputs, while your workload has long outputs and modest concurrency. Re-run the benchmark with your own token distribution. Does a bigger GPU always lower cost per token? No. A bigger GPU only helps if you can fill it. If you are already at low utilization, a larger accelerator raises your hourly cost while the tokens-per-hour barely moves, making the unit cost worse. Where do continuous batching and quantization overlap? Quantization frees memory, continuous batching turns that freed memory into extra concurrent sequences. Used together they compound. Used alone on an idle GPU, neither fixes the underlying utilization problem. How do I know if I am compute-bound or memory-bound? Watch KV-cache usage and GPU compute utilization together. If cache usage hits the ceiling while compute sits below 100 percent, you are memory-bound and should attack the cache GQA models, FP8 cache, shorter contexts before buying compute. What is a realistic utilization target? For bursty interactive traffic, sustained 50 to 70 percent is a reasonable and honest target. Claiming 95 percent usually means you are either batching offline jobs or quoting a benchmark, not a production SLA. All throughput and cost figures in this post are labeled illustrative and are meant to show the shape of the relationships. Replace them with your own measured values before making a budget decision.