# What Does It Actually Cost to Self-Host an LLM? The Batching Math Nobody Shows You

> Source: <https://dev.to/wolfsea2357/what-does-it-actually-cost-to-self-host-an-llm-the-batching-math-nobody-shows-you-4pmp>
> Published: 2026-10-06 00:43:30+00:00

Self-hosting an open-weight LLM is rarely expensive because the GPU is expensive. It is expensive because most people run the GPU at single-digit utilization. The headline rental price of an accelerator is fixed per hour, so your true cost per million tokens is set almost entirely by how many tokens you push through that hour. Continuous batching, a realistic input/output token mix, and honest utilization numbers move the price by more than 10x. This post gives you a back-of-the-envelope model you can plug your own numbers into, plus the three server settings that actually change the result.

When someone says "an H100 costs about 2 to 3 dollars per hour," they have quoted an input to the problem, not the answer. What you sell or consume is tokens, not GPU-seconds. The conversion factor between them is throughput, and throughput is not a property of the GPU alone. It is a property of the GPU plus the model plus the request pattern plus your server configuration.

The single formula that matters:

```
cost_per_1M_tokens = (gpu_hourly_cost / tokens_per_hour) * 1_000_000
tokens_per_hour    = throughput_tokens_per_sec * 3600 * utilization
```

Two deployments renting the identical GPU at the identical hourly price can differ by an order of magnitude in `cost_per_1M_tokens`, purely because one achieves 800 tokens per second at 70 percent utilization and the other achieves 90 tokens per second at 15 percent utilization. The hardware invoice is the same. The unit economics are not.

A naive server processes one request, waits for it to finish generating, then starts the next. The GPU spends most of its time idle between the compute-heavy prefill and the memory-bound decode steps. Continuous batching, popularized by the vLLM project, instead keeps a running set of in-flight sequences and injects new requests into the batch as soon as slots free up, token by token. The vLLM docs describe this as iteration-level scheduling, and it is the main reason a well-tuned open server reaches throughput that feels implausible if you have only ever benchmarked batch size 1.

Here is the intuition in a tiny simulator. It is not a GPU model, it just shows why aggregate throughput climbs with concurrency until something saturates:

``` python
def tokens_per_hour(single_stream_tps, max_concurrency, scaling_efficiency):
    # scaling_efficiency < 1 captures memory-bandwidth and KV-cache limits
    effective = single_stream_tps * max_concurrency * scaling_efficiency
    return effective * 3600

for c in (1, 8, 32, 64):
    tph = tokens_per_hour(42, c, scaling_efficiency=0.55)
    cost = (2.5 / tph) * 1_000_000   # $2.5/hr GPU, illustrative
    print(f"concurrency={c:>3}  tok/hr={tph/1e6:5.2f}M  $/1M={cost:6.2f}")
```

The numbers above are illustrative, not measured, but the shape is real: moving from one in-flight request to a few dozen can turn a 20-dollar price per million tokens into something close to 1 dollar. The cover chart shows the same relationship. The curve flattens once you hit a bottleneck, which is usually KV-cache memory rather than raw compute.

The quiet ceiling is the KV cache. Every active sequence stores keys and values for all previous tokens, and that memory grows with context length and batch size. When the cache fills, the scheduler either queues new requests or preempts running ones, and your effective concurrency stops climbing.

A rough KV-cache size estimate for a transformer:

```
kv_bytes_per_token = 2 * num_layers * num_kv_heads * head_dim * bytes_per_elem
# factor of 2 is for keys AND values
```

Two levers shrink this and therefore raise the concurrency ceiling:

`num_kv_heads`, which is why many recent open models adopt it.` bytes_per_elem`. vLLM exposes this through `kv_cache_dtype`.
If your p99 latency is fine but throughput plateaus early, you are almost certainly cache-bound, not compute-bound. Measure `gpu_cache_usage` in the server metrics before you reach for a bigger GPU.

Weight quantization (for example 4-bit GGUF for llama.cpp, or AWQ and GPTQ for GPU serving) does two separate things, and people conflate them.

What it does not reliably do is cut cost on a GPU you were already underutilizing. If you serve one user at a time, a 4-bit model on an idle GPU still produces an embarrassing price per million tokens, because the denominator (tokens per hour) is still tiny. Quantization pays off when you combine it with batching so the freed memory becomes extra concurrency.

GGUF on CPU is a related but different trade. llama.cpp can serve quantized models with no GPU at all, which is attractive for low-traffic internal tools where a GPU would sit idle. The crossover point is traffic volume: below some requests-per-minute threshold, a CPU box you already own beats a rented GPU you barely use. Above it, the GPU wins decisively because its tokens-per-hour ceiling is so much higher.

You need four inputs, three of which you can measure and one you must assume:

```
gpu_hourly = 2.5          # your actual rental or amortized price
measured_tps = 650        # steady-state tokens/sec under realistic load
utilization = 0.60        # fraction of the hour the GPU is actually busy
overhead = 1.15           # networking, retries, idle head start, replicas

tph = measured_tps * 3600 * utilization
cost_per_1M = (gpu_hourly / tph) * 1_000_000 * overhead
print(round(cost_per_1M, 2), "USD per 1M tokens")
```

The one input people fake is `measured_tps`. Do not copy a vendor benchmark run at batch 256 with 128-token outputs if your workload is batch 4 with 1,000-token outputs. Output-heavy workloads spend far more time in the bandwidth-bound decode phase, so their tokens-per-hour is lower and their cost is higher. Benchmark with your own input/output length distribution or the number is fiction.

After tuning many open-model deployments, the settings with the largest effect on cost per token are consistently these:

`max_num_seqs` (or the equivalent max concurrency). Set it too low and you leave throughput on the table. Set it too high and you thrash the KV cache and trigger preemption. Sweep it against your real traffic.`gpu_memory_utilization`. vLLM pre-allocates the KV cache from this fraction. Nudging it from 0.80 to 0.90 can meaningfully raise the concurrency ceiling, as long as you keep headroom for activation spikes.
Everything else (speculative decoding, chunked prefill, tensor parallel degree) matters, but these three are where the first large wins live.

**Is self-hosting cheaper than a hosted API?**

It depends entirely on utilization. A hosted API amortizes one GPU across thousands of tenants, so it runs at high utilization you would struggle to match. Self-hosting wins on cost only when your sustained traffic keeps your own GPU busy, or when data residency and latency, not price, are the reason.

**Why is my cost per token so much higher than the blog posts I read?**

Almost always because those posts report peak throughput at large batch sizes with short outputs, while your workload has long outputs and modest concurrency. Re-run the benchmark with your own token distribution.

**Does a bigger GPU always lower cost per token?**

No. A bigger GPU only helps if you can fill it. If you are already at low utilization, a larger accelerator raises your hourly cost while the tokens-per-hour barely moves, making the unit cost worse.

**Where do continuous batching and quantization overlap?**

Quantization frees memory, continuous batching turns that freed memory into extra concurrent sequences. Used together they compound. Used alone on an idle GPU, neither fixes the underlying utilization problem.

**How do I know if I am compute-bound or memory-bound?**

Watch KV-cache usage and GPU compute utilization together. If cache usage hits the ceiling while compute sits below 100 percent, you are memory-bound and should attack the cache (GQA models, FP8 cache, shorter contexts) before buying compute.

**What is a realistic utilization target?**

For bursty interactive traffic, sustained 50 to 70 percent is a reasonable and honest target. Claiming 95 percent usually means you are either batching offline jobs or quoting a benchmark, not a production SLA.

All throughput and cost figures in this post are labeled illustrative and are meant to show the shape of the relationships. Replace them with your own measured values before making a budget decision.
