# KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly

> Source: <https://dev.to/wolfsea2357/kv-cache-quantization-in-llm-serving-fp8-and-int8-tradeoffs-the-silent-configjson-trap-and-how-5acc>
> Published: 2026-10-06 19:10:48+00:00

The KV cache is often the real memory ceiling in LLM serving, not the weights. Quantizing it to FP8 or INT8 roughly halves KV memory, which lets you serve longer contexts or more concurrent requests on the same GPU. The catch: the quality impact is real but small when measured correctly, and it is easy to measure it wrong. A single `kv_cache_dtype` flag, or a scaling factor baked into a model's `config.json`, can quietly change your outputs while every dashboard still shows green. This post covers where the savings come from, what actually degrades, the silent-config trap, and a fair measurement protocol.

When a transformer generates tokens autoregressively, it stores the key and value tensors for every past token so it does not recompute them. That store is the KV cache, and it grows linearly with sequence length and batch size.

A rough per-token KV size:

```
bytes_per_token = 2 (K and V)
               * num_layers
               * num_kv_heads
               * head_dim
               * dtype_bytes
```

For a mid-size model with 32 layers, 8 KV heads (grouped-query attention), and head_dim 128, in FP16 that is:

```
2 * 32 * 8 * 128 * 2 = 131,072 bytes  (128 KB per token)
```

At 8K context that is 1 GB per sequence, before you batch anything. Weights are a fixed cost you pay once; the KV cache is a per-request cost that scales with your traffic. That is why it is the highest-leverage place to spend a quantization budget. Drop the dtype from 2 bytes to 1 byte and you halve that 128 KB to 64 KB per token, which directly buys you longer contexts or a bigger batch.

Both formats store each cached value in 8 bits, so the memory savings are effectively identical, about 2x versus FP16/BF16. The difference is in how they represent numbers and what that does to accuracy.

**FP8** (typically `e4m3` or `e5m2`) keeps an exponent field, so it covers a wide dynamic range with graceful precision falloff. The vLLM docs note that FP8 KV cache is supported on CUDA 11.8 and later and on ROCm, and that `e4m3` is the common default on NVIDIA hardware. Because attention keys and values can have occasional large-magnitude outliers, the exponent bits help: an outlier does not blow out the whole tensor's scale.

**INT8** uses a uniform grid with a per-tensor or per-channel scale. When the distribution is well behaved, INT8 can actually be more precise than FP8 inside its range because it does not spend bits on an exponent. The risk is outliers: one large value forces a coarse scale on everything else. INT8 KV usually wants calibration to pick good scales.

Practical reading:

Less than people fear, if you measure the right thing. KV quantization only touches the cache, not the weights or the activations on the compute path, so the error is confined to attention's read of past tokens. In practice:

The headline number ("FP8 KV costs X percent") is meaningless without naming the task, the context length, and whether you measured a distribution or a single greedy run.

Here is the failure that burns teams. You benchmark model A and model B with what looks like the identical serving command, same flags, same prompts, and you get different accuracy. The flags were identical. The conditions were not.

Two things commonly hide inside a model directory rather than on your command line:

**A checkpoint-level KV cache dtype.** Some published checkpoints ship a quantization config in `config.json` (for example a `quantization_config` or a KV-specific field) that sets the cache to FP8 even when your launch command says nothing about it. Your serving engine reads it and honors it. Your benchmark harness never sees it.

**Baked-in KV scaling factors.** FP8 KV can use static per-layer scales stored in the checkpoint. If one checkpoint has calibrated scales and another falls back to default scales, their attention precision differs even though both report "fp8 KV" in the logs. vLLM exposes `--calculate-kv-scales` to compute scales at runtime for the `e4m3` path precisely so you are not silently depending on whatever shipped in the file.

The lesson that keeps recurring: identical flags are not identical conditions. The authoritative record of what ran is the resolved config inside the model folder plus the engine's startup log, not the command you typed.

A one-line guard before trusting any KV benchmark:

```
# Dump the resolved KV settings the engine actually used
python - <<'PY'
import json, glob
for f in glob.glob("models/**/config.json", recursive=True):
    c = json.load(open(f))
    qc = c.get("quantization_config", {})
    print(f, "| kv:", c.get("kv_cache_dtype"), "| qc_kv:", qc.get("kv_cache_dtype"))
PY
```

Then confirm against the server's own report:

```
vllm serve <model> --kv-cache-dtype fp8_e4m3 --calculate-kv-scales 2>&1 \
  | grep -iE "kv.cache|kv_scale|quantiz"
```

If the two disagree, the file wins, and your benchmark was comparing two different things.

Treat it as an A/B where the only thing allowed to change is the KV dtype. Everything else gets pinned.

```
# Fair KV-quant comparison: pin everything except kv_cache_dtype
from vllm import LLM, SamplingParams

COMMON = dict(
    model="your/model",
    max_model_len=8192,
    gpu_memory_utilization=0.90,
    seed=1234,            # pin the seed
    enforce_eager=True,   # avoid graph-capture variance while measuring
)

sp = SamplingParams(temperature=0.0, max_tokens=512)  # greedy, fixed budget

baseline = LLM(**COMMON, kv_cache_dtype="auto")     # FP16/BF16 KV
fp8      = LLM(**COMMON, kv_cache_dtype="fp8_e4m3", calculate_kv_scales=True)
```

Rules that make the result trustworthy:

`max_tokens` that is too small truncates the chain and manufactures a fake quality gap that has nothing to do with the KV cache.
A minimal results table makes the decision obvious:

| Config | KV bytes/token | Max batch at 8K | Long-ctx retrieval | Reasoning pass@1 | 
|---|---|---|---|---|
| FP16 KV | 128 KB | 1x | 94.1 +/- 0.5 | 71.2 +/- 0.6 | 
| FP8 e4m3 KV | 64 KB | ~2x | 93.8 +/- 0.5 | 70.9 +/- 0.6 | 

(Numbers above are illustrative placeholders; fill them from your own run.)

If the accuracy columns overlap within two standard errors and the batch column doubles, FP8 KV is a free win for that workload. If the long-context column drops outside the error bars, you have found the ceiling for that task and should keep FP16 KV there.

**Does KV cache quantization quantize the model weights too?**

No. It only changes how cached keys and values are stored. Weights and the active compute path are untouched unless you separately apply weight quantization. That is why the quality impact is usually smaller than full weight quantization.

**FP8 or INT8, which should I pick first?**

Start with FP8 `e4m3` on recent NVIDIA or AMD GPUs. It is robust to attention outliers and often needs no calibration. Reach for INT8 when FP8 hardware paths are unavailable or when you have invested in good per-channel calibration.

**Why did two checkpoints give different accuracy with the same command?**

Almost always a setting inside the model folder. A `config.json` KV dtype or baked-in KV scales override your intent silently. Dump the resolved config and the engine startup log before trusting the comparison.

**What does `--calculate-kv-scales` do?**

For the FP8 `e4m3` path it computes the KV scaling factors at runtime instead of relying on whatever scales shipped in the checkpoint, which keeps your comparison honest and avoids depending on an unknown baked-in value.

**Will KV quantization make serving faster, or just smaller?**

Primarily smaller, which indirectly raises throughput: a smaller KV cache means a larger batch or longer context fits, and higher batch occupancy is where serving throughput actually comes from. Do not expect a large per-token latency drop by itself.

**How do I know my eval is sensitive enough to detect a regression?**

Include a task you know FP16 passes and a weak baseline fails. If FP8 KV and FP16 KV score identically on everything, verify the eval can discriminate at all before concluding there is no loss.

`kv_cache_dtype` options (docs.vllm.ai)`--calculate-kv-scales` and quantization support matrix (docs.vllm.ai)
