KV Cache Quantization in LLM Serving: FP8 and INT8 Tradeoffs, the Silent config.json Trap, and How to Measure It Fairly A developer detailed how quantizing the KV cache to FP8 or INT8 roughly halves per-token memory in LLM serving, from about 128 KB per token in FP16 to 64 KB, enabling longer contexts or larger batches on the same GPU. The writeup warns that checkpoint-level KV cache dtypes and baked-in scaling factors in a model's config.json can silently change outputs even when serving flags are identical, and outlines a fair measurement protocol that names the task, context length, and whether a distribution or single greedy run was used. The KV cache is often the real memory ceiling in LLM serving, not the weights. Quantizing it to FP8 or INT8 roughly halves KV memory, which lets you serve longer contexts or more concurrent requests on the same GPU. The catch: the quality impact is real but small when measured correctly, and it is easy to measure it wrong. A single kv cache dtype flag, or a scaling factor baked into a model's config.json , can quietly change your outputs while every dashboard still shows green. This post covers where the savings come from, what actually degrades, the silent-config trap, and a fair measurement protocol. When a transformer generates tokens autoregressively, it stores the key and value tensors for every past token so it does not recompute them. That store is the KV cache, and it grows linearly with sequence length and batch size. A rough per-token KV size: bytes per token = 2 K and V num layers num kv heads head dim dtype bytes For a mid-size model with 32 layers, 8 KV heads grouped-query attention , and head dim 128, in FP16 that is: 2 32 8 128 2 = 131,072 bytes 128 KB per token At 8K context that is 1 GB per sequence, before you batch anything. Weights are a fixed cost you pay once; the KV cache is a per-request cost that scales with your traffic. That is why it is the highest-leverage place to spend a quantization budget. Drop the dtype from 2 bytes to 1 byte and you halve that 128 KB to 64 KB per token, which directly buys you longer contexts or a bigger batch. Both formats store each cached value in 8 bits, so the memory savings are effectively identical, about 2x versus FP16/BF16. The difference is in how they represent numbers and what that does to accuracy. FP8 typically e4m3 or e5m2 keeps an exponent field, so it covers a wide dynamic range with graceful precision falloff. The vLLM docs note that FP8 KV cache is supported on CUDA 11.8 and later and on ROCm, and that e4m3 is the common default on NVIDIA hardware. Because attention keys and values can have occasional large-magnitude outliers, the exponent bits help: an outlier does not blow out the whole tensor's scale. INT8 uses a uniform grid with a per-tensor or per-channel scale. When the distribution is well behaved, INT8 can actually be more precise than FP8 inside its range because it does not spend bits on an exponent. The risk is outliers: one large value forces a coarse scale on everything else. INT8 KV usually wants calibration to pick good scales. Practical reading: Less than people fear, if you measure the right thing. KV quantization only touches the cache, not the weights or the activations on the compute path, so the error is confined to attention's read of past tokens. In practice: The headline number "FP8 KV costs X percent" is meaningless without naming the task, the context length, and whether you measured a distribution or a single greedy run. Here is the failure that burns teams. You benchmark model A and model B with what looks like the identical serving command, same flags, same prompts, and you get different accuracy. The flags were identical. The conditions were not. Two things commonly hide inside a model directory rather than on your command line: A checkpoint-level KV cache dtype. Some published checkpoints ship a quantization config in config.json for example a quantization config or a KV-specific field that sets the cache to FP8 even when your launch command says nothing about it. Your serving engine reads it and honors it. Your benchmark harness never sees it. Baked-in KV scaling factors. FP8 KV can use static per-layer scales stored in the checkpoint. If one checkpoint has calibrated scales and another falls back to default scales, their attention precision differs even though both report "fp8 KV" in the logs. vLLM exposes --calculate-kv-scales to compute scales at runtime for the e4m3 path precisely so you are not silently depending on whatever shipped in the file. The lesson that keeps recurring: identical flags are not identical conditions. The authoritative record of what ran is the resolved config inside the model folder plus the engine's startup log, not the command you typed. A one-line guard before trusting any KV benchmark: Dump the resolved KV settings the engine actually used python - <<'PY' import json, glob for f in glob.glob "models/ /config.json", recursive=True : c = json.load open f qc = c.get "quantization config", {} print f, "| kv:", c.get "kv cache dtype" , "| qc kv:", qc.get "kv cache dtype" PY Then confirm against the server's own report: vllm serve