{"slug": "keeping-vllm-s-prefix-cache-warm-between-agent-turns", "title": "Keeping vLLM's Prefix Cache Warm Between Agent Turns", "summary": "Tuning vLLM 0.28.0's prefix cache raised the share of prompt tokens served from cache from 55% to 95% on a local Qwen3.8-27B coding agent, cutting average time to first token from 26-28 seconds to 7.3 seconds and worst-case wait from 514 seconds to 54 seconds, according to a first-person account of the setup. The stack runs the W4A16 AutoRound quant across two RTX 3090s with tensor parallelism, an oh-my-pi agent, and the Bifrost LLM gateway, with a CPU offload tier holding evicted KV blocks in 48 GiB of /dev/shm. The author reports that combining the two cards with tensor parallelism instead of running a replica per card reduced KV restores to about two an hour, and that Qwen3.8-27B's hybrid linear-attention layers force the cache to work in 2048-token blocks rather than the usual 16.", "body_md": "# Keeping vLLM's Prefix Cache Warm Between Agent Turns\n\n## From 55% to 95% cached [🔗](#from-55-to-95-cached)\n\nI’ve been playing with a few different ways to host Qwen3.8 locally. I’m aiming for something that can replace Claude Code for most of my tasks. Right now the 27B runs on two RTX 3090s under vLLM, and it’s usable, but the first day was rough. In the morning the wait before the first word averaged about half a minute, and some replies took several minutes. The time went into rereading large parts of every prompt from scratch, because the cached copy from the previous turn had been thrown away or no longer matched. By the evening the server was keeping almost everything from one turn to the next.\n\n|  | Morning | Evening | \n|---|---|---|\n| Prompt tokens served from cache | 55% | 95% | \n| Average wait before the first word | 26-28 s | 7.3 s | \n| Worst wait | 514 s | 54 s | \n\nThe agent-side settings from that day are in [Tuning a Local Coding\nAgent](https://doug.sh/posts/tuning-a-local-coding-agent-oh-my-pi/).\n\nThe stack is [oh-my-pi](https://github.com/can1357/oh-my-pi) talking to\n[Bifrost](https://github.com/maximhq/bifrost), an LLM gateway that also aggregates the MCP tool servers,\ntalking to vLLM 0.28.0, serving a W4A16 AutoRound quant across both cards:\n\n```\n--tensor-parallel-size 2 --max-model-len 262144 --max-num-seqs 16\n--max-num-batched-tokens 8192 --long-prefill-token-threshold 2048\n--kv-transfer-config '{\"kv_connector\":\"OffloadingConnector\", ...}'\n```\n\nBetween the two columns I also combined the two cards with tensor parallelism instead of running a replica on each, capped thinking at 6,000 tokens, changed the batch token budget, and turned on speculative decoding and CPU offload.\n\nA coding agent resends the whole conversation every turn. Reading it runs at 700-900 tokens per second per request on the combined engine, so a 120,000-token turn takes about 150 seconds from scratch. If the server still has the previous turn’s work in its KV cache, it reads only the new 2,000 tokens or so and starts in 3 seconds.\n\n## What the server keeps, and where [🔗](#what-the-server-keeps-and-where)\n\nThe GPU prefix cache holds the results of reading tokens, indexed by a hash of each block together with every block before it. A lookup walks the prompt from the start and stops at the first block that doesn’t match. It matches a prefix rather than diffing two prompts, so a change at token 12,000 of a 120,000-token prompt throws away 108,000 tokens of finished work.\n\nThe CPU tier copies blocks evicted from GPU memory to host RAM (48 GiB in `/dev/shm` here) and\ncopies them back on demand. While I was still running two replicas, each had room for two or three\nlong conversations, so contexts were evicted all day, and a restore over PCIe brought a whole\n42,000-token context back in under half a second where reading it cold on that replica took 24\nseconds. Once I combined the cards with tensor parallelism the restores dropped to about two an hour.\nAnything in neither tier is read again from scratch.\n\nvLLM also documents a filesystem tier and a peer-to-peer connector. I looked at them for sharing KV between the two per-GPU replicas I started with, and combining the cards removed the need. The filesystem tier also never evicts anything.\n\n## Measuring it [🔗](#measuring-it)\n\nWatch `prompt_tokens_cached_total / prompt_tokens_total`, the share of prompt tokens the server didn’t\nhave to read. The GPU hit counter alone looked healthy on my morning run while that share was 55%.\n\n## What the hybrid model changes [🔗](#what-the-hybrid-model-changes)\n\nQwen3.8-27B is a hybrid: most of its layers use a linear attention variant that keeps a small fixed amount of state instead of a KV cache that grows with the prompt. vLLM can only resume those layers from a saved copy of that state, and it only saves one where a prefill step ends on a 2048-token boundary, so on this model the cache works in 2048-token blocks instead of the usual 16.\n\nA 2048-token block also changes what the scheduler can fit: a request with less than 2048 tokens of\nbatch budget left in a step gets nothing that step, and the scheduler stops looking at the queue\nbehind it. With\n`--max-num-batched-tokens 4096` long requests sat queued behind a single prefill. 8192 with\n`--long-prefill-token-threshold 2048` fits three or four prefill chunks in a step instead of one, and a\nbigger budget than that stalled decoding for 13-19 seconds a step.\n\nThe same saved states are why the CPU tier stopped helping. vLLM keeps only a few of them per conversation, mainly the one where the last request ended, and drops the rest. The next turn starts exactly there, so the GPU cache serves it. A restore from the CPU tier only helps if a saved state still exists at the boundary where the restored blocks end, and one rarely does. The combined pool is also far bigger than either replica’s and peaked at 37% full during the evening run, so little is evicted in the first place. One hour of agent traffic on the combined setup looked like this:\n\n| Over one hour |  | \n|---|---|\n| Copied to RAM | ~40 GB in ~700 copies | \n| Restored from RAM | ~250 MB in 2 restores | \n| Tokens asked for | ~800,000 | \n| Tokens served | ~12,000 (1.5% of those asked for) | \n| Cost | ~9 s of copying, 48 GiB of RAM | \n\nOther hours on it looked the same. I leave it on because it costs so little, but on this model it can’t do much.\n\n## What breaks the prefix [🔗](#what-breaks-the-prefix)\n\n### Tool schemas that shuffle their keys [🔗](#tool-schemas-that-shuffle-their-keys)\n\nEvery hour or so, a few turns came back 2-25% cached with 90-180 seconds of reading, on an idle engine with the KV pool almost empty, from prompts 94-96% identical to the previous turn.\n\nReplaying pairs of consecutive turns from a traffic capture against an idle server reproduced the misses, so it wasn’t eviction. Tokenizing both prompts and diffing them showed they matched for exactly 12,622 tokens and then diverged inside the tool definitions. Four MCP tools had arrived with the keys of their parameter schema in a different order, on 8% of consecutive turns in the capture. The schemas were identical once sorted. The Qwen template renders the tool definitions at the top of the system turn, before the system prompt itself, so one shuffled schema invalidated everything behind it.\n\nThe shuffle comes from the gateway: mcp-go decodes each MCP server’s tool list into a Go map, which\nhas no key order, and Bifrost copies that map back out with a `range` loop, which Go deliberately\nrandomizes. Bifrost refreshes its tool list every 10 minutes, so the order it serves can change on any\nrefresh, and the order changes in the capture were spaced at multiples of 10 minutes.\n\nI sent fixes for both and they’re merged: [bifrost#7170](https://github.com/maximhq/bifrost/pull/7170)\nsorts the schema keys during conversion, and [mcp-go#984](https://github.com/mark3labs/mcp-go/pull/984)\nrecords the key order while decoding so a gateway can keep the server’s own order. Before that, my\nworkaround was in the chat template. Where the Qwen template serializes each tool’s parameters, I\nchanged\n\n```\ntool.function.parameters | tojson\n```\n\nto\n\n```\ntool.function.parameters | tojson(sort_keys=True)\n```\n\nso the keys come out in the same order whatever order they arrive in. Everything else I changed that day had taken the cached share from 55% to 78%. The template edit on its own took it from 78% to 95%, and the average wait from 21 seconds to 7.3. Hosted APIs match a byte-identical prefix that includes the tools block too, so the same gateway bug would have cost the cached read discount there.\n\n### Live status text near the top of the prompt [🔗](#live-status-text-near-the-top-of-the-prompt)\n\noh-my-pi lists each subagent’s state (running, idle, parked) roughly 29,000 characters into the system prompt, so every state change invalidated the main agent’s prompt from that point on. After one 24-minute subagent run, the main agent came back to a 150,000-token prompt with 14% of it cached. The block is hardcoded in oh-my-pi’s prompt template, so the fix belongs upstream, and I haven’t asked yet why it’s there.\n\n### Compaction [🔗](#compaction)\n\nWhen the conversation gets too long, the agent replaces old history with a summary. A summary changes\nthe prompt, so the turn after compaction reads all of it again. The only lever is how often it runs:\ntell the agent the model’s real context window and raise the threshold, both covered in\n[Tuning a Local Coding Agent](https://doug.sh/posts/tuning-a-local-coding-agent-oh-my-pi/).\n\n### Per-agent text before shared text [🔗](#per-agent-text-before-shared-text)\n\nEach subagent’s prompt has its own id and peer list about two thirds of the way through the system\nprompt, so sibling subagents share only the first 60-70% of their prefix. The main agent and subagents\nalso list their tools in different orders, so they diverge. Putting the shared parts (tool definitions\nand base instructions) first and the per-agent parts last would let every agent share the same first\n15,000 tokens or so. As with the status block, the order may be\n[Chesterton’s fence](https://en.wikipedia.org/wiki/Chesterton%27s_fence), and I haven’t asked.\n\n### What doesn’t break it [🔗](#what-doesnt-break-it)\n\nAppending tool results, which is most of an agent’s traffic, extends the prefix and leaves everything before it cached. Big tool outputs still cost reading time, and capping them is a separate problem.\n\n## Benchmark with the cache you will run with [🔗](#benchmark-with-the-cache-you-will-run-with)\n\nSpeculative decoding looked bad in the morning: it doubled the writing speed per stream (63 tokens per second against 27), but its slower reading made overlapping requests queue, and the median wait went from 29 seconds to 121. At 95% cached there’s hardly any reading left to slow down, and it became the better mode: 24-32 tokens per second per request against ~22, with a shorter wait.\n\n## A checklist [🔗](#a-checklist)\n\n1. Measure the share of prompt tokens served from cache, not the number of cache hits.\n2. When a turn is slow, tokenize it and the previous turn and find the first token that differs.\n3. Reproduce the miss against an idle server before blaming eviction.\n4. Serialize everything that goes into the prompt the same way every time. Sort JSON keys.\n5. Once that holds, put the parts of the prompt that change at the end: tool definitions and fixed instructions first, then history, then anything per-agent or live.\n6. Compact as rarely as your agent allows.\n7. Report the cache hit rate with any benchmark number.\n\n*I had help with this one. Anthropic’s Claude and the local Qwen models described here helped me dig\nthrough the traffic captures, write the replay script and draft this post. I ran the experiments,\nchecked the numbers and rewrote anything that sounded like a chatbot, so the mistakes are mine.*\n\n## Sources [🔗](#sources)\n\n- [vLLM automatic prefix caching](https://docs.vllm.ai/en/stable/features/automatic_prefix_caching/)\n- [vLLM KV offloading](https://docs.vllm.ai/en/stable/features/kv_offloading_usage/)\n- vLLM prefix caching on hybrid models with speculative decoding: [#53670](https://github.com/vllm-project/vllm/issues/53670) and[#53504](https://github.com/vllm-project/vllm/issues/53504)\n- vLLM `mamba_cache_mode` and`prefix_cache_retention_interval` :[vllm/config/cache.py](https://github.com/vllm-project/vllm/blob/main/vllm/config/cache.py)\n- Bifrost MCP tool schema key order: [issue #7169](https://github.com/maximhq/bifrost/issues/7169) and[PR #7170](https://github.com/maximhq/bifrost/pull/7170)\n- mcp-go property order: [issue #983](https://github.com/mark3labs/mcp-go/issues/983) and[PR #984](https://github.com/mark3labs/mcp-go/pull/984)\n- [oh-my-pi](https://github.com/can1357/oh-my-pi)\n- [Qwen3.8-27B model card and chat template](https://huggingface.co/Qwen/Qwen3.8-27B)\n- [Model weights](https://huggingface.co/dbirks/Qwen3.8-27B-W4A16-AutoRound)\n- [syv-ai RTX 3090 vLLM image](https://github.com/syv-ai/qwen38-27b-rtx3090)", "url": "https://wpnews.pro/news/keeping-vllm-s-prefix-cache-warm-between-agent-turns", "canonical_source": "https://doug.sh/posts/vllm-kv-cache-agents/", "published_at": "2026-09-15 00:00:00+00:00", "updated_at": "2026-10-01 01:17:26.582583+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-agents", "mlops", "ai-tools"], "entities": ["vLLM 0.28.0", "Qwen3.8-27B", "oh-my-pi", "Bifrost", "RTX 3090", "W4A16 AutoRound", "OffloadingConnector", "Claude Code"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/keeping-vllm-s-prefix-cache-warm-between-agent-turns", "markdown": "https://wpnews.pro/news/keeping-vllm-s-prefix-cache-warm-between-agent-turns.md", "text": "https://wpnews.pro/news/keeping-vllm-s-prefix-cache-warm-between-agent-turns.txt", "jsonld": "https://wpnews.pro/news/keeping-vllm-s-prefix-cache-warm-between-agent-turns.jsonld"}}