00:00
2026-09-15
doug.sh
ai-infrastructure
Keeping vLLM's Prefix Cache Warm Between Agent Turns
Tuning vLLM 0.28.0's prefix cache raised the share of prompt tokens served from cache from 55% to 95% on a local Qwen3.8-27B coding agent, cutting average time to first token from 26-28 seconds to 7.3…