cd/entity/vLLM 0.28.0· home› entities› vLLM 0.28.0
grep -l @vllm 0.28.0 /news/*.json | wc -l → 1

vLLM 0.28.0

mentions 1 type Person feed RSS

// recent coverage 1 mentions

00:00
2026-09-15
doug.sh
ai-infrastructure

Keeping vLLM's Prefix Cache Warm Between Agent Turns

Tuning vLLM 0.28.0's prefix cache raised the share of prompt tokens served from cache from 55% to 95% on a local Qwen3.8-27B coding agent, cutting average time to first token from 26-28 seconds to 7.3…

// co-occurs with top 7 entities
// topics top 5 topics