18:31
2026-09-14
j9s.io
ai-infrastructure
Improving Throughput by Optimising KV Cache Efficiency for Agentic Workloads
Agentic workloads run roughly double the context of chat conversations by turn 10 and require a median of 15,146 prefill tokens and 501 decode tokens per turn, according to an analysis of a custom harβ¦