04:00
2026-07-20
arxiv.org
large-language-models
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
A new study from arXiv finds that query-aware prompt compression, which invalidates prefix caches on every call, can be more expensive than naive caching under realistic hit rates on Anthropic's Sonneβ¦