Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
A new study from arXiv finds that query-aware prompt compression, which invalidates prefix caches on every call, can be more expensive than naive caching under realistic hit rates on Anthropic's Sonne…