13:02
2026-07-26
github.com
artificial-intelligence
Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%
OpenLake, an open-source storage engine for offloading LLM KV caches from GPU memory to RAM and NVMe, cuts GPU time by 48.2% for long-context inference, reducing a 1,169-second workload to 606 secondsβ¦