cd /news/artificial-intelligence/show-hn-external-kv-cache-offloading… · home topics artificial-intelligence article
[ARTICLE · art-74280] src=github.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Show HN: External KV Cache Offloading Cuts Long Horizon Inference Costs by 50%

OpenLake, an open-source storage engine for offloading LLM KV caches from GPU memory to RAM and NVMe, cuts GPU time by 48.2% for long-context inference, reducing a 1,169-second workload to 606 seconds. The project achieves 1.72× lossless KV compression and 66× faster time-to-first-token at 128K context, dropping from 44 seconds to 0.6 seconds. OpenLake provides connectors for vLLM and SGLang without modifying inference engines.

read1 min views1 publishedJul 26, 2026

Hey HN, we’re the developers of OpenLake, an open source storage engine for off LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory.

A single 256K token conversation on Gemma 4 31B produces approximately 43GB of KV state, more than half the memory of an 80GB H100. The problem becomes even harder across a cluster: a prefix cached on one GPU host is unavailable when the next request lands on a different GPU, forcing the new GPU to repeat work the fleet has already completed.

Once the KV cache is offloaded, network bandwidth becomes a major constraint on read latency. To move less data across the wire, we built deferred materialization: a custom CUDA kernel that losslessly compresses KV blocks before they leave GPU memory and decompresses them on the GPU after retrieval. In our tests, this achieved:

  • 1.72× lossless KV compression. - Approximately 600GB/s decompression throughput on an H100 - 80GB/s of effective KV throughput over a physical 50GB/s link

At 128K context, retrieving cached KV reduces TTFT from 44 seconds to 0.6 seconds, a 66× improvement. Across the complete workload, GPU time reduces from 1,169 seconds to 606 seconds, saving 48.2% of GPU cost.

OpenLake is written in Rust and uses io_uring with one pinned runtime per physical core. We provide connectors for vLLM and SGLang so the cache can be enabled without modifying the inference engine itself.

I would love to hear how others are handling KV reuse across GPU hosts, especially for long contexts, and get to know your thoughts.

Thanks!

GitHub: [https://github.com/openlake-project/openlake](https://github.com/openlake-project/openlake)

Here is our blog: [https://cloud.theopenlake.com/blog/taming-the-beast-managing...](https://cloud.theopenlake.com/blog/taming-the-beast-managing-100-tb-of-kv-cache-on-open-source-inference)

Comments URL: [https://news.ycombinator.com/item?id=49057767](https://news.ycombinator.com/item?id=49057767)

Points: 10

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openlake 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-external-kv-…] indexed:0 read:1min 2026-07-26 ·