01:33
2026-09-09
github.com
artificial-intelligence
Zero-Copy KV-Cache Migration Protocol (81.6ms Latency)
An open-core protocol for streaming LLM KV-cache states across datacenters cuts time-to-first-token latency to 81.73 ms and reduces GPU VRAM compute overhead by up to 95%, according to benchmarks releβ¦