15:04
2026-08-11
github.com
artificial-intelligence
Show HN: Proxima serves 4x more requests with no hardware change on vLLM
Proxima, an out-of-tree vLLM plugin implementing STAR-KV low-rank KV cache compression, serves 4x more concurrent requests at 8192 context and boots at 16384 context where plain vLLM refuses to start,โฆ