Show HN: Proxima serves 4x more requests with no hardware change on vLLM
Proxima, an out-of-tree vLLM plugin implementing STAR-KV low-rank KV cache compression, serves 4x more concurrent requests at 8192 context and boots at 16384 context where plain vLLM refuses to start,β¦