06:39
2026-07-31
github.com
artificial-intelligence
HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA
HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and relateβ¦