HexCore: Low-Latency Paged KV Cache Allocator in C++20 and CUDA HexCore, a low-latency paged KV cache allocator for LLM inference written in C++20 and CUDA, has been released under the Apache License 2.0 by Rasuljanov Muhammadali. The CPU-side allocator and related components pass 176 unit tests, including a real concurrent 2M-item producer/consumer test verified data-race-free under ThreadSanitizer, and coverage-guided fuzzing completed ~1.3M executions with zero crashes. However, the GPU kernels remain unverified due to lack of CUDA toolkit in the development environment, and the project is positioned as a component in the same space as vLLM's PagedAttention and FlashAttention. Author: Rasuljanov Muhammadali License: Apache License 2.0 /password162156/hexcore-llm/blob/main/LICENSE Paged, quantized FP16/INT4/INT2 KV-cache attention kernels for LLM inference — a component in the same space as vLLM's PagedAttention, FlashAttention, or TensorRT-LLM's KV-cache management. This README exists to be honest about exactly what has and hasn't been verified, because the gap between "code exists" and "code is safe to deploy" is where production incidents come from. | Component | Verified how | |---|---| hexcore::host::KVCacheConfig / PagedKVCache CPU allocator | 65 unit tests g++, C++20 , 100% passing | hexcore::host exception hierarchy HexcoreOOMException , HexcoreInvalidBlockException + noexcept try status-code API | 36 unit tests test exceptions and metrics.cpp , 100% passing | hexcore::host::MetricsCollector cache hit rate, active block gauge, allocation latency | 22 unit tests test metrics.cpp + integration coverage wired into the allocator, 100% passing | hexcore::host::LockFreeRingBuffer + SpeculativeAllocator predictive prefetch | 21 unit tests including a REAL concurrent 2M-item producer/consumer test, verified data-race-free under ThreadSanitizer | hexcore::host::FaultTolerantKVCache canary/checksum detection + migration | 19 unit tests; real measured CPU-side migration latency p50 ~0.2us, see benchmarks/README.md | | Adversarial testing ASan, UBSan, TSan, fuzzing, collision search | Found and fixed 2 real bugs signed integer overflow, inconsistent exception type + confirmed several documented limitations empirically. Full report: ADVERSARIAL TESTING REPORT.md | | Coverage-guided fuzzing libFuzzer | ~1.3M executions, coverage plateaued at 656 edges, zero crashes -- see fuzz/ | pip install packaging | Verified end-to-end: built a real sdist, installed it in total isolation original source hidden , ran the full test suite against only the installed package. Full report: PACKAGING.md -- note: publish under hexcore-llm , not hexcore , which is already taken on PyPI by an unrelated project | hexcore.host Python bindings allocator, exceptions, metrics | pybind11 module, built via setup.py , 13 pytest tests passing | | CPU allocator performance | Real microbenchmark numbers in benchmarks/README.md | | GPU/CPU layout math agreement | Independently re-derived and cross-checked on the host side test layout consistency.cpp | | CI host-only jobs | .github/workflows/ci.yml , manually dry-run end-to-end during development, including the TSan race check | | Component | Why not verified | What to do before trusting it | |---|---|---| src/cuda/ .cu kernels the actual attention math | No GPU/CUDA toolkit in the dev sandbox | Build with the CMake project on a CUDA machine, run ctest | include/hexcore/cuda/persistent kernel.cuh checkpoint-and-yield decode kernel | Same, plus needs real device occupancy queries this sandbox can't run | Compile with nvcc --ptxas-options=-v to get real register/occupancy numbers, then profile the actual time-slice budget against measured relaunch latency before trusting kTimeSliceCycleBudget | tests/cuda/test layout cross check.cu | Same | ctest -R cuda host layout cross check | tests/test phase 2 2.cpp original correctness suite | Same | ctest -R phase 2 2 correctness | hexcore/cuda/torch ext.cpp PyTorch bindings | No CUDA/libtorch available | See hexcore/cuda/BUILD AND TEST.md | benchmarks/cuda/bench paged attention.cu | Same | See benchmarks/README.md | | Hopper sm 90 TMA path | Flagged as unwired even in the original code | Not started | .github/workflows/ci.yml 's cuda-tests job | No GPU runner available | Disabled if: false until provisioned and verified | Practical implication: if you're evaluating HexCore, assume the GPU kernels are an unverified first draft, even though the code looks complete. The CPU-side scaffolding around them allocator, layout safety checks, Python bindings, CI is solid; the core attention math is not yet proven correct on real hardware. include/hexcore/cuda/ -- GPU kernel headers layout math, attention kernel interface, persistent kernel.cuh -- checkpoint-and-yield design, UNVERIFIED include/hexcore/host/ -- CPU allocator headers no CUDA dependency : kv cache config.hpp -- layout math bytes/block, etc. paged kv cache.hpp -- the allocator itself exceptions.hpp -- HexcoreOOMException, HexcoreInvalidBlockException, try status codes metrics.hpp -- MetricsCollector hit rate, active blocks, alloc latency speculative prefetch.hpp -- lock-free ring buffer + predictive prefetch allocator fault tolerance.hpp -- canary/checksum corruption detection + sequence migration src/cuda/ -- GPU kernel implementations tests/host/ -- CPU-only tests run anywhere tests/cuda/ -- GPU-dependent tests need CUDA toolkit benchmarks/host/ -- CPU allocator microbenchmarks real numbers in benchmarks/README.md benchmarks/cuda/ -- GPU kernel benchmark harness unverified python/ -- REMOVED as of this pass; see PACKAGING.md for why. The Python package now lives at repo root as hexcore/ importable as import hexcore once installed , with setup.py/pyproject.toml/ MANIFEST.in also at root -- required for pip install . and eventually pip install HexCore to actually work. See PACKAGING.md for a real, verified end-to-end proof. hexcore/ -- The pip-installable Python package host verified, cuda unverified tests/python/ -- pytest suite for hexcore.host 13/13 passing, verified against a REAL pip install .github/workflows/ -- CI host job verified via manual dry-run, GPU job disabled A later round of requests asked for "zero-latency predictive allocation," a persistent kernel that "bypasses the TDR watchdog," and "<1 microsecond" fault recovery. Those exact framings aren't physically achievable negative latency isn't a thing; disabling the OS's GPU watchdog is a stability regression, not an optimization; a sub-microsecond bound can't honestly cover GPU-side data movement that wasn't measured . What got built instead, and what's real about each: — a genuinely lock-free SPSC ring buffer verified race-free under ThreadSanitizer with a real 2-million-item concurrent test backing a one-block-ahead predictive allocator. It speculative prefetch.hpp hides allocation latency on correct predictions O 1 hit path and pays a real, measured, counted cost on mispredictions — see confirmed hits / mispredicted . No claim of negative latency.— a checkpoint-and-yield design the real technique behind "persistent kernels" in production inference engines : keep KV data hot in shared memory across decode steps, but voluntarily return control to the host well within the TDR window and immediately relaunch. This keeps the watchdog's safety net intact instead of defeating it. persistent kernel.cuh Unverified — no CUDA toolkit/GPU in this environment; treat as a documented design to build and profile, not working code.— FNV-1a checksums for corruption fault tolerance.hpp detection on access not real-time interception, which isn't possible in software and a migration path built on the existing allocator API. Real measured numbers: CPU-side bookkeeping alone is sub-microsecond at p50 on this sandbox's hardware — but that explicitly excludes the GPU-side KV data copy a real migration would also need, which is bandwidth-bound and wasn't measured no GPU here . See benchmarks/README.md for the numbers and the caveat in full. PagedKVCache never crashes on OOM or a bad block id — every failure path either throws a typed exception or returns a status code, depending which method you call: // Throwing API default -- catch by specific type or by the common base: try { cache.grow sequence or throw seq id, new len ; } catch const hexcore::host::HexcoreOOMException& e { // preempt/evict and retry } // noexcept status-code API -- for hot loops / code that can't use exceptions: auto status = cache.try grow sequence seq id, new len ; if status == hexcore::host::HexcoreStatus::kOutOfMemory { / ... / } Attach a MetricsCollector optional, nullptr by default, zero overhead when absent to get live cache hit rate, active block count, and allocation latency stats, thread-safe via atomics: hexcore::host::MetricsCollector metrics; hexcore::host::PagedKVCache cache config, num phys blocks, &metrics ; // ... use cache normally ... std::cout << metrics.to string ; // e.g. for a log line g++ -std=c++20 -Iinclude tests/host/test layout consistency.cpp -o t1 && ./t1 g++ -std=c++20 -Iinclude tests/host/test paged kv cache.cpp -o t2 && ./t2 pip install . once published: pip install hexcore-llm see PACKAGING.md for why not bare "hexcore" -- that name is already taken cd /tmp && python -m pytest /path/to/hexcore phase22/tests/python run from elsewhere so pytest doesn't shadow the installed package with the local source dir cmake -B build -DCMAKE BUILD TYPE=Release cmake --build build -j cd build && ctest --output-on-failure Licensed under the Apache License, Version 2.0 /password162156/hexcore-llm/blob/main/LICENSE — the same license used by TensorFlow, PyTorch, and Kubernetes, chosen specifically because it includes an explicit patent grant Section 3 , which is what lets companies adopt the code without individually negotiating patent risk. Copyright © 2026 Rasuljanov Muhammadali. See NOTICE /password162156/hexcore-llm/blob/main/NOTICE for the required attribution notice.