cd /news/ai-infrastructure/line-rate-post-quantum-byzantine-con… · home topics ai-infrastructure article
[ARTICLE · art-134897] src=github.com ↗ pub= topic=ai-infrastructure verified=true sentiment=· neutral

Line-Rate Post-Quantum Byzantine Consensus via L1D Balanced Ternary

Justin Grimm published a reference implementation for a 64-byte L1D cache-resident frame that compresses a 128-node consensus vote bitmask to 26 bytes using balanced ternary (3^5 = 243 ≤ 256), letting the full synchronous descriptor fit in one 64-byte cache line. The accompanying eBPF/XDP driver was benchmarked at 100GbE line-rate (99.2 Mpps aggregate across 8 queues, 12.4 Mpps/core) with p50 = 40.0 ns and p99.9 = 60.0 ns, below the 80.5 ns frame budget, while ML-DSA-44 (NIST FIPS 204) signature verification is offloaded over 2MB hugepage lock-free SPSC rings. The repository, released under the MIT License and targeting SOSP/OSDI/ISCA/ASPLOS, includes a Triton kernel that decompresses 5-trit packed bytes into FP16 ternary weights with zero shared-memory bank conflicts.

read2 min views1 publishedSep 20, 2026
Line-Rate Post-Quantum Byzantine Consensus via L1D Balanced Ternary
Image: Michielbdejong (auto-discovered)

This repository contains the reference eBPF/XDP drivers, Triton GPU lookup kernels, and microbenchmarking suites for the paper:

"Radix Economy and Balanced Ternary Microarchitectures: Resolving the Memory Wall in Line-Rate Post-Quantum Consensus and Nanoscale Computing"

Target Venues: SOSP / OSDI / ISCA / ASPLOS

64-Byte L1D Cache-Resident Frame : Compresses a 128-node consensus vote bitmask to 26 bytes ($3^5 = 243 \le 256$ ), enabling the entire synchronous descriptor (epoch + BLAKE3 accumulator + status flags) to fit within exactly one 64-byte L1D cache line. 2. Deterministic Wire Latency : Evaluated via eBPF XDP at 100GbE line-rate (99.2 Mpps aggregate across 8 queues, 12.4 Mpps/core) with$p50 = 40.0\text{ ns}$ and$p99.9 = 60.0\text{ ns}$ , safely below the$80.5\text{ ns}$ frame budget. 3. State-Crypt Separation : Decouples the 64-byte synchronous consensus frame from asynchronous ML-DSA-44 (NIST FIPS 204) signature verification offloaded over 2MB hugepage lock-free SPSC rings. 4. Conflict-Free GPU Decompression : Triton kernel maps 5-trit packed bytes into FP16 ternary weights with zero shared-memory bank conflicts using single-cycle hardware broadcast addressing.

  • xdp_ternary_filter.c - Production eBPF XDP C driver for line-rate packet parsing, SipHash-2-4 pre-authentication, monotonic epoch tracking, and fast-path quorum accumulation.
  • triton_lut_kernel.py - Triton GPU kernel for high-throughput 5-trit decompression on Tensor Cores.
  • benchmark_harness.py - Microarchitectural evaluation reproducing the latency and throughput ablations across 64-byte ternary and 96-byte binary frames.
  • LICENSE - MIT License.
clang -O2 -target bpf -c xdp_ternary_filter.c -o xdp_ternary_filter.o

ip link set dev eth0 xdpgeneric obj xdp_ternary_filter.o sec xdp
python3 triton_lut_kernel.py
python3 benchmark_harness.py
@article{grimm2026radix,
  title={Radix Economy and Balanced Ternary Microarchitectures: Resolving the Memory Wall in Line-Rate Post-Quantum Consensus and Nanoscale Computing},
  author={Grimm, Justin},
  year={2026}
}

MIT License - Copyright (c) 2026 Justin Grimm.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @justin grimm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/line-rate-post-quant…] indexed:0 read:2min 2026-09-20 ·