cd /news/ai-infrastructure/show-hn-reflex-a-gguf-cuda-inference… · home topics ai-infrastructure article
[ARTICLE · art-138414] src=github.com ↗ pub= topic=ai-infrastructure verified=true sentiment=↑ positive

Show HN: Reflex – a GGUF/CUDA inference engine tuned for cold-start latency

Lateos AI released Reflex, an open-source GGUF-native Rust and CUDA inference engine that compiles every CUDA kernel ahead-of-time with nvcc at build time to eliminate the multi-second JIT tax on first use. On an RTX A6000, Reflex's system1 subcommand scored three candidates for the prompt "The capital of France is" in 8524.253 ms from process start to result, assigning " Paris" a probability of 0.997350 versus 0.002239 for " London" and 0.000411 for " Berlin". The project targets cold-start workloads such as serverless/FaaS, single-shot CLI calls, batch jobs and edge devices, and its system1 path currently supports dense and MoE Qwen3 only, while generate supports Qwen3, the Qwen3.5 hybrid mixer and DeepSeek-V2/V3 MLA.

read11 min views1 publishedSep 23, 2026
Show HN: Reflex – a GGUF/CUDA inference engine tuned for cold-start latency
Image: Michielbdejong (auto-discovered)

A high-performance, GGUF-native Rust & CUDA inference engine optimized for cold-start latency and real-time "System 1" agent decision loops — process launch to first token, not sustained server throughput.

Every CUDA kernel is compiled ahead-of-time by nvcc at build time and shipped inside the binary — never compiled at runtime via NVRTC — so there's no multi-second JIT tax on first use, the way there is with a runtime-compilation design. That's the whole bet: be the fastest way to turn a cold process into one output token, then get out of the way.

Requires the CUDA toolkit (nvcc on PATH, or CUDA_PATH/ CUDA_HOME set) and an NVIDIA GPU. REFLEX_SKIP_CUDA=1 cargo build skips kernel compilation for editing/type-checking on a machine without CUDA (no subcommand will actually run kernels in that mode).

git clone https://github.com/lateos-ai/reflex.git
cd reflex
cargo build --release

cargo run --release --bin reflex -- smoke

cargo run --release --bin reflex -- generate <path-to-gguf> "Once upon a time"

reflex is a single binary with subcommands: generate (load a GGUF and generate tokens), system1 (single-pass, non-autoregressive candidate scoring — the "System 1" decision-loop path), smoke (the AOT-pipeline check above), bench (warm-latency microbenchmark), check (byte-exact-vs-reference correctness check, CI-scriptable), and stdio/ uds (local JSON-line IPC, both need --features ipc). Run reflex <subcommand> with no further arguments to see that subcommand's own usage.

system1 takes a local GGUF path, so download the file first with the hf CLI (pip install -U huggingface_hub), then point system1 at it — this scores each --candidate against the prompt in a single pass, with no autoregressive decode loop:

hf download Qwen/Qwen3-0.6B-GGUF Qwen3-0.6B-Q8_0.gguf --local-dir .

cargo run --release --bin reflex -- system1 Qwen3-0.6B-Q8_0.gguf \
  "The capital of France is" \
  --candidate " Paris" --candidate " London" --candidate " Berlin"

Real output from this exact command (RTX A6000):

REFLEX_SYSTEM1_CANDIDATE_OK idx=0 text=" Paris" token_ids=[12095] score=17.407064 probability=0.997350
REFLEX_SYSTEM1_CANDIDATE_OK idx=1 text=" London" token_ids=[7148] score=11.308186 probability=0.002239
REFLEX_SYSTEM1_CANDIDATE_OK idx=2 text=" Berlin" token_ids=[19846] score=9.612655 probability=0.000411
REFLEX_SYSTEM1_OK process_start_to_result_ms=8524.253 num_candidates=3 best_idx=0 best_text=" Paris" entropy=0.028154

probability is relative to this candidate set only, not a vocab-wide probability — see Model::system1_evaluate's doc comment in src/model.rs. system1 currently supports dense/MoE Qwen3 only; the Qwen3.5 hybrid mixer and DeepSeek-V2/V3 (MLA) are rejected with a clear error (generate supports all four architectures — swap in a DeepSeek GGUF the same way for a generate run instead).

generate can also pull a GGUF straight from the Hub itself, via this project's own Rust hf-hub integration — --model <org/repo:file.gguf> or --quickstart, both requiring cargo build --features download:

cargo run --release --features download --bin reflex -- generate --quickstart "Once upon a time"

Closing the steady-state-throughput gap with llama.cpp/vLLM is a kernel-optimization race against projects with a multi-year head start — Rust as a language doesn't change who wins it. What none of llama.cpp, vLLM, or candle are built for or measured against is a cold invocation — serverless/FaaS, single-shot CLI/dev-tool calls, batch/cron jobs, edge devices that wake on demand. A naive runtime-JIT design pays a real, measured multi-second tax on first kernel use; vLLM took ~235s to become ready (CUDA graph capture) before serving one request; llama.cpp avoids both because its kernels are compiled by nvcc at build time, not at process start.

Target metric: energy-to-first-token from cold start (joules, process launch to first generated token) — a real, underexplored gap. Existing energy benchmarks measure warm/steady-state joules-per-token, not full-lifecycle cold-start cost.

Target models: Qwen and DeepSeek families.

All comparisons are cold-start (process launch to first token/result), same ThunderCompute A6000, n=3, external wall-clock (/usr/bin/time -v — process launch to exit, not just Reflex's own internal timer). Full methodology, disclosed caveats, and per-run numbers for every comparison below are in DECISIONS.md and HISTORY.md.

vs. Result Caveat
llama.cpp ~1.3–1.4x faster (4.71–5.05s vs. 6.45–6.56s) Both AOT-compiled — doesn't exercise the JIT-tax claim below
vLLM ~24–52x faster (4.71–5.05s vs. 121–244s, depending ontorch.compile cache state) Installed vLLM has no GGUF support; ran against an HF safetensors checkpoint instead, disclosed
Ollama Directly competitive when it doesn't stall (~6–7s), but its bundled llama-server intermittently hits an internal GPU-discovery-watchdog timeout (~55–62s) Wraps llama.cpp's own runtime — tests packaging/daemon overhead, not the AOT-vs-JIT bet
TypeSafe Jev , cold-start-to-decision Reflex loses, ~10–60x slower Different deployment model: Jev is an always-warm managed API; this measures a genuine cold local process launch
TypeSafe Jev , warm/compute-only Competitive, within ~1.3–2x (19.4ms vs. Jev's cited 10–15ms) Jev's figures are self-reported/published, not independently reproduced here

The llama.cpp/Ollama/Jev "loses" results above are reported as-is, not smoothed over — see DECISIONS.md's benchmark-methodology entries for why each comparison is framed the way it is.

Every CUDA kernel is compiled ahead of time (build.rs invokes nvcc, see build.rs and src/kernels_cuda/), never at runtime via NVRTC. src/aot.rs loads the precompiled PTX/cubin at process start via the CUDA driver API. Default mode emits portable PTX (small driver-side JIT-to-SASS cost at load); set REFLEX_CUDA_ARCH=sm_XX to compile straight to a cubin for one target architecture (true zero-JIT, at the cost of needing a matching cubin per deployment target). Which one actually wins on real hardware is unverified — that's the first thing to measure, not assume.

Run cargo run --bin reflex -- smoke on a real GPU instance as the very first real-hardware step: it proves the AOT pipeline works end to end and reports actual process-start-to-first-result wall clock on the simplest possible kernel, before any model-architecture work begins.

Linux build prerequisite for --features download/ ipc/ python (--all-features included): these pull in hf-hub, whose ureq HTTP client needs libssl-dev + pkg-config on the build host, or cargo build fails with openssl-sys unable to find an OpenSSL installation. Not needed for the default feature-less build. On Ubuntu/Debian:

sudo apt-get install -y libssl-dev pkg-config

(Discovered on a fresh ThunderCompute instance during the MVP-release adoption round — not needed on the Windows dev machine that round otherwise developed on, since native-tls uses a different TLS backend there.)

All four steps below are done and real-hardware-verified (see STATUS.md for current state, HISTORY.md for the full verification write-up of each):

  1. Dense Qwen3 — the best-understood, most well-documented architecture to build against first; proves the AOT-compilation + cold-start-benchmark harness works at all.
  2. Qwen3-MoE
  3. Qwen3.5 hybrid Gated DeltaNet mixer
  4. DeepSeek-V2/V3 MLA — deliberately last; a genuinely different (compressed latent-KV) caching strategy, not an incremental GQA extension.

These are permanent constraints on this engine, not just current-MVP scope — the whole reason Reflex exists is to win a narrower bet (cold-start energy/latency) than sustained-server throughput. A broad serving feature set re-inherits the exact throughput/serving race that's unwinnable against llama.cpp/vLLM/SGLang's head start. Multi-tenancy and persistent state belong in the host orchestrator, not in this engine:

  • batch_size is always 1. No request queue, no continuous batching, no PagedAttention-style dynamic allocation, no context preemption. Horizontal scaling (many concurrent jobs) is the orchestrator's job — spin up NReflex processes across GPU slices/time-slices — not this engine's, ever.
  • No internal multi-tenant LoRA router/scheduler.
  • No internal NVMe/S3 KV-cache manager or cache-hit logic.
  • No concurrent HTTP/gRPC server , no request auth/rate-limiting, no autoscaling decision-making. If a warm-context mode ever exists (see Phase 4 below), it accepts one job at a time, strictly sequentially — never a thread pool.

No in-core HTTP/gRPC server, ever, not deferred. This is not a separate exception to the rule above — a concurrent HTTP listener is the exact same violation ("batch_size always 1... never a thread pool") under a different name, and an adoption/UX ask asking for one doesn't get to reopen it. If HTTP access to this engine is ever genuinely needed, the pattern is a separate, optional sidecar binary (e.g. system1-openai-adapter) that talks to this core engine over local IPC only — the core engine itself never grows a network socket. Building that sidecar is out of scope for now; this paragraph only records the escape-hatch pattern so a future HTTP ask gets routed there instead of back into this engine.

For local, non-network ergonomics, this engine may instead expose: a stdio JSON-line mode (reflex stdio, one JSON request per stdin line, fully processed before the next line is read) and a Unix Domain Socket mode (reflex uds <path>, Unix-only, one connection fully processed before the next is accepted) — both strictly sequential, never a thread pool, mirroring the same request/response protocol. A shared-memory ring-buffer transport was considered and deliberately deferred — crash-safety and synchronization design is disproportionate complexity for the ergonomics it would buy — recorded here as a future-work idea only, not designed.

Once the model-architecture MVP proves the engine handles the target model families at all, the next axis is making the cold-start path itself faster and adoptable — without ever crossing into building a serving platform. The framing: let vLLM win the warm-throughput race; Reflex wins by being the fastest way to turn cold compute into one output token, then getting out of the way. All four phases below are done — see HISTORY.md for the full per-round write-up of each:

  • Phase 1 — Single-shot CLI: process launch -> one forward pass -> exit.

  • Phase 2 — Fast IO : weights upload to the GPU once (not re-uploaded per kernel call), device-resident activations through a whole layer, and on-GPU dequant kernels for the block types this project's fixtures use for the bulk of weight bytes. This is the work behind the llama.cpp benchmark result above.

  • Phase 3 — State I/O :--export-kv <file> /--import-kv <file> for raw K/V-cache dump/load/resume, across all four architectures. Reflex stays ignorant ofwhere that file lives (NVMe, an S3-backed FUSE mount, tmpfs) — that's the orchestrator's job, not this engine's.

  • Phase 4 — Embeddability :--lora <path> (load-time adapter application, no runtime hot-swap multiplexer) and a Rust C-FFI surface (src/ffi.rs ,include/reflex_engine.h ) so an external orchestrator can embed Reflex directly instead ofexec -ing a binary.

  • src/gguf.rs — GGUF metadata/tensor-directory parsing (mmap-based).

  • src/dequant.rs ,src/dequant_iq.rs ,src/dequant_iq_tables.rs — standard and i-quant dequantization, verified byte-exact againstgguf-py .

  • src/tokenizer.rs — verified against real sentencepiece/BPE references.

None of these care how kernels get compiled — they're pure host-side GGUF/tokenizer logic. There is deliberately no NVRTC runtime-compile-and-load path anywhere in this codebase — that's the thing this project's AOT design replaces, not reuses.

A multi-stage Dockerfile is included: the builder stage has the full CUDA devel toolkit (nvcc) to compile the AOT kernels; the runtime stage only needs the CUDA runtime libraries, since every kernel byte is embedded directly into the compiled binary at build time — the runtime image never runs nvcc and never needs the devel toolkit.

docker build --build-arg REFLEX_CUDA_ARCH=sm_86 -t reflex .

docker run --rm --gpus all -v /path/to/models:/models \
  reflex /models/Qwen3-0.6B-Q4_K_M.gguf "Once upon a time"

For any other subcommand (system1/ smoke/ bench/ check/ stdio/ uds), override the entrypoint:

docker run --rm --gpus all -v /path/to/models:/models \
  --entrypoint /usr/local/bin/reflex reflex \
  system1 /models/Qwen3-0.6B-Q4_K_M.gguf "Q: ...? A:" --candidate " Yes" --candidate " No"

The CUDA major/minor version in both Docker stages must stay consistent with Cargo.toml's pinned cudarc feature ("cuda-12000", i.e. CUDA 12.x) — a mismatch is a build-time/runtime library version mismatch this Dockerfile can't catch for you.

Reflex is a single-shot CLI, not a server (see Non-goals above) — the natural Kubernetes primitive is a Job, one cold-start invocation per Pod, never a Deployment/ Service. A minimal example running generate against a GGUF baked into a volume, requesting one GPU via the standard NVIDIA device plugin:

apiVersion: batch/v1
kind: Job
metadata:
  name: reflex-generate
spec:
  backoffLimit: 0
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: reflex
          image: reflex:latest
          args: ["/models/Qwen3-0.6B-Q4_K_M.gguf", "Once upon a time"]
          resources:
            limits:
              nvidia.com/gpu: 1
          volumeMounts:
            - name: models
              mountPath: /models
              readOnly: true
      volumes:
        - name: models
          persistentVolumeClaim:
            claimName: reflex-models

For system1/ bench/ check/other subcommands, set command: ["/usr/local/bin/reflex"] and put the subcommand as the first entry in args, same as the Docker override above. This is exactly the "orchestrator's job" this engine intentionally stays out of — Reflex itself never grows a scheduler, a request queue, or a batch_size > 1; Kubernetes (or cron, or a FaaS platform) is where that concurrency/scheduling belongs.

── more in #ai-infrastructure 4 stories · sorted by recency
── more on @reflex 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/show-hn-reflex-a-ggu…] indexed:0 read:11min 2026-09-23 ·