# Fast single-box DeepSeek v4.1 Flash runtime

> Source: <https://tangled.org/astrra.space/ds4-recipe>
> Published: 2026-10-07 10:59:14+00:00

# Recipe: a fast single-box DeepSeek V4.1 Flash runtime[#](#recipe-a-fast-single-box-deepseek-v41-flash-runtime)

for humans: this repo is intended to be read by your agent, not you. it is pure claudeslop, but it does contain everything you need to get an extremely fast deepseek v4.1 flash up and running on your hardware. expect to use quite a lot of tokens on the runtime, and ideally use something like opus 5.5 since it's extremely nice at making llm runtimes, quite unsurprisingly.

ask your agent to guide you through, this repo contains a guide for them on where and how to start. make sure to tell them what you're actually going to use the runtime for, since a lot of the decisions are about hyper-specialization, therefore the end consumer determines whether some optimizations are applicable or not. also i recommend discussing your tolerance for accuracy-affecting changes, from "should match upstream logits as much as possible" to "eh, a bit is fine if the model doesn't go insane".

imo it's a good idea to use the [https://github.com/antirez/ds4](https://github.com/antirez/ds4) codebase as a starting point, and then just shaping it to fit this doc. that's what i did, anyways.

some max performance numbers in the docs are outdated since the agents benchmarking my runtime sometimes forgot to write their actual results down, but the perf i'm getting rn is around 450tps perfill and 15tps decode (with DSpark).

For LLMs:

This folder is a specification of ideas, not a blueprint. It covers everything that turns a straightforward DeepSeek V4.1 Flash implementation into a fast, memory-bounded, SSD-streaming serving runtime on one AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives). For each piece it says what was built, why, how it works, how much it measurably bought, under which conditions, and which traps cost us time.

It deliberately names no functions, files, settings or module boundaries. Two implementations built from this recipe can look nothing alike internally and still be equally correct and equally fast.

## Who this is for[#](#who-this-is-for)

A capable team of implementing agents (or humans) building a new runtime from scratch, possibly on different hardware. Besides this folder you need:

| Reference | What it defines | 
|---|---|
| Hugging Face repo `deepseek-ai/DeepSeek-V4.1-Flash` | Weights, config, tokenizer and chat template, reference inference code (including the vision tower and image processor) | 
| DeepSeek-V4.1-Flash technical report, [arXiv 2609.19969](https://arxiv.org/abs/2609.19969) | Architecture: Causal Encoder-Decoder (CED, §2.2), CSA2 with cross-layer KV/index reuse and the hierarchical sparse indexer (§2.3), single-pass mHC, Engram, DSpark, FP4 main KV (§2.4), persistent KV management and SWA Bounded Replay (§3.2), multimodal input (§2.1.1) | 
| DSpark report, [arXiv 2607.05147](https://arxiv.org/abs/2607.05147) | Semi-autoregressive draft, confidence head and post-hoc calibration, hardware-aware prefix scheduler (§3.2, §5.2) | 
| Mixture of Cache-Conditional Experts, [arXiv 2412.00099](https://arxiv.org/abs/2412.00099) | Cache-aware routing for MoE inference when only part of the experts fit in memory | 
| Suffix Cache Reuse, [`facebookresearch/context-language-models/suffix_cache_reuse`](https://github.com/facebookresearch/context-language-models/tree/main/suffix_cache_reuse) | Reusing cached KV for text that reappears at a different position | 
| antirez/ds4 | [github.com/antirez/ds4](https://github.com/antirez/ds4) | 
| DeepSeek's request/encoding toolkit linked from the Hugging Face model card ("deepseek-recipe") | The production image preprocessing that 14-vision.md matches | 

The references define the model. This recipe documents only the deltas: it never re-specifies the architecture, and it summarizes a baseline detail only when a delta would be unreadable without it.

## Headline results on the reference host[#](#headline-results-on-the-reference-host)

All numbers are measurements on the reference machine, with context. They
calibrate expectations; they are not promises for other hardware. Full table:
`90-results-ledger.md`.

| Quantity | Value | Context | 
|---|---|---|
| Cold prefill | ~354–391 tok/s | 100k real text 354 tok/s; 32k cold agent prompt 86.0 s (381 tok/s); 32k book with the long-context selection levers 391 tok/s | 
| Decode, DSpark default vs ordinary | 11.89 vs 8.74 tok/s (code), 9.67 vs 9.91 (narrative) | 512-step greedy, last matched comparison (at promotion). Later verification-path work was screened DSpark vs DSpark only (its greedy screens compound to roughly +25% code / +23% prose); the current margin over ordinary decode was not re-measured | 
| Decode in deep agent turns | ~11–13 tok/s | 400–560k context, real agent turns, DSpark on | 
| Prompt reuse | 98.75% of prompt tokens served from KV | one week of production traffic (disk KV cache + suffix cache reuse + live sessions) | 
| Memory | one 116.46e9-byte target | everything mandatory admitted first; ~5,000 of 15,360 routed experts resident | 

## How to use this recipe[#](#how-to-use-this-recipe)

1. **Read `00-start-here-environment.md` first and do what it says.** Probe your
own machine, then ask the user what resources they have and want the
runtime to use (memory budget, drives, context length, sessions, features,
accuracy policy). Many techniques here exist because of one specific
hardware ratio. The decision matrix there tells you which ones apply to you
and how to re-derive their parameters from your measurements. Build only
what applies.
2. **Read `01-correctness-and-measurement.md` before writing any optimization.** Every speedup here was admitted through the exactness and measurement
discipline it describes. Without it the numbers in this folder are not
reproducible and many "wins" are noise.
3. **Build a correct, slow baseline from the references and keep it.** It is
the oracle every fast path is compared against, bit for bit, for the rest of
the project.
4. **Implement the subsystem docs in dependency order** (below). Each doc lists
its prerequisites, every technique's measured value and exactness class,
how to validate it, its pitfalls, what is still unknown, and what was tried
without paying.

## Dependency order and swarm split[#](#dependency-order-and-swarm-split)

``` php
graph TD
  E[00 environment] --> M[01 correctness & measurement]
  E --> P[02 platform]
  M --> W[03 weights & formats]
  P --> W
  W --> S[04 storage IO]
  W --> L[05 memory ledger]
  S --> C[06 expert cache & routing]
  L --> C
  C --> PF[07 prefill pipeline]
  C --> D[09 decode kernels]
  A[08 attention & selection] --> PF
  A --> D
  D --> DS[10 DSpark]
  L --> DS
  PF --> KV[12 disk KV cache]
  KV --> SCR[13 suffix cache reuse]
  DS --> SV[11 multi-session serving]
  KV --> SV
  L --> SV
  SV --> V[14 vision]
  SV --> SM[15 sampling & logit bias]
  SV --> T[16 telemetry, optional]
```

Reasonable parallel slices for a swarm, once 00–03 are settled and the reference oracle runs:

| Slice | Docs | Shared contract it must honor | 
|---|---|---|
| Storage and residency | 04, 05, 06 | The memory ledger is the only source of slot counts; expert residency is a deterministic function of ledger arithmetic and request history | 
| Kernels | 07, 08, 09 | Every kernel is gated bit-exact against the reference path chosen for its phase; summation orders are written down and pinned | 
| Speculation | 10 | Verification reproduces single-token decode arithmetic bit for bit; draft memory is mandatory ledger state | 
| Persistence and reuse | 12, 13 | Checkpoints are self-contained and validated before anything is mutated; reuse never silently changes arithmetic without a tagged approval | 
| Serving | 11, 15 | No active request is preempted; per-request state (bias set, routing transaction, images) is pinned for the request's lifetime | 
| Vision (required) | 14 | Image rows enter the same ledger, KV keys, checkpoints, priority queue and prefill yields as text; the tower is validated against the reference modules before any speed work | 
| Optional observability | 16 | Observation never changes arithmetic; capture copies from existing boundaries | 

Integrate early: the hardest bugs in this project were disagreements between two paths that each looked correct alone (verification vs decode fold orders, restore vs fresh, batched vs single member).

## Conventions[#](#conventions)

- **Exactness tags** on every technique (canonical definitions in 01):
  - `[exact]` — outputs bit-identical to the reference configuration:
full-vocabulary logits byte-equal at every step, identical token IDs,
identical KV/checkpoint bytes, identical work counters.
  - `[reference-defining]` — a free choice of floating-point evaluation order,
library kernel or GEMM plan, made once (e.g. a K-split geometry, an expert
fold order, a library GEMM on prepared BF16 weights). Not a quality
tradeoff; once chosen it*is* the reference, and every path that must agree
reproduces it bit for bit. Timed GEMM plan selection is reference-defining
per process: recommended practice is timed plans in production and pinned
plans for every experiment and byte-equality gate.
  - `[approx — approved]` — deliberately changes outputs inside an accepted
quality budget; the doc states how quality was validated and what moved
(cache-conditional routing, shared residency in batched rounds, suffix
cache reuse, crop + SWA bounded replay, logit/phrase bias).
  - `[policy]` — changes what work happens when or where (memory, scheduling,
caching, IO) without changing arithmetic. Caution: with cache-conditional
routing on, expert residency feeds routing, so cache and memory policy can
change outputs.
- **Prefill and decode are separate arithmetic paths.** Each is judged against
its own reference. Verification rows for speculative decoding belong to the
decode path.
- **Vision is required, not optional.** V4.1 is a multimodal model; every
build accepts images from its first serving release (14-vision.md).**DSpark is on by default** : build it unless the owner explicitly rejects it
(10-dspark-speculative-decoding.md). Only the telemetry/dashboard doc (16) is
optional as a whole; other feature choices are settled with the owner in 00.
- **Numbers** are measurements from the reference host with their context
(workload, metric, before → after). Unmeasured claims are marked. Throughput
there depends on text, expert misses, drive temperature and runtime
variation. Measurements taken before a 2026-09-16 correction used a
synthetic repetitive prompt that flattered decode by ~11% and prefill by ~7%;
docs label such figures.
- **"Strix Halo specific"** marks facts that only hold on this platform; each is
paired with the principle behind it and how to re-measure it on yours.
- **Pitfalls** sections hold correctness traps, incidents and design-forcing
measurements. They are the most expensive knowledge in this folder.
- **Known unknowns** sections list what was never measured or settled, and the
measurement that would settle it.
- **"Tried here, did not pay (low trust)"** appendices list experiments that did
not help in our configuration. Many were measured once, some under conditions
later shown to be wrong. They are not proofs that an idea fails.

## Document map[#](#document-map)

| Doc | Subject | 
|---|---|
| `00-start-here-environment.md` | Environment discovery, user resource interview, decision matrix, parameter derivation | 
| `01-correctness-and-measurement.md` | Exactness doctrine and tags, validation gates, benchmark method, determinism, quality studies for approved approximations, experiment safety | 
| `02-platform-strix-halo.md` | gfx1151 / UMA / GTT / OS / toolchain specifics and their generalizations | 
| `03-weights-and-formats.md` | Converting the checkpoint into runtime stores: precision, layout, drive placement, prepared expansions, lossless head packing, identity | 
| `04-storage-io.md` | N-drive replicated weight store: direct async reads, scheduling, rate estimation, hedging, failure and unplug model, Engram reads | 
| `05-memory-ledger.md` | One byte target, mandatory vs elastic state, phase budgets, tensor pool, admission and reclaim, multi-session memory | 
| `06-expert-cache-and-routing.md` | Expert slot cache, eviction keys, pinning, near-pick eviction, cache-conditional routing, miss handling, staged prefetch | 
| `07-prefill-pipeline.md` | Chunk cohorts and leases, CED scheduling, prefill GEMMs and MoE kernels, overlaps, host routing, long-context prefill | 
| `08-attention-index-selection.md` | Exact wide attention tiles, decode attention, CSA2 index scoring and top-k selection, candidate pools, long-context levers | 
| `09-decode-kernels.md` | Single-token decode: byte model, matvec geometries, fusions, routed-expert kernels, head, step composition | 
| `10-dspark-speculative-decoding.md` | Decode-exact verification, exact multi-row kernels, confidence/cost scheduling and calibration, draft residency, verify-round overlap | 
| `11-multi-session-serving.md` | Slots, priority, request lifecycle, fair prefill sharing, nested work in yields, batched multi-session rounds, cancellation, embeddings | 
| `12-kv-disk-cache.md` | Persistent on-disk KV cache: keys, anchoring, delta checkpoints, eviction, restore validation, images, crop + replay reuse | 
| `13-suffix-cache-reuse.md` | SCR on V4.1: relocation, RoPE master, splicing, planner and cost model, multi-source import, quality | 
| `14-vision.md` | Vision tower kernels, preprocessing, encoder lifecycle, embedding caches, image identity and spans | 
| `15-sampling-logit-bias-and-tool-calls.md` | Sampling defaults, logit bias, phrase entries and backtrack, prompt/tool-call rendering, optional tool-call guards, output encoding | 
| `16-optional-telemetry-dashboard.md` | Optional: telemetry store, dashboard, structured logs, exact-boundary internal-state capture | 
| `90-results-ledger.md` | Every adopted technique with its tag and measured effect, in one place |
