{"slug": "fast-single-box-deepseek-v4-1-flash-runtime", "title": "Fast single-box DeepSeek v4.1 Flash runtime", "summary": "A new recipe for running DeepSeek V4.1 Flash on a single AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives) reports roughly 450 tokens per second prefill and 15 tokens per second decode with DSpark, according to the runtime's author. The document specifies a memory-bounded, SSD-streaming serving runtime built on the antirez/ds4 codebase and references the DeepSeek-V4.1-Flash technical report (arXiv 2609.19969), the DSpark report (arXiv 2607.05147), and Mixture of Cache-Conditional Experts (arXiv 2412.00099). The author notes some max-performance numbers in the docs are outdated because benchmarking agents sometimes failed to record their actual results.", "body_md": "# Recipe: a fast single-box DeepSeek V4.1 Flash runtime[#](#recipe-a-fast-single-box-deepseek-v41-flash-runtime)\n\nfor humans: this repo is intended to be read by your agent, not you. it is pure claudeslop, but it does contain everything you need to get an extremely fast deepseek v4.1 flash up and running on your hardware. expect to use quite a lot of tokens on the runtime, and ideally use something like opus 5.5 since it's extremely nice at making llm runtimes, quite unsurprisingly.\n\nask your agent to guide you through, this repo contains a guide for them on where and how to start. make sure to tell them what you're actually going to use the runtime for, since a lot of the decisions are about hyper-specialization, therefore the end consumer determines whether some optimizations are applicable or not. also i recommend discussing your tolerance for accuracy-affecting changes, from \"should match upstream logits as much as possible\" to \"eh, a bit is fine if the model doesn't go insane\".\n\nimo it's a good idea to use the [https://github.com/antirez/ds4](https://github.com/antirez/ds4) codebase as a starting point, and then just shaping it to fit this doc. that's what i did, anyways.\n\nsome max performance numbers in the docs are outdated since the agents benchmarking my runtime sometimes forgot to write their actual results down, but the perf i'm getting rn is around 450tps perfill and 15tps decode (with DSpark).\n\nFor LLMs:\n\nThis folder is a specification of ideas, not a blueprint. It covers everything that turns a straightforward DeepSeek V4.1 Flash implementation into a fast, memory-bounded, SSD-streaming serving runtime on one AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives). For each piece it says what was built, why, how it works, how much it measurably bought, under which conditions, and which traps cost us time.\n\nIt deliberately names no functions, files, settings or module boundaries. Two implementations built from this recipe can look nothing alike internally and still be equally correct and equally fast.\n\n## Who this is for[#](#who-this-is-for)\n\nA capable team of implementing agents (or humans) building a new runtime from scratch, possibly on different hardware. Besides this folder you need:\n\n| Reference | What it defines | \n|---|---|\n| Hugging Face repo `deepseek-ai/DeepSeek-V4.1-Flash` | Weights, config, tokenizer and chat template, reference inference code (including the vision tower and image processor) | \n| DeepSeek-V4.1-Flash technical report, [arXiv 2609.19969](https://arxiv.org/abs/2609.19969) | Architecture: Causal Encoder-Decoder (CED, §2.2), CSA2 with cross-layer KV/index reuse and the hierarchical sparse indexer (§2.3), single-pass mHC, Engram, DSpark, FP4 main KV (§2.4), persistent KV management and SWA Bounded Replay (§3.2), multimodal input (§2.1.1) | \n| DSpark report, [arXiv 2607.05147](https://arxiv.org/abs/2607.05147) | Semi-autoregressive draft, confidence head and post-hoc calibration, hardware-aware prefix scheduler (§3.2, §5.2) | \n| Mixture of Cache-Conditional Experts, [arXiv 2412.00099](https://arxiv.org/abs/2412.00099) | Cache-aware routing for MoE inference when only part of the experts fit in memory | \n| Suffix Cache Reuse, [`facebookresearch/context-language-models/suffix_cache_reuse`](https://github.com/facebookresearch/context-language-models/tree/main/suffix_cache_reuse) | Reusing cached KV for text that reappears at a different position | \n| antirez/ds4 | [github.com/antirez/ds4](https://github.com/antirez/ds4) | \n| DeepSeek's request/encoding toolkit linked from the Hugging Face model card (\"deepseek-recipe\") | The production image preprocessing that 14-vision.md matches | \n\nThe references define the model. This recipe documents only the deltas: it never re-specifies the architecture, and it summarizes a baseline detail only when a delta would be unreadable without it.\n\n## Headline results on the reference host[#](#headline-results-on-the-reference-host)\n\nAll numbers are measurements on the reference machine, with context. They\ncalibrate expectations; they are not promises for other hardware. Full table:\n`90-results-ledger.md`.\n\n| Quantity | Value | Context | \n|---|---|---|\n| Cold prefill | ~354–391 tok/s | 100k real text 354 tok/s; 32k cold agent prompt 86.0 s (381 tok/s); 32k book with the long-context selection levers 391 tok/s | \n| Decode, DSpark default vs ordinary | 11.89 vs 8.74 tok/s (code), 9.67 vs 9.91 (narrative) | 512-step greedy, last matched comparison (at promotion). Later verification-path work was screened DSpark vs DSpark only (its greedy screens compound to roughly +25% code / +23% prose); the current margin over ordinary decode was not re-measured | \n| Decode in deep agent turns | ~11–13 tok/s | 400–560k context, real agent turns, DSpark on | \n| Prompt reuse | 98.75% of prompt tokens served from KV | one week of production traffic (disk KV cache + suffix cache reuse + live sessions) | \n| Memory | one 116.46e9-byte target | everything mandatory admitted first; ~5,000 of 15,360 routed experts resident | \n\n## How to use this recipe[#](#how-to-use-this-recipe)\n\n1. **Read `00-start-here-environment.md` first and do what it says.** Probe your\nown machine, then ask the user what resources they have and want the\nruntime to use (memory budget, drives, context length, sessions, features,\naccuracy policy). Many techniques here exist because of one specific\nhardware ratio. The decision matrix there tells you which ones apply to you\nand how to re-derive their parameters from your measurements. Build only\nwhat applies.\n2. **Read `01-correctness-and-measurement.md` before writing any optimization.** Every speedup here was admitted through the exactness and measurement\ndiscipline it describes. Without it the numbers in this folder are not\nreproducible and many \"wins\" are noise.\n3. **Build a correct, slow baseline from the references and keep it.** It is\nthe oracle every fast path is compared against, bit for bit, for the rest of\nthe project.\n4. **Implement the subsystem docs in dependency order** (below). Each doc lists\nits prerequisites, every technique's measured value and exactness class,\nhow to validate it, its pitfalls, what is still unknown, and what was tried\nwithout paying.\n\n## Dependency order and swarm split[#](#dependency-order-and-swarm-split)\n\n``` php\ngraph TD\n  E[00 environment] --> M[01 correctness & measurement]\n  E --> P[02 platform]\n  M --> W[03 weights & formats]\n  P --> W\n  W --> S[04 storage IO]\n  W --> L[05 memory ledger]\n  S --> C[06 expert cache & routing]\n  L --> C\n  C --> PF[07 prefill pipeline]\n  C --> D[09 decode kernels]\n  A[08 attention & selection] --> PF\n  A --> D\n  D --> DS[10 DSpark]\n  L --> DS\n  PF --> KV[12 disk KV cache]\n  KV --> SCR[13 suffix cache reuse]\n  DS --> SV[11 multi-session serving]\n  KV --> SV\n  L --> SV\n  SV --> V[14 vision]\n  SV --> SM[15 sampling & logit bias]\n  SV --> T[16 telemetry, optional]\n```\n\nReasonable parallel slices for a swarm, once 00–03 are settled and the reference oracle runs:\n\n| Slice | Docs | Shared contract it must honor | \n|---|---|---|\n| Storage and residency | 04, 05, 06 | The memory ledger is the only source of slot counts; expert residency is a deterministic function of ledger arithmetic and request history | \n| Kernels | 07, 08, 09 | Every kernel is gated bit-exact against the reference path chosen for its phase; summation orders are written down and pinned | \n| Speculation | 10 | Verification reproduces single-token decode arithmetic bit for bit; draft memory is mandatory ledger state | \n| Persistence and reuse | 12, 13 | Checkpoints are self-contained and validated before anything is mutated; reuse never silently changes arithmetic without a tagged approval | \n| Serving | 11, 15 | No active request is preempted; per-request state (bias set, routing transaction, images) is pinned for the request's lifetime | \n| Vision (required) | 14 | Image rows enter the same ledger, KV keys, checkpoints, priority queue and prefill yields as text; the tower is validated against the reference modules before any speed work | \n| Optional observability | 16 | Observation never changes arithmetic; capture copies from existing boundaries | \n\nIntegrate early: the hardest bugs in this project were disagreements between two paths that each looked correct alone (verification vs decode fold orders, restore vs fresh, batched vs single member).\n\n## Conventions[#](#conventions)\n\n- **Exactness tags** on every technique (canonical definitions in 01):\n  - `[exact]` — outputs bit-identical to the reference configuration:\nfull-vocabulary logits byte-equal at every step, identical token IDs,\nidentical KV/checkpoint bytes, identical work counters.\n  - `[reference-defining]` — a free choice of floating-point evaluation order,\nlibrary kernel or GEMM plan, made once (e.g. a K-split geometry, an expert\nfold order, a library GEMM on prepared BF16 weights). Not a quality\ntradeoff; once chosen it*is* the reference, and every path that must agree\nreproduces it bit for bit. Timed GEMM plan selection is reference-defining\nper process: recommended practice is timed plans in production and pinned\nplans for every experiment and byte-equality gate.\n  - `[approx — approved]` — deliberately changes outputs inside an accepted\nquality budget; the doc states how quality was validated and what moved\n(cache-conditional routing, shared residency in batched rounds, suffix\ncache reuse, crop + SWA bounded replay, logit/phrase bias).\n  - `[policy]` — changes what work happens when or where (memory, scheduling,\ncaching, IO) without changing arithmetic. Caution: with cache-conditional\nrouting on, expert residency feeds routing, so cache and memory policy can\nchange outputs.\n- **Prefill and decode are separate arithmetic paths.** Each is judged against\nits own reference. Verification rows for speculative decoding belong to the\ndecode path.\n- **Vision is required, not optional.** V4.1 is a multimodal model; every\nbuild accepts images from its first serving release (14-vision.md).**DSpark is on by default** : build it unless the owner explicitly rejects it\n(10-dspark-speculative-decoding.md). Only the telemetry/dashboard doc (16) is\noptional as a whole; other feature choices are settled with the owner in 00.\n- **Numbers** are measurements from the reference host with their context\n(workload, metric, before → after). Unmeasured claims are marked. Throughput\nthere depends on text, expert misses, drive temperature and runtime\nvariation. Measurements taken before a 2026-09-16 correction used a\nsynthetic repetitive prompt that flattered decode by ~11% and prefill by ~7%;\ndocs label such figures.\n- **\"Strix Halo specific\"** marks facts that only hold on this platform; each is\npaired with the principle behind it and how to re-measure it on yours.\n- **Pitfalls** sections hold correctness traps, incidents and design-forcing\nmeasurements. They are the most expensive knowledge in this folder.\n- **Known unknowns** sections list what was never measured or settled, and the\nmeasurement that would settle it.\n- **\"Tried here, did not pay (low trust)\"** appendices list experiments that did\nnot help in our configuration. Many were measured once, some under conditions\nlater shown to be wrong. They are not proofs that an idea fails.\n\n## Document map[#](#document-map)\n\n| Doc | Subject | \n|---|---|\n| `00-start-here-environment.md` | Environment discovery, user resource interview, decision matrix, parameter derivation | \n| `01-correctness-and-measurement.md` | Exactness doctrine and tags, validation gates, benchmark method, determinism, quality studies for approved approximations, experiment safety | \n| `02-platform-strix-halo.md` | gfx1151 / UMA / GTT / OS / toolchain specifics and their generalizations | \n| `03-weights-and-formats.md` | Converting the checkpoint into runtime stores: precision, layout, drive placement, prepared expansions, lossless head packing, identity | \n| `04-storage-io.md` | N-drive replicated weight store: direct async reads, scheduling, rate estimation, hedging, failure and unplug model, Engram reads | \n| `05-memory-ledger.md` | One byte target, mandatory vs elastic state, phase budgets, tensor pool, admission and reclaim, multi-session memory | \n| `06-expert-cache-and-routing.md` | Expert slot cache, eviction keys, pinning, near-pick eviction, cache-conditional routing, miss handling, staged prefetch | \n| `07-prefill-pipeline.md` | Chunk cohorts and leases, CED scheduling, prefill GEMMs and MoE kernels, overlaps, host routing, long-context prefill | \n| `08-attention-index-selection.md` | Exact wide attention tiles, decode attention, CSA2 index scoring and top-k selection, candidate pools, long-context levers | \n| `09-decode-kernels.md` | Single-token decode: byte model, matvec geometries, fusions, routed-expert kernels, head, step composition | \n| `10-dspark-speculative-decoding.md` | Decode-exact verification, exact multi-row kernels, confidence/cost scheduling and calibration, draft residency, verify-round overlap | \n| `11-multi-session-serving.md` | Slots, priority, request lifecycle, fair prefill sharing, nested work in yields, batched multi-session rounds, cancellation, embeddings | \n| `12-kv-disk-cache.md` | Persistent on-disk KV cache: keys, anchoring, delta checkpoints, eviction, restore validation, images, crop + replay reuse | \n| `13-suffix-cache-reuse.md` | SCR on V4.1: relocation, RoPE master, splicing, planner and cost model, multi-source import, quality | \n| `14-vision.md` | Vision tower kernels, preprocessing, encoder lifecycle, embedding caches, image identity and spans | \n| `15-sampling-logit-bias-and-tool-calls.md` | Sampling defaults, logit bias, phrase entries and backtrack, prompt/tool-call rendering, optional tool-call guards, output encoding | \n| `16-optional-telemetry-dashboard.md` | Optional: telemetry store, dashboard, structured logs, exact-boundary internal-state capture | \n| `90-results-ledger.md` | Every adopted technique with its tag and measured effect, in one place |", "url": "https://wpnews.pro/news/fast-single-box-deepseek-v4-1-flash-runtime", "canonical_source": "https://tangled.org/astrra.space/ds4-recipe", "published_at": "2026-10-07 10:59:14+00:00", "updated_at": "2026-10-07 11:20:02.369339+00:00", "lang": "en", "topics": ["large-language-models", "ai-infrastructure", "mlops", "ai-tools", "ai-research"], "entities": ["DeepSeek V4.1 Flash", "AMD Strix Halo", "antirez/ds4", "DSpark", "DeepSeek", "arXiv 2609.19969", "arXiv 2607.05147", "gfx1151"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/fast-single-box-deepseek-v4-1-flash-runtime", "markdown": "https://wpnews.pro/news/fast-single-box-deepseek-v4-1-flash-runtime.md", "text": "https://wpnews.pro/news/fast-single-box-deepseek-v4-1-flash-runtime.txt", "jsonld": "https://wpnews.pro/news/fast-single-box-deepseek-v4-1-flash-runtime.jsonld"}}