cd /news/large-language-models/fast-single-box-deepseek-v4-1-flash-… · home › topics › large-language-models › article
[ARTICLE · art-146752] src=tangled.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Fast single-box DeepSeek v4.1 Flash runtime

A new recipe for running DeepSeek V4.1 Flash on a single AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives) reports roughly 450 tokens per second prefill and 15 tokens per second decode with DSpark, according to the runtime's author. The document specifies a memory-bounded, SSD-streaming serving runtime built on the antirez/ds4 codebase and references the DeepSeek-V4.1-Flash technical report (arXiv 2609.19969), the DSpark report (arXiv 2607.05147), and Mixture of Cache-Conditional Experts (arXiv 2412.00099). The author notes some max-performance numbers in the docs are outdated because benchmarking agents sometimes failed to record their actual results.

read10 min views1 publishedOct 7, 2026
Fast single-box DeepSeek v4.1 Flash runtime
Image: source

for humans: this repo is intended to be read by your agent, not you. it is pure claudeslop, but it does contain everything you need to get an extremely fast deepseek v4.1 flash up and running on your hardware. expect to use quite a lot of tokens on the runtime, and ideally use something like opus 5.5 since it's extremely nice at making llm runtimes, quite unsurprisingly.

ask your agent to guide you through, this repo contains a guide for them on where and how to start. make sure to tell them what you're actually going to use the runtime for, since a lot of the decisions are about hyper-specialization, therefore the end consumer determines whether some optimizations are applicable or not. also i recommend discussing your tolerance for accuracy-affecting changes, from "should match upstream logits as much as possible" to "eh, a bit is fine if the model doesn't go insane".

imo it's a good idea to use the https://github.com/antirez/ds4 codebase as a starting point, and then just shaping it to fit this doc. that's what i did, anyways.

some max performance numbers in the docs are outdated since the agents benchmarking my runtime sometimes forgot to write their actual results down, but the perf i'm getting rn is around 450tps perfill and 15tps decode (with DSpark).

For LLMs:

This folder is a specification of ideas, not a blueprint. It covers everything that turns a straightforward DeepSeek V4.1 Flash implementation into a fast, memory-bounded, SSD-streaming serving runtime on one AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives). For each piece it says what was built, why, how it works, how much it measurably bought, under which conditions, and which traps cost us time.

It deliberately names no functions, files, settings or module boundaries. Two implementations built from this recipe can look nothing alike internally and still be equally correct and equally fast.

Who this is for# #

A capable team of implementing agents (or humans) building a new runtime from scratch, possibly on different hardware. Besides this folder you need:

Reference What it defines
Hugging Face repo deepseek-ai/DeepSeek-V4.1-Flash Weights, config, tokenizer and chat template, reference inference code (including the vision tower and image processor)
DeepSeek-V4.1-Flash technical report, arXiv 2609.19969 Architecture: Causal Encoder-Decoder (CED, §2.2), CSA2 with cross-layer KV/index reuse and the hierarchical sparse indexer (§2.3), single-pass mHC, Engram, DSpark, FP4 main KV (§2.4), persistent KV management and SWA Bounded Replay (§3.2), multimodal input (§2.1.1)
DSpark report, arXiv 2607.05147 Semi-autoregressive draft, confidence head and post-hoc calibration, hardware-aware prefix scheduler (§3.2, §5.2)
Mixture of Cache-Conditional Experts, arXiv 2412.00099 Cache-aware routing for MoE inference when only part of the experts fit in memory
Suffix Cache Reuse, facebookresearch/context-language-models/suffix_cache_reuse Reusing cached KV for text that reappears at a different position
antirez/ds4 github.com/antirez/ds4
DeepSeek's request/encoding toolkit linked from the Hugging Face model card ("deepseek-recipe") The production image preprocessing that 14-vision.md matches

The references define the model. This recipe documents only the deltas: it never re-specifies the architecture, and it summarizes a baseline detail only when a delta would be unreadable without it.

Headline results on the reference host# #

All numbers are measurements on the reference machine, with context. They calibrate expectations; they are not promises for other hardware. Full table: 90-results-ledger.md.

Quantity Value Context
Cold prefill ~354–391 tok/s 100k real text 354 tok/s; 32k cold agent prompt 86.0 s (381 tok/s); 32k book with the long-context selection levers 391 tok/s
Decode, DSpark default vs ordinary 11.89 vs 8.74 tok/s (code), 9.67 vs 9.91 (narrative) 512-step greedy, last matched comparison (at promotion). Later verification-path work was screened DSpark vs DSpark only (its greedy screens compound to roughly +25% code / +23% prose); the current margin over ordinary decode was not re-measured
Decode in deep agent turns ~11–13 tok/s 400–560k context, real agent turns, DSpark on
Prompt reuse 98.75% of prompt tokens served from KV one week of production traffic (disk KV cache + suffix cache reuse + live sessions)
Memory one 116.46e9-byte target everything mandatory admitted first; ~5,000 of 15,360 routed experts resident

How to use this recipe# #

  1. Read 00-start-here-environment.md first and do what it says. Probe your own machine, then ask the user what resources they have and want the runtime to use (memory budget, drives, context length, sessions, features, accuracy policy). Many techniques here exist because of one specific hardware ratio. The decision matrix there tells you which ones apply to you and how to re-derive their parameters from your measurements. Build only what applies.
  2. Read 01-correctness-and-measurement.md before writing any optimization. Every speedup here was admitted through the exactness and measurement discipline it describes. Without it the numbers in this folder are not reproducible and many "wins" are noise.
  3. Build a correct, slow baseline from the references and keep it. It is the oracle every fast path is compared against, bit for bit, for the rest of the project.
  4. Implement the subsystem docs in dependency order (below). Each doc lists its prerequisites, every technique's measured value and exactness class, how to validate it, its pitfalls, what is still unknown, and what was tried without paying.

Dependency order and swarm split# #

graph TD
  E[00 environment] --> M[01 correctness & measurement]
  E --> P[02 platform]
  M --> W[03 weights & formats]
  P --> W
  W --> S[04 storage IO]
  W --> L[05 memory ledger]
  S --> C[06 expert cache & routing]
  L --> C
  C --> PF[07 prefill pipeline]
  C --> D[09 decode kernels]
  A[08 attention & selection] --> PF
  A --> D
  D --> DS[10 DSpark]
  L --> DS
  PF --> KV[12 disk KV cache]
  KV --> SCR[13 suffix cache reuse]
  DS --> SV[11 multi-session serving]
  KV --> SV
  L --> SV
  SV --> V[14 vision]
  SV --> SM[15 sampling & logit bias]
  SV --> T[16 telemetry, optional]

Reasonable parallel slices for a swarm, once 00–03 are settled and the reference oracle runs:

Slice Docs Shared contract it must honor
Storage and residency 04, 05, 06 The memory ledger is the only source of slot counts; expert residency is a deterministic function of ledger arithmetic and request history
Kernels 07, 08, 09 Every kernel is gated bit-exact against the reference path chosen for its phase; summation orders are written down and pinned
Speculation 10 Verification reproduces single-token decode arithmetic bit for bit; draft memory is mandatory ledger state
Persistence and reuse 12, 13 Checkpoints are self-contained and validated before anything is mutated; reuse never silently changes arithmetic without a tagged approval
Serving 11, 15 No active request is preempted; per-request state (bias set, routing transaction, images) is pinned for the request's lifetime
Vision (required) 14 Image rows enter the same ledger, KV keys, checkpoints, priority queue and prefill yields as text; the tower is validated against the reference modules before any speed work
Optional observability 16 Observation never changes arithmetic; capture copies from existing boundaries

Integrate early: the hardest bugs in this project were disagreements between two paths that each looked correct alone (verification vs decode fold orders, restore vs fresh, batched vs single member).

Conventions# #

  • Exactness tags on every technique (canonical definitions in 01):
    • [exact] — outputs bit-identical to the reference configuration: full-vocabulary logits byte-equal at every step, identical token IDs, identical KV/checkpoint bytes, identical work counters.
    • [reference-defining] — a free choice of floating-point evaluation order, library kernel or GEMM plan, made once (e.g. a K-split geometry, an expert fold order, a library GEMM on prepared BF16 weights). Not a quality tradeoff; once chosen itis the reference, and every path that must agree reproduces it bit for bit. Timed GEMM plan selection is reference-defining per process: recommended practice is timed plans in production and pinned plans for every experiment and byte-equality gate.
    • [approx — approved] — deliberately changes outputs inside an accepted quality budget; the doc states how quality was validated and what moved (cache-conditional routing, shared residency in batched rounds, suffix cache reuse, crop + SWA bounded replay, logit/phrase bias).
    • [policy] — changes what work happens when or where (memory, scheduling, caching, IO) without changing arithmetic. Caution: with cache-conditional routing on, expert residency feeds routing, so cache and memory policy can change outputs.
  • Prefill and decode are separate arithmetic paths. Each is judged against its own reference. Verification rows for speculative decoding belong to the decode path.
  • Vision is required, not optional. V4.1 is a multimodal model; every build accepts images from its first serving release (14-vision.md).DSpark is on by default : build it unless the owner explicitly rejects it (10-dspark-speculative-decoding.md). Only the telemetry/dashboard doc (16) is optional as a whole; other feature choices are settled with the owner in 00.
  • Numbers are measurements from the reference host with their context (workload, metric, before → after). Unmeasured claims are marked. Throughput there depends on text, expert misses, drive temperature and runtime variation. Measurements taken before a 2026-09-16 correction used a synthetic repetitive prompt that flattered decode by ~11% and prefill by ~7%; docs label such figures.
  • "Strix Halo specific" marks facts that only hold on this platform; each is paired with the principle behind it and how to re-measure it on yours.
  • Pitfalls sections hold correctness traps, incidents and design-forcing measurements. They are the most expensive knowledge in this folder.
  • Known unknowns sections list what was never measured or settled, and the measurement that would settle it.
  • "Tried here, did not pay (low trust)" appendices list experiments that did not help in our configuration. Many were measured once, some under conditions later shown to be wrong. They are not proofs that an idea fails.

Document map# #

Doc Subject
00-start-here-environment.md Environment discovery, user resource interview, decision matrix, parameter derivation
01-correctness-and-measurement.md Exactness doctrine and tags, validation gates, benchmark method, determinism, quality studies for approved approximations, experiment safety
02-platform-strix-halo.md gfx1151 / UMA / GTT / OS / toolchain specifics and their generalizations
03-weights-and-formats.md Converting the checkpoint into runtime stores: precision, layout, drive placement, prepared expansions, lossless head packing, identity
04-storage-io.md N-drive replicated weight store: direct async reads, scheduling, rate estimation, hedging, failure and unplug model, Engram reads
05-memory-ledger.md One byte target, mandatory vs elastic state, phase budgets, tensor pool, admission and reclaim, multi-session memory
06-expert-cache-and-routing.md Expert slot cache, eviction keys, pinning, near-pick eviction, cache-conditional routing, miss handling, staged prefetch
07-prefill-pipeline.md Chunk cohorts and leases, CED scheduling, prefill GEMMs and MoE kernels, overlaps, host routing, long-context prefill
08-attention-index-selection.md Exact wide attention tiles, decode attention, CSA2 index scoring and top-k selection, candidate pools, long-context levers
09-decode-kernels.md Single-token decode: byte model, matvec geometries, fusions, routed-expert kernels, head, step composition
10-dspark-speculative-decoding.md Decode-exact verification, exact multi-row kernels, confidence/cost scheduling and calibration, draft residency, verify-round overlap
11-multi-session-serving.md Slots, priority, request lifecycle, fair prefill sharing, nested work in yields, batched multi-session rounds, cancellation, embeddings
12-kv-disk-cache.md Persistent on-disk KV cache: keys, anchoring, delta checkpoints, eviction, restore validation, images, crop + replay reuse
13-suffix-cache-reuse.md SCR on V4.1: relocation, RoPE master, splicing, planner and cost model, multi-source import, quality
14-vision.md Vision tower kernels, preprocessing, encoder lifecycle, embedding caches, image identity and spans
15-sampling-logit-bias-and-tool-calls.md Sampling defaults, logit bias, phrase entries and backtrack, prompt/tool-call rendering, optional tool-call guards, output encoding
16-optional-telemetry-dashboard.md Optional: telemetry store, dashboard, structured logs, exact-boundary internal-state capture
90-results-ledger.md Every adopted technique with its tag and measured effect, in one place
── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek v4.1 flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/fast-single-box-deep…] indexed:0 read:10min 2026-10-07 · —