10:59
2026-10-07
tangled.org
large-language-models
Fast single-box DeepSeek v4.1 Flash runtime
A new recipe for running DeepSeek V4.1 Flash on a single AMD Strix Halo machine (gfx1151 GPU, 128 GB unified memory, several NVMe drives) reports roughly 450 tokens per second prefill and 15 tokens pe…