18:49
2026-07-05
gist.github.com
large-language-models
Running DeepSeek-V4-Flash (284B MoE) on a 64GB Strix Halo via SSD expert-streaming
A developer successfully ran DeepSeek-V4-Flash, a 284B-parameter mixture-of-experts model, on a 64GB AMD Strix Halo laptop by streaming cold experts from SSD via mmap. The model achieves ~1.9 tok/s deβ¦