05:57
2026-09-16
github.com
ai-infrastructure
Prefill a 284B model on Nvidia. Decode it on Apple Silicon. Over plain 10GbE
A prefill/decode disaggregation setup running DeepSeek-V4-Flash β a 284B total / 13B active model with 256 routed experts β bridged a 700,630-token cold prompt end to end in 11 minutes 26 seconds overβ¦