18:28
2026-09-18
arxiv.org
large-language-models
SSD-Llama: SSD-Native Inference for Trillion-Parameter Moe on a Consumer PC
Researchers submitted SSD-LLaMA to arXiv on 16 Sep 2026, an SSD-native local Mixture-of-Experts inference system that runs trillion-parameter models at over 1 token/s on a single RTX 5090 with no moreβ¦