22:37
2026-09-01
github.com
machine-learning
Moe expert offloading on a 2-core Celeron with 2.7GB RAM
A developer's on-hardware test shows that prefetching mixture-of-experts (MoE) expert weights from disk ahead of matmuls can improve token generation speed on a severely resource-constrained machine, …