12:12
2026-09-02
github.com
artificial-intelligence
Show HN: PulsarForge – run a 744B MoE model on 32GB RAM with zero GPU (pure C)
PulsarForge, a CPU-only LLM inference engine written in C11, runs a 744-billion-parameter GLM-5.2 MoE model on a 2018 laptop with 32GB RAM and a USB SSD, achieving 9.3 seconds per token interactive sp…