Colibri Runs a 744B Model on 25GB RAM — Here’s the Architecture
On July 1, 2026, developer JustVugg released Colibri, a 14,700-line pure-C inference engine that runs GLM-5.2, a 744-billion-parameter Mixture-of-Experts model, on 25GB of RAM with no GPU, using an LR…