Meta dropped a 30-billion-parameter model made for consumer-type machines Monday, the same day its CEO declared that the future of AI must be open.
Muse Glimmer is an Apache 2.0 licensed distillation of Muse Spark with a default 128K token context window.
"Muse Glimmer was optimized for local deployment, and designed to run at practical speeds on consumer hardware without sacrificing quality," Meta said in documentation that encourages DIY quantization.
Glimmer was shrunk to under 20GB, with quantization to about 4-bit precision, in order to leave headroom for the KV cache, image encoder, and decoding drafter to run simultaneously on the kind of memory available to a recent-model PC or Mac.
Meta said it had validated that compression had not degraded agentic task use, which is the entire point of the model it pitched as a drop-in replacement for at least some offsite inference.
"Point your existing tooling at a local Muse Glimmer server and keep building — you own the weights, the runtime, and the data," said Meta.
Zuckerberg's vision
Get the full story: Subscribe for free #
Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events.
[Subscribe now](https://www.thestack.technology/membership/)
Already a member? [Sign in](https://www.thestack.technology/signin/)