Muse Glimmer: Meta says its 30B newcomer beats Qwen3.6 on local jobs Meta released Muse Glimmer, a 30-billion-parameter Apache 2.0 licensed model optimized for local deployment, on Monday, claiming it outperforms Qwen3.5 on local tasks. The model, a distillation of Muse Spark with a 128K token context window, is designed to run on consumer hardware with under 20GB memory footprint and 4-bit quantization. Meta said the compression preserves agentic task performance, positioning Glimmer as a drop-in replacement for some offsite inference, and CEO Mark Zuckerberg declared the future of AI must be open. agentic ai /tag/agentic-ai/ Meta dropped a 30-billion-parameter model made for consumer-type machines Monday, the same day its CEO declared that the future of AI must be open. Muse Glimmer is an Apache 2.0 licensed distillation of Muse Spark with a default 128K token context window. "Muse Glimmer was optimized for local deployment, and designed to run at practical speeds on consumer hardware without sacrificing quality," Meta said in documentation https://dev.meta.ai/docs/muse-glimmer?ref=thestack.technology that encourages DIY quantization. Glimmer was shrunk to under 20GB, with quantization to about 4-bit precision, in order to leave headroom for the KV cache, image encoder, and decoding drafter to run simultaneously on the kind of memory available to a recent-model PC or Mac. Meta said it had validated that compression had not degraded agentic task use, which is the entire point of the model it pitched as a drop-in replacement for at least some offsite inference. "Point your existing tooling at a local Muse Glimmer server and keep building — you own the weights, the runtime, and the data," said Meta. Zuckerberg's vision Get the full story: Subscribe for free Join peers managing over $100 billion in annual IT spend and subscribe to unlock full access to The Stack’s analysis and events. Subscribe now https://www.thestack.technology/membership/ Already a member? Sign in https://www.thestack.technology/signin/