Hugging Face and Meta’s launch artwork for Muse Glimmer. Image: Hugging Face and Meta
Muse Glimmer is Meta’s free, open-weight AI model built to run AI agents on your own computer: a 30-billion-parameter model that reads text and images, writes code, calls tools and fits on a single 24GB graphics card or a well-equipped Mac. Meta Superintelligence Labs released it on August 10, 2026 under the Apache 2.0 licence, so anyone can download it, use it commercially and change it. Here is what it is, what you need to run it, the exact commands for Ollama, LM Studio and llama.cpp, and how it stacks up against Qwen 3.8 and Gemma 4.
Jump to:
[What is Muse Glimmer?](#what)
[Glimmer vs Muse Spark](#spark)
[Hardware you need](#hardware)
[Run it with Ollama](#ollama)
[LM Studio](#lm-studio)
[llama.cpp and DFlash](#llama-cpp)
[Best settings](#settings)
[vs Qwen 3.8 and Gemma 4](#compare)
[Where it works](#where)
[FAQ](#faq)
What is Muse Glimmer? #
Muse Glimmer is a dense 30B model (about 29.6 billion parameters, including a 1.8 billion-parameter vision encoder) that Meta describes as “an open model built for always-on local agents”. Its model card lists the basics:
- Input and output: text and images in, text out. Video is handled as individual frames; there is no audio.
- Context window: 131,072 tokens (Ollama rounds it to 128K).
- Languages: trained on data from more than 100 languages, though Meta says it hasn’t evaluated all of them.
- Knowledge cutoff: January 4, 2026.
- Built for: local AI agents, coding agents, tool use and function calling, multimodal reasoning and grading other models’ answers.
Meta says it is tuned for reliable tool calls, long tasks and recovering from failures, with “self-managed memory over hours-long sessions”. It is not meant for under-18s, and Meta recommends adding guardrails such as human confirmation before any irreversible action.
Muse Glimmer vs Muse Spark #
Glimmer is the small, open sibling of Muse Spark, the larger model behind Meta’s Muse AI agent. Meta pre-trained Glimmer “on Muse Spark’s outputs using logit distillation”, a way of teaching a small model to imitate a big one. The difference is where they run: Muse Spark is only available through Meta’s apps and its paid Model API, while Glimmer’s weights are free to download and run offline. Developers entering Meta’s $1 million AI hackathon can use either.
What hardware do you need to run Muse Glimmer? #
The full-precision model needs about 60GB of memory, but Meta’s own 4-bit versions shrink it to fit a 24GB or 32GB card with little loss. Meta’s quantisation guide says to “start here” with the 17GB version:
| Version | Download | Made for | Accuracy loss* |
|---|---|---|---|
| K-Quant-17GB (Q4_K_M) | 16.8GB | 24GB graphics card | 1.0% |
| K-Quant-Dynamic (Q4_K_XL) | 19.7GB | 32GB graphics card | 0.2% |
| Full precision (BF16) | about 60GB | 64GB of video memory | none |
*Meta’s average across 15 benchmarks. Image input needs a 1.4GB vision file on top; the optional DFlash speed-up file is 1.6GB.
Meta says the 17GB version with image support and the full 131,072-token context uses 19.0GiB, which “fits a 24 GiB card with headroom”, so an RTX 4090 or 5090 class card works. On a Mac, LM Studio says you need “at least 26 GB of RAM” for the smallest version, which in practice means a Mac with 32GB of unified memory or more. If you are short of memory, Meta suggests a smaller context or skipping the vision file. Our guide to running local AI on a Mac mini explains how unified memory affects which models fit.
How to run Muse Glimmer with Ollama #
Ollama is the quickest route on Windows, Mac and Linux. Install it from ollama.com, then open a terminal and run:
- Windows, Linux or an Intel Mac:
ollama run muse-glimmer(an 18GB download). - Apple Silicon Macs:
ollama run muse-glimmer:30b-mlx(19GB). Ollama says its MLX engine gives the best performance on Apple chips, with DFlash and image input supported.
Ollama’s tag list has 15 versions, ranging from 17GB to 65GB, plus -dflash builds that bundle the speed-up drafter (the 4-bit one is 20GB). Every version has a 128K context window and accepts images.
Ollama can also plug Glimmer straight into coding agents. Its library page lists ollama launch claude --model muse-glimmer for Claude Code, and the same pattern for OpenCode, Hermes Agent and OpenClaw, so the agent runs on your machine with no API bill.
How to run Muse Glimmer in LM Studio #
LM Studio is the easiest option if you’d rather not use a terminal. Install it from lmstudio.ai, search for “Muse Glimmer” in the model browser and download it. Its model page lists a 25.77GB GGUF download, says Glimmer models “support tool use, vision input, and reasoning”, and repeats the 26GB memory minimum. Once it’s loaded you can chat with it, drop in images, or switch on LM Studio’s local server so other apps can use it like the OpenAI API.
llama.cpp, vLLM and the DFlash speed boost #
For more control, Hugging Face’s [launch post](https://huggingface.co/blog/muse-glimmer) gives these llama.cpp commands:
- **Start a server with a web chat at localhost:8080:**`llama serve -hf meta-models/Muse-Glimmer-30B-GGUF`
- **Add the DFlash drafter:**`llama serve -hf meta-models/Muse-Glimmer-30B-GGUF --spec-type draft-dflash --spec-draft-n-max 15`
DFlash is speculative decoding: a small drafter guesses 16 tokens at a time and Glimmer checks them in one pass, with what Meta calls “identical output quality”. Meta measured these speeds:
| Hardware | Normal | With DFlash | Speed-up |
|---|---|---|---|
| Nvidia RTX 5090 | 74.9 tokens/sec | 233.4 tokens/sec | 3.1x |
| Apple M5 Max | 26.6 tokens/sec | 50.2 tokens/sec | 1.8x |
| Apple M4 Max | 23.7 tokens/sec | 37.8 tokens/sec | 1.5x |
Meta’s figures, one request at a time; the RTX test used llama.cpp and the Macs used ExecuTorch.
Servers with a 64GB-plus card can run the full model with vllm serve meta-models/Muse-Glimmer-30B, which gives an OpenAI-compatible API on port 8000. Meta says there is “no API key and no gated access” on any of its downloads.
Best Muse Glimmer settings #
Meta’s prompting guide recommends a temperature of 1.0, top_p of 0.95 and top_k of 64. Glimmer thinks before it answers, and you set how hard with a reasoning strength of low, medium, high or xhigh. The default is high, and Meta suggests high or xhigh for coding, agents and hard problems, and lower settings when you want speed. Two more tips from Meta: give it plenty of room to reply, since a low token limit can cut off its reasoning before the answer, and note that it makes one tool call per turn, with no parallel tool calls.
Muse Glimmer vs Qwen 3.8 vs Gemma 4 #
All three are free, Apache 2.0 models of about the same size, but they are not equal. Alibaba’s Qwen3.8-27B came out a few days after Glimmer, in mid-August 2026, and its model card compares itself directly with Glimmer and beats it on every test it lists. Meta’s own table compares Glimmer with Google’s Gemma 4 31B and the older Qwen3.6-27B, and there Glimmer wins on most agent and coding tests.
| | Muse Glimmer 30B | Qwen3.8-27B | Gemma 4 31B |
|---|---|---|---|
| SWE-bench Pro (coding) | 51.2 | 61.7 | 36.9 |
| Terminal-Bench 2.1 | 51.7 | 73.0 | 43.4 | | OSWorld-Verified (computer use) | 65.9 | 84.3 | 58.5 | | MCP Atlas (tool use) | 75.5 | not published | 54.2 | | GPQA Diamond (science) | 83.5 | 89.2 | 85.7 | | Context window | 131K | 262K (up to 1M) | 256K | | Input | text, images | text, images, video | text, images |
Scroll sideways to see all columns.
Glimmer and Gemma 4 scores from Meta’s model card; Qwen3.8 scores from Alibaba’s. Each company ran its own tests, so treat close results with caution.
In short: Qwen3.8-27B is the stronger all-rounder on paper and has a much longer memory, Gemma 4 31B is Google’s option with the widest language support (more than 140 languages in training), and Glimmer’s strengths are tool use, multi-step agent work and an official 4-bit build with a ready-made speed-up. Meta’s own safety table also shows Glimmer resisting prompt-injection attacks better than Qwen3.6, though not quite as well as Gemma 4. The best test is your own work: all three run in Ollama and LM Studio, so try each on a real task.
Where Muse Glimmer works #
Because you download the weights, Glimmer works anywhere you can reach Hugging Face, Ollama or LM Studio, including the US, UK, EU, Canada, Australia, India and Germany, and once downloaded it runs fully offline. Meta lists no country restrictions, and unlike Meta’s Muse app, which is not yet in the UK or EU, there is no waitlist. If your computer can’t run it, Together AI, Fireworks AI and OpenRouter host it for a per-token fee. Meta’s usage policy still applies however you use it.
Last checked: October 10, 2026, against Meta’s model card and docs, Ollama’s library and LM Studio’s model page.
Frequently asked questions #
Is Muse Glimmer free?
Yes. Meta released the weights free under the Apache 2.0 licence, which allows commercial use. Running it on your own computer costs nothing; hosted versions on Together AI, Fireworks AI and OpenRouter charge per token.
How do I run Muse Glimmer in Ollama?
Install Ollama and run ollama run muse-glimmer (an 18GB download). On an Apple Silicon Mac use ollama run muse-glimmer:30b-mlx instead, which uses Ollama’s faster MLX engine.
How much VRAM does Muse Glimmer need?
Meta’s 17GB 4-bit version fits a 24GB graphics card, using 19.0GiB with image support and the full context. The higher-quality Dynamic version targets 32GB, and the full model needs about 64GB.
Can Muse Glimmer run on a Mac?
Yes, on Apple Silicon with enough memory. LM Studio says the smallest version needs at least 26GB of RAM, so a Mac with 32GB or more of unified memory. Meta measured 50.2 tokens a second on an M5 Max with DFlash.
Is Muse Glimmer better than Qwen 3.8?
Not on the published numbers. Alibaba’s Qwen3.8-27B card scores higher than Glimmer on SWE-bench Pro, Terminal-Bench 2.1, OSWorld-Verified and GPQA Diamond. Glimmer’s strengths are tool use and agent work, where Meta’s own tests show it ahead of Gemma 4 and Qwen3.6.
Is Muse Glimmer the same as Meta Muse?
No. Muse is Meta’s AI agent app, powered by the larger Muse Spark model. Glimmer is a smaller model distilled from Muse Spark that you download and run yourself.
Can Muse Glimmer power Claude Code?
Yes, through Ollama. Run ollama launch claude --model muse-glimmer to point Claude Code at Glimmer on your own machine. Ollama lists the same command for OpenCode, Hermes Agent and OpenClaw.
Does Muse Glimmer work in the UK and EU?
Yes. The weights are a free download with no gated access and Meta lists no country restrictions, so it works in the UK, EU and everywhere else you can reach Hugging Face, Ollama or LM Studio.
What is Muse Glimmer’s context window?
131,072 tokens, shown as 128K in Ollama. Qwen3.8-27B (262K) and Gemma 4 31B (256K) have longer windows.
Sources: Meta: Introducing Muse Glimmer; Muse Glimmer model card; Meta developer docs; Hugging Face; Ollama; LM Studio; Qwen3.8-27B model card; Gemma 4 31B model card.