Muse Glimmer 30B Meta released Muse Glimmer 30B, its first open-weights model in the Muse family, under the Apache 2.0 license, featuring a 1.8B vision encoder and a 128K+ context window for coding, agentic workflows, and visual reasoning. Ollama ships the model day-0 with an MLX-optimized tag for Apple Silicon, reporting 1.5x-1.8x faster performance with DFlash, and offers quantized versions (Q4_K_M, Q8_0, FP16) for local deployment on high-end hardware. The model is positioned as a smaller, open alternative to the proprietary Muse Spark family, with independent benchmarks still pending. Muse Glimmer 30B enthusiast Meta’s first open-weights model in the Muse family , released under Apache 2.0. Muse Glimmer 30B is a dense multimodal model with a built-in 1.8B vision encoder and a 128K+ context window . It is designed for coding, agentic workflows, and visual reasoning. Local deployment. Ollama ships the model day-0 with an MLX-optimized tag for Apple Silicon: muse-glimmer:30b-mlx . Ollama reports that the MLX build runs 1.5x-1.8x faster with DFlash than the baseline implementation. Other common tags include muse-glimmer:30b , muse-glimmer:30b-q4 K M , muse-glimmer:30b-q8 0 , and muse-glimmer:30b-fp16 . Agent launch flows. The Ollama registry tags the model for direct launch with Claude Code, Codex, Pi, Hermes, and other agent tools that use the Ollama API. Honest framing. Glimmer is positioned as a smaller, open alternative to the proprietary Muse Spark family. The 30B dense size means the full-precision model requires substantial unified memory, but quantized versions fit on high-end Apple Silicon Macs and modern NVIDIA/AMD GPUs. Independent coding and reasoning benchmarks are still rolling in; treat early claims as preliminary until replicated. - 30.0B - 128k - apache 2.0 - Aug 2026 Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required Can you run it? - reference rigs | Rig | Q4 K M | Q8 0 | FP16 | |---|---|---|---| | NVIDIA Jetson Orin NX 16GB | | no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing no - cloud cloud-pricing Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud Live per-provider pricing, throughput and uptime. Click a column to sort. | Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value | |---|---|---|---|---|---|---|---|---| | Sub | - | - | - | - | - | - | $20.00/mo Pro | | | Sub | - | - | - | - | - | - | $100.00/mo Max | Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing. Detailed API pricing page + JSON endpoint → /models/muse-glimmer-30b/pricing Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.