They claim to have optimized it for exactly the kind of things I'm looking for in a local model:
End-to-end Agentic Task Completion.Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish.Reliable Tool Use.The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows.Multi-Step Reasoning.Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows. [...]
Here's [a pelican](https://gist.github.com/simonw/f20d4cd0ea7596990f7910ead616493e) which I generated using LM Studio's [18.16 GB version of the model](https://lmstudio.ai/models/muse-glimmer):
I also tried it out with my [llm-coding-agent](https://github.com/simonw/llm-coding-agent) plugin, running against a fresh checkout of Datasette with the prompt:
how does auth work?
Here's the response, at the end of a long transcript showing all of the tool calls it made to explore the codebase.
I really like this size of model, because if a machine has 32 GB of RAM or more (mine has 128GB) it leaves plenty of space for running other applications at the same time.
Via [Hacker News](https://news.ycombinator.com/item?id=49241679)
Tags: [ai](https://simonwillison.net/tags/ai), [generative-ai](https://simonwillison.net/tags/generative-ai), [llama](https://simonwillison.net/tags/llama), [local-llms](https://simonwillison.net/tags/local-llms), [llms](https://simonwillison.net/tags/llms), [meta](https://simonwillison.net/tags/meta), [llm-release](https://simonwillison.net/tags/llm-release)