{"slug": "muse-spark-1-3-a-review", "title": "Muse Spark 1.3 - A Review", "summary": "A developer's review of Meta's Muse Spark 1.3 LLM and its harness highlights strengths in code generation and value but criticizes skill-following inconsistencies and a restrictive sandbox. The model is seen as competitive with frontier models like Claude and Grok, though the ecosystem remains rough around the edges.", "body_md": "In this post I'll talk about my brief experience with `muse`\n\n, Meta's LLM harness for developers, as well as Muse Spark 1.3, their latest frontier-level model.\n\nI'll start with the bad, just because I like to end with the positive :)\n\nIt's not that good following skills. If the skill has `disable-model-invocation`\n\n, sometimes it refuses to launch it, even if you manually call it. I think it happens when you call the skill mid-sentence, but it's not consistent.\n\nIt's also not as good as other models at following skill instructions. It seems to get confused more often.\n\nFor example, I have one skill that will address an issue from GitHub to PR.\n\nIn Claude (Opus 5) and Cursor (Grok 4.6) it works perfectly. The first step is [grilling](https://www.aihero.dev/skills-grill-me) the issue, after that's finished, the next step is autonomous, plan, implement with TDD, review and open PR.\n\nWith Muse Spark 1.3, sometimes the skill will not continue and I have to nudge it for the next step, just saying something like \"continue\" is enough, but surely is annoying.\n\nThe output is not great. Sometimes it will show me raw markdown, sometimes not. It's not consistent.\n\nHaving a sandbox is good, but in this case, it's a bit too restrictive. For example, I'm working with a Firebase project and I want to use the emulators.\n\nWell, too bad. The sandbox doesn't allow you to run files outside your workspace or use external ports.\n\nThat would be great if I could add exceptions or some kind of configuration, but you can't. You are basically forced into `--yolo`\n\nmode if you don't want to be prompted on repeat for the same things over and over.\n\nWhat's sad is that even if you want to give them access, the models will just get stuck asking for permissions for the same thing over and over again and eventually they will just be stuck doing nothing.\n\nNot everything is bad, of course. With a bit of effort I think it's actually quite usable.\n\nThe main reason I decided to try the model. The subscription plan is relatively generous for it's price. With the $15 plan you can code non-stop for a couple hours, in my experience around 2 to 4 hours, before you reach the 5-hour limit.\n\nBut of course, it depends on your usage. If you have a lot of concurrent sessions, it will last less.\n\nAlso the weekly usage is a bit low, I'd say around 4-5 times the 5 hour limit (~25h).\n\nI think compared to other plans though it's probably one of the best values. For $50 you can get ~33% more usage, which should be enough for ~11.5hs of coding a day. Quite decent for that price tag.\n\nMy biggest hurdles with the model have been in the actual development flow. When it comes to general intelligence and code, I think it's quite good. It feels around Opus 4.8 / Grok 4.6. Of course, that's totally subjective.\n\nI think the code it generates is very similar to most other frontier models. I don't think anyone can really tell what model generated what output nowadays anyway, they are all quite competent.\n\nUnlike Claude which speaks a weird dialect of english, called *claudish* by most devs, the output from this model seems to be much more readable in general.\n\nIn my experience I rarely need to tell it to re-word what it just said.\n\nI think the model speed is quite good. I think the time to first token is average, but after that it feels pretty fast.\n\nOverall, I think this whole ecosystem is still a bit rough around the edges, but if you are willing to work around that, it's an usable model and harness.\n\nProbably it will be much better in a couple months. It's clear they are playing catch-up at this point. But they (Meta) seems to have the tools (model, data centers, engineers), so they have a real shot at creating a good product.\n\nI don't think there's a real reason to swap from, say, Codex, Claude or Cursor unless you want to try something new, or maybe are looking for a cheaper alternative.\n\nLet me know what you think in the comments, I'd love to hear other people's opinions about this model/harness!", "url": "https://wpnews.pro/news/muse-spark-1-3-a-review", "canonical_source": "https://dev.to/gosukiwi/muse-spark-13-a-review-3h7", "published_at": "2026-09-04 03:56:36+00:00", "updated_at": "2026-09-04 04:26:53.470918+00:00", "lang": "en", "topics": ["large-language-models", "developer-tools", "ai-products"], "entities": ["Meta", "Muse Spark 1.3", "Claude", "Grok", "Cursor", "Codex"], "alternates": {"html": "https://wpnews.pro/news/muse-spark-1-3-a-review", "markdown": "https://wpnews.pro/news/muse-spark-1-3-a-review.md", "text": "https://wpnews.pro/news/muse-spark-1-3-a-review.txt", "jsonld": "https://wpnews.pro/news/muse-spark-1-3-a-review.jsonld"}}