{"slug": "muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding", "title": "Muse Spark 1.3: Cheaper Than GPT-5.6, Better at Coding", "summary": "Meta released Muse Spark 1.3 on September 2, scoring 75.4 on the DeepSWE v1.1 benchmark, surpassing Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0), with a cost per task of $0.55, 42% below GPT-5.6 Sol's $0.95. The model also uses 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2, but its top benchmark results come from a max reasoning configuration that is only in limited partner preview, with no general availability date announced.", "body_md": "Meta dropped Muse Spark 1.3 on September 2 with two numbers worth paying attention to: a 75.4 on [DeepSWE v1.1](https://artificialanalysis.ai/articles/muse-spark-1-3) — the benchmark that measures autonomous software engineering end-to-end — placing it above Claude Opus 5 (74.0) and GPT-5.6 Sol (73.0), and a cost-per-task of $0.55, which is 42% below GPT-5.6 Sol’s $0.95. That combination — leading coding benchmarks at substantially lower cost — is the case for switching your agent infrastructure to 1.3.\n\n## The Efficiency Gain Is the Real Story\n\nBenchmark scores get the headlines, but what actually moves production costs is token and tool-call efficiency. Meta’s engineers measured Muse Spark 1.3 completing coding work with **20% fewer tool calls** and **25% fewer tokens** than 1.2. For high-volume coding agents, that is a direct cost reduction requiring zero code changes — just a model version bump.\n\nTo make it concrete: if you run 10,000 coding tasks per day at $0.55 per task, you are spending $5,500. The same workload under Muse Spark 1.2 (implied ~$0.73 per task) would run closer to $7,300. The 1.3 upgrade pays for itself in the first hour of production traffic.\n\n## What the Benchmarks Actually Measure\n\nThree numbers are worth examining:\n\n**DeepSWE v1.1 — 75.4.** This benchmark runs real software engineering tasks autonomously: open a repo, find the bug, fix it, pass the tests. Muse Spark 1.3 leads here. Claude Opus 5 scores 74.0, GPT-5.6 Sol 73.0. A two-point lead may sound small, but on a benchmark that directly simulates agentic coding workloads, it is not.**Terminal-Bench 2.1 — 88.8.** Ties GPT-5.6 Sol and beats Claude Opus 5 (86.7) on shell and terminal execution tasks.**Long-context retrieval (MRCR 512K-1M) — 98.5.** GPT-5.6 Sol scores 73.8. If your agent needs to hold an entire monorepo in context, Muse Spark 1.3’s one-million-token window is not just bigger — it is more reliable.\n\n## The Catch You Need to Know\n\nMeta’s published benchmark table uses the **max reasoning configuration**. That configuration is in limited partner preview only and has no announced general availability date. What you can actually use today is xhigh mode. [VentureBeat noted](https://venturebeat.com/technology/meta-says-muse-spark-1-3-has-frontier-performance-but-its-best-results-come-from-a-model-developers-cant-broadly-use-yet) that Meta has not disclosed how xhigh performs relative to max on DeepSWE or Terminal-Bench, and Artificial Analysis currently lists no public API provider for max at all.\n\nBenchmarking against a config that most developers cannot access is a marketing choice, not an engineering one. The xhigh model is still competitive — but read the headline numbers accordingly.\n\n## Pricing Comparison\n\n| Model | DeepSWE v1.1 | Output ($/M tokens) | Cost per Task |\n|---|---|---|---|\n| Muse Spark 1.3 | 75.4 | $4.25 | $0.55 |\n| GPT-5.6 Sol | 73.0 | Higher | $0.95 |\n| Claude Opus 5 | 74.0 | ~$25.50 | ~$3.30 |\n\nPer the [DataCamp pricing breakdown](https://www.datacamp.com/blog/muse-spark-1-3), Muse Spark 1.3 input is $1.25 per million tokens, cached input is $0.15 per million tokens, and output is $4.25 per million tokens. The cached input pricing is especially relevant for agents that repeatedly reference the same system prompt or codebase context across calls.\n\n## What Changed in How It Behaves\n\nBeyond the numbers, 1.3 introduces behavioral changes relevant to anyone building agents. The model maintains multiple workflows in a single long thread — useful for agent frameworks where context continuity across tool calls matters. It detects gaps in its own plans and asks for clarification rather than making assumptions, and confirms before executing consequential changes. Code output is cleaner and less verbose than 1.2, which reduces downstream parsing work.\n\n## Which Model Belongs in Your Stack\n\n**Muse Spark 1.3:** High-volume coding agents, tasks where errors are catchable (unit tests, CI pipelines), cost-sensitive deployments, large-codebase retrieval tasks.**Claude Opus 5:** High-stakes single operations where a wrong output is expensive to recover from, or when you need peak quality regardless of cost.**GPT-5.6 Sol:** Browsing-heavy agents, long professional document workflows.\n\n## Meta Is Now a Real Frontier Competitor\n\nThis is the fourth Muse Spark release in five months. Whatever you think of Meta’s AI strategy, the iteration speed is real. The company that spent years releasing open weights models — and faced early criticism for not competing at the frontier — is now shipping proprietary models that lead coding benchmarks and undercut competitors on cost. [Read the official announcement](https://research.meta.ai/blog/introducing-muse-spark-1-3) for full benchmark methodology.\n\nFor developers evaluating AI infrastructure, the question is no longer whether Meta belongs in the conversation. It does. Run a benchmark on your specific task distribution using xhigh mode — the pricing makes the test essentially free. Then decide.", "url": "https://wpnews.pro/news/muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding", "canonical_source": "https://byteiota.com/muse-spark-1-3-coding-agent/", "published_at": "2026-09-04 10:09:21+00:00", "updated_at": "2026-09-04 10:25:03.135938+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-tools"], "entities": ["Meta", "Muse Spark 1.3", "Claude Opus 5", "GPT-5.6 Sol", "DeepSWE v1.1", "Terminal-Bench 2.1", "MRCR 512K-1M", "VentureBeat"], "alternates": {"html": "https://wpnews.pro/news/muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding", "markdown": "https://wpnews.pro/news/muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding.md", "text": "https://wpnews.pro/news/muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding.txt", "jsonld": "https://wpnews.pro/news/muse-spark-1-3-cheaper-than-gpt-5-6-better-at-coding.jsonld"}}