AI releases are coming in thick and fast as we approach what’s been called the “foothills of the singularity”.
Meta has released Muse Spark 1.3, the latest update to the model family that Meta Superintelligence Labs (MSL) has been shipping at a relentless pace since April. Mark Zuckerberg announced the release himself, calling it frontier performance that’s “almost too cheap to meter,” and describing it as the biggest single jump the team has made on coding and agentic work so far. The model is live now in Muse Code, Meta’s coding agent, and through the Meta Model API, with open weights for the Muse Spark line said to be coming soon.
On a company-published benchmark table spanning agentic, long-context, and coding categories, Muse Spark 1.3 (max) beats Anthropic’s Opus 5 (max) and OpenAI’s GPT 5.6 Sol (max) on several evaluations, while trailing both on others — a mixed but genuinely competitive picture for a model family that didn’t exist five months ago.
This is Muse Spark’s fourth major release since its April debut, following Muse Spark 1.1 in July and Muse Spark 1.2 in August. That’s four frontier-model releases in roughly five months, a cadence that would be aggressive by the standards of any AI lab right now.
Muse Spark 1.3 Benchmarks #
Meta’s comparison table pits Muse Spark 1.3 (max) against its own predecessor Muse Spark 1.2 (xhigh), OpenAI’s GPT 5.6 Sol (max), and Anthropic’s Opus 5 (max), across three broad buckets: agentic tasks, long-context handling, and coding.
Agentic performance is where the results are most mixed. On GDPVal-AA v2, a knowledge-work benchmark, Muse Spark 1.3 scores 1754, ahead of GPT 5.6 Sol’s 1710 but behind Opus 5’s 1824. On JobBench (professional tool use) and OSWorld 2.0 (agentic computer use), Muse Spark 1.3 posts 64.9 and 66.9 respectively, comfortably ahead of GPT 5.6 Sol on both and just short of Opus 5. On AutomationBench, a test of end-to-end business workflows, it scores 49.4, again close behind Opus 5’s 50.3. But on DeepSearchQA (agentic browsing) and the internal Agentic IF Index (instruction following), GPT 5.6 Sol pulls clearly ahead, scoring 93.0 and 60.5 against Muse Spark 1.3’s 89.4 and 57.8.
Long context is Muse Spark 1.3’s strongest category by a wide margin. On MRCR 256K-512K and MRCR 512K-1M, it scores 98.5 and 98.1, well clear of GPT 5.6 Sol’s 91.5 and 73.8 and of its own predecessor Muse Spark 1.2, which scored 66.3 and 55.5. Opus 5 has no listed score on either test in Meta’s table. This tracks with Meta’s broader push on context handling — Muse Spark 1.1 was already built to manage a million-token context window on its own, deciding what to retain, retrieve, or compress as a session runs long, and 1.3 appears to have sharpened that further.
Coding is the category Zuckerberg singled out, and the numbers back that up relative to Muse Spark 1.2. On DeepSWE v1.1 (long-horizon agentic coding), Muse Spark 1.3 jumps to 75.4 from 1.2’s 55.0, edging past Opus 5’s 74.0. On SWEAtlas CodeBase QnA, it scores 59.4, ahead of both GPT 5.6 Sol (53.5) and Opus 5 (52.7). On Terminal-Bench 2.1, it ties GPT 5.6 Sol at 88.8, ahead of Opus 5’s 86.7.
Taken together, the table shows a model that has closed most of the gap to GPT 5.6 Sol and Opus 5 on agentic and coding work, leads outright on long-context tasks, and still cedes ground to GPT 5.6 Sol specifically on browsing and instruction-following. As with any lab-published benchmark set, it’s worth keeping in mind that these are Meta’s own numbers, run on Meta’s own harness — independent evaluators like Artificial Analysis, which had early access to Muse Spark 1.1 and Muse Spark 1.2, will likely publish their own third-party numbers on 1.3 in the coming days.
A Fast-Moving Model Family #
Muse Spark’s rise has been notably quick. The original Muse Spark launched in April as MSL’s first model, built after Meta poached Scale AI’s Alexandr Wang to lead the effort. It scored 43 on the Artificial Analysis Intelligence Index at launch. Three months later, Muse Spark 1.1 pushed that to 51, and Zuckerberg marked the occasion by posting on X for the first time in three years. In August, Muse Spark 1.2 announced plans for open weights alongside the smaller, locally-runnable Muse Glimmer model, positioning Muse Spark 1.2 as potentially the most capable open-weights model out of an American lab, if and when the weights land.
Pricing details for 1.3 haven’t been published yet, but Zuckerberg’s “too cheap to meter” framing continues a theme from earlier releases — Muse Spark 1.2 was already priced at $1.25 per million input tokens and $4.25 per million output tokens, undercutting much of the frontier field.
Muse Spark 1.3 rolls out as developers can access it immediately through Muse Code and the Meta Model API, with a broader rollout to Meta AI, Instagram, and Facebook expected to follow. A “max reasoning” mode is reportedly still going through internal safety testing before it ships.