# Meta upgrades Muse Spark for long-horizon coding without raising API prices

> Source: <https://www.infoworld.com/article/4218106/meta-upgrades-muse-spark-for-long-horizon-coding-without-raising-api-prices.html>
> Published: 2026-09-03 12:04:21+00:00

Meta is upgrading its Muse Spark family with a new model aimed at long-horizon coding tasks and agentic software development while keeping its standard API pricing unchanged, giving development teams access to improved capabilities without paying more per token.

The new Muse Spark 1.3 model claims to complete coding tasks using about 25% fewer tokens and 20% fewer tool calls than its predecessor, potentially giving development teams more efficient AI-assisted coding at the same per-token price, the company wrote in a [blog post](https://research.meta.ai/blog/introducing-muse-spark-1-3).

The efficiency gains come from Muse Spark 1.3 taking fewer turns when they aren’t needed and producing less verbose responses, while improving its ability to sustain longer-running tasks by generating its own context, correcting gaps in its plan, following complex instructions, managing multiple workflows and recognizing its limitations, it added.

For developers, the combination of improved capabilities and unchanged pricing could boost productivity while encouraging teams to use AI for a broader range of coding tasks, said [Pareekh Jain](https://pareekh.com/about/), principal analyst at Pareekh Consulting.

“Since it is better at understanding complex prompts and following instructions, developers can get better results with fewer retries and potentially use less AI for the same tasks, while also making it easier to extend the model to a broader range of coding and software development workloads with reduced human supervision due to its improved reliability,” Jain said.

These workloads could include understanding large codebases, debugging, changing multiple files, using developer tools, and testing code, Jain added.

That, in turn, could potentially reduce inference costs related to AI-assisted software development for CIOs.

“CIOs potentially get more work for the same token price if Muse Spark 1.3 completes tasks with fewer tokens, tool calls and retries,” Jain further said.

However, that reduction in inference costs may not translate into lower overall AI spending.

Primarily because, most enterprise teams will use the more capable model for longer running and more complex tasks, offsetting the reduction in inference costs, said [Abhishek Satapathy](https://www.linkedin.com/in/abhisekhsatapathy/), principal analyst at Avasant.

That scenario is more likely as enterprises tend to increase their use of a resource when efficiency improvements make it cheaper to consume, a dynamic reflected in the Jevons Paradox, Jain said.

Enterprises should therefore look at the cost of successfully completing a task rather than the model’s per-token price, Jain added.

Meta’s benchmarks of the new model, which were conducted internally, also come with caveats when viewed against independent evaluations.

While Meta claims that Muse Spark 1.3 uses fewer tokens than its predecessor, independent evaluations have produced different results, raising questions about how consistently those efficiency gains will translate across workloads.

Artificial Analysis, an AI benchmarking platform, [found](https://artificialanalysis.ai/models/muse-spark-1-3-xhigh) that Muse Spark 1.3 generated about 100 million tokens while completing its Intelligence Index evaluation, compared with a 71 million-token median across the models it evaluated.

While these results do not directly replicate Meta’s internal coding evaluation, which measures token use on specific coding tasks, they highlight how model efficiency can vary significantly depending on the workload and evaluation method, said [Manoj Chandra Jha](https://www.linkedin.com/in/manoj-chandra-jha-b5ab0a13/), principal analyst at Nord-IQ Research.

“The differing results reinforce the need for enterprises to test the model against their own workloads before budgeting for expected efficiency gains,” Jha added.

Eesel AI, an AI productivity platform**,** also cautioned users against Meta’s claims about Muse Spark 1.3’s performance.

“Meta’s scorecard compares 1.3’s max mode (still gated behind safety testing at launch) against 1.2’s xhigh mode, so part of the jump is a reasoning-tier change, not a clean generational leap,” it wrote in a [blog post](https://www.eesel.ai/blog/muse-spark-1-3).

That makes it harder for enterprises to determine how much of the reported performance gain comes from improvements in 1.3 itself and how much comes from the higher reasoning configuration used in Meta’s comparison, Jha said.

Redditors have also [expressed skepticism](https://www.reddit.com/r/ClaudeAI/comments/1w5mbo4/meta_releases_muse_spark_13_matching_fable_5_w_10/) over Meta’s benchmarks, terming it “classic benchmaxxing”, the practice of optimizing a model or workload to score high on a benchmark rather than a real-world use case.

However, analysts remained optimistic about the new model’s value and its adoption.

“This is Meta playing the value card as the new model undercuts Anthropic and OpenAI’s flagship pricing by 4-to-6x, while still producing results that are competitive with more expensive models,” said Jha.

That should potentially position it as a “strong” alternative for enterprises running complex software development workloads at scale, echoed Satapathy.
