Meta upgrades Muse Spark for long-horizon coding without raising API prices Meta has released Muse Spark 1.3, an upgraded AI model for long-horizon coding and agentic software development, claiming it uses about 25% fewer tokens and 20% fewer tool calls than its predecessor while keeping API prices unchanged. The efficiency gains could reduce inference costs for CIOs, but analysts warn that enterprises may offset savings by using the model for more complex tasks, and independent evaluations by Artificial Analysis show higher token usage than Meta's internal benchmarks, underscoring the need for workload-specific testing. Meta is upgrading its Muse Spark family with a new model aimed at long-horizon coding tasks and agentic software development while keeping its standard API pricing unchanged, giving development teams access to improved capabilities without paying more per token. The new Muse Spark 1.3 model claims to complete coding tasks using about 25% fewer tokens and 20% fewer tool calls than its predecessor, potentially giving development teams more efficient AI-assisted coding at the same per-token price, the company wrote in a blog post https://research.meta.ai/blog/introducing-muse-spark-1-3 . The efficiency gains come from Muse Spark 1.3 taking fewer turns when they aren’t needed and producing less verbose responses, while improving its ability to sustain longer-running tasks by generating its own context, correcting gaps in its plan, following complex instructions, managing multiple workflows and recognizing its limitations, it added. For developers, the combination of improved capabilities and unchanged pricing could boost productivity while encouraging teams to use AI for a broader range of coding tasks, said Pareekh Jain https://pareekh.com/about/ , principal analyst at Pareekh Consulting. “Since it is better at understanding complex prompts and following instructions, developers can get better results with fewer retries and potentially use less AI for the same tasks, while also making it easier to extend the model to a broader range of coding and software development workloads with reduced human supervision due to its improved reliability,” Jain said. These workloads could include understanding large codebases, debugging, changing multiple files, using developer tools, and testing code, Jain added. That, in turn, could potentially reduce inference costs related to AI-assisted software development for CIOs. “CIOs potentially get more work for the same token price if Muse Spark 1.3 completes tasks with fewer tokens, tool calls and retries,” Jain further said. However, that reduction in inference costs may not translate into lower overall AI spending. Primarily because, most enterprise teams will use the more capable model for longer running and more complex tasks, offsetting the reduction in inference costs, said Abhishek Satapathy https://www.linkedin.com/in/abhisekhsatapathy/ , principal analyst at Avasant. That scenario is more likely as enterprises tend to increase their use of a resource when efficiency improvements make it cheaper to consume, a dynamic reflected in the Jevons Paradox, Jain said. Enterprises should therefore look at the cost of successfully completing a task rather than the model’s per-token price, Jain added. Meta’s benchmarks of the new model, which were conducted internally, also come with caveats when viewed against independent evaluations. While Meta claims that Muse Spark 1.3 uses fewer tokens than its predecessor, independent evaluations have produced different results, raising questions about how consistently those efficiency gains will translate across workloads. Artificial Analysis, an AI benchmarking platform, found https://artificialanalysis.ai/models/muse-spark-1-3-xhigh that Muse Spark 1.3 generated about 100 million tokens while completing its Intelligence Index evaluation, compared with a 71 million-token median across the models it evaluated. While these results do not directly replicate Meta’s internal coding evaluation, which measures token use on specific coding tasks, they highlight how model efficiency can vary significantly depending on the workload and evaluation method, said Manoj Chandra Jha https://www.linkedin.com/in/manoj-chandra-jha-b5ab0a13/ , principal analyst at Nord-IQ Research. “The differing results reinforce the need for enterprises to test the model against their own workloads before budgeting for expected efficiency gains,” Jha added. Eesel AI, an AI productivity platform , also cautioned users against Meta’s claims about Muse Spark 1.3’s performance. “Meta’s scorecard compares 1.3’s max mode still gated behind safety testing at launch against 1.2’s xhigh mode, so part of the jump is a reasoning-tier change, not a clean generational leap,” it wrote in a blog post https://www.eesel.ai/blog/muse-spark-1-3 . That makes it harder for enterprises to determine how much of the reported performance gain comes from improvements in 1.3 itself and how much comes from the higher reasoning configuration used in Meta’s comparison, Jha said. Redditors have also expressed skepticism https://www.reddit.com/r/ClaudeAI/comments/1w5mbo4/meta releases muse spark 13 matching fable 5 w 10/ over Meta’s benchmarks, terming it “classic benchmaxxing”, the practice of optimizing a model or workload to score high on a benchmark rather than a real-world use case. However, analysts remained optimistic about the new model’s value and its adoption. “This is Meta playing the value card as the new model undercuts Anthropic and OpenAI’s flagship pricing by 4-to-6x, while still producing results that are competitive with more expensive models,” said Jha. That should potentially position it as a “strong” alternative for enterprises running complex software development workloads at scale, echoed Satapathy.