Meta Muse Spark dominates health AI benchmarks while offering free access that undercuts every paid frontier model Meta's Muse Spark scored 42.8 on HealthBench Hard, more than double Gemini 3.1 Pro's 20.6, while remaining free on meta.ai and pricing its API at $1.25 per million input tokens and $4.25 per million output tokens, undercutting OpenAI and Anthropic. Developed by Meta Superintelligence Labs under Chief AI Officer Alexandr Wang, the model also ranks fourth on the Artificial Analysis Intelligence Index v4.0 with a score of 52, behind Gemini 3.1 Pro and GPT-5.4 at 57 and Claude Opus 4.6 at 53. Meta's Muse Spark has scored 42.8 on HealthBench Hard, more than double Gemini 3.1 Pro's 20.6, while remaining free on meta.ai and pricing its API at roughly a quarter of what OpenAI and Anthropic charge. The model that was supposed to be Meta's catch-up shot has turned into something more disruptive. Muse Spark, developed by Meta Superintelligence Labs under Chief AI Officer Alexandr Wang, the former Scale AI CEO who joined Meta through a $14.3 billion deal giving Meta a 49% stake in the data labeling company, doesn't beat the field on everything. But where it does win, it wins decisively. On HealthBench Hard, the most demanding of OpenAI's physician-curated medical benchmarks, Muse Spark scored 42.8. GPT-5.4 scored 40.1. Gemini 3.1 Pro and Grok 4.2 both landed in the low 20s. The gap isn't close. That health benchmark gap matters more than the headline number suggests. According to data published by BenchLM.ai, every other frontier model tested on HealthBench Hard trailed Muse Spark by at least two points, and Gemini, which many health systems have been quietly piloting over the last eighteen months, scored less than half of Muse Spark's figure. Meta, which worked with over 1,000 physicians to curate the model's training data according to TechCrunch's April 2026 reporting on the model's debut, turned a traditionally weak area for social media companies into a genuine advantage. It's worth being precise about where Muse Spark actually sits. On the Artificial Analysis Intelligence Index v4.0, it ranks fourth overall with a score of 52, behind Gemini 3.1 Pro and GPT-5.4 at 57 and Claude Opus 4.6 at 53. So don't mistake this for a total upset. What Muse Spark has done is pick its lanes: health AI, token efficiency, and, with the July 9 release of Muse Spark 1.1, coding and agentic work. On those sub-evals, it beats Gemini. On abstract reasoning and some agentic tasks, Gemini still has the edge. The pricing is where this gets complicated for the rest of the industry. Muse Spark is free for consumers on meta.ai and through the Meta AI app. The new Muse Spark 1.1 API, which Meta opened on July 9, costs $1.25 per million input tokens and $4.25 per million output tokens. For reference, Claude Opus 4.6 runs at $15 per million input and $75 per million output. New API accounts get $20 in free credits. Frankly, if those numbers hold, a meaningful portion of the AI startup market will reprice its infrastructure assumptions within the next two quarters. The token efficiency angle reinforces this. Muse Spark runs on approximately 58 million output tokens across the benchmark suite, compared to 157 million for Claude Opus 4.6, according to figures published by Artificial Analysis. A frontier model producing comparable or superior results on specific verticals while consuming roughly a third of the tokens isn't a minor operational footnote. It's the number a CFO notices. Wang's team has been explicit that Muse Spark 1.1 was trained specifically for coding and agentic performance, and CNBC reported that Meta views coding improvement as the engine for general agent capability. The model supports a 1 million token context window with active compaction, meaning it manages what it retains and what it compresses during long sessions rather than forcing the developer to handle that ceiling manually. That's a concrete advantage. The health AI wedge is the longer story Med-tech and life sciences companies can't ignore a gap that size. Health AI is a sector where switching costs are punishing: regulatory approvals, clinical validation studies, data integration contracts all have to be renegotiated, revalidated, and re-signed before you can move a model. Once a hospital system or pharma company commits to a vendor, they don't move quickly - inertia is the default, and the default is expensive to break. If Muse Spark's health benchmark lead persists through the next round of evals, Meta will have a credible conversation with enterprise buyers that it simply couldn't have had a year ago. What hasn't changed is the tension inside Meta's own identity. As Artificial Intelligence News noted in April, Muse Spark is a real departure from Meta's open-source tradition. Llama, Meta's open-weights family, built the company enormous goodwill with developers. Muse Spark is closed. Wang's Lab is building a competitive commercial model, not a research artifact to release on Hugging Face. Whether those two things can coexist inside one company over the next few years is a real question, and Meta hasn't answered it. But the near-term picture is clear enough. A free frontier model that nearly doubles Gemini's score on the hardest health benchmark, available at a quarter of the API price of its closest peers, is not a niche product targeting one vertical. It's a cost argument aimed at everyone building on top of AI right now. Also read: OpenAI raises its compute bet to $750 billion but its own CFO isn't sure it can pay the bill https://startupfortune.com/openai-raises-its-compute-bet-to-750-billion-but-its-own-cfo-isnt-sure-it-can-pay-the-bill/ • Samsung launches three foldables at once in London and bets Gemini AI can make them mainstream https://startupfortune.com/samsung-launches-three-foldables-at-once-in-london-and-bets-gemini-ai-can-make-them-mainstream/ • Monday.com cuts 620 jobs and raises its margin outlook in the same breath https://startupfortune.com/mondaycom-cuts-620-jobs-and-raises-its-margin-outlook-in-the-same-breath/