# A Mystery Model Called Ox Alpha Just Topped Coding Benchmarks For Free

> Source: <https://startupfortune.com/a-mystery-model-called-ox-alpha-just-topped-coding-benchmarks-for-free/>
> Published: 2026-08-21 10:37:51+00:00

*A free, anonymous AI model called Ox Alpha showed up on OpenRouter and OpenCode this week, posted eye-catching coding benchmark numbers, and nobody will say who built it.*

OpenRouter listed the model under the provider name "stealth" on August 20, with the model ID stealth/ox-alpha. No developer name. No company. Just a note that a third-party provider chose to stay anonymous during this preview, and that OpenRouter itself is not the model's owner, only the router passing requests to it. That alone is enough to get developers talking. What's pulling them in is the free access and the numbers.

Here's what's actually confirmed. Ox Alpha runs a 1,048,576-token context window with a max output of 131,072 tokens, and it takes text, images, and video as input. It costs nothing right now, zero charge for prompt or completion tokens, according to OpenRouter's own listing. OpenCode, the terminal-based coding agent, picked it up the same day and posted on X that it's free for the next week through its OpenCode Zen plan, with what it called "near unlimited usage" and claimed capacity for 100 trillion tokens a day. Zero data retention, the company added. That's a real offer, not a teaser tier.

The model is being marketed as a reasoning system built for long-horizon software engineering: multi-step agentic coding, terminal work, sustained tasks that run over many turns. Its backers, whoever they are, describe a latent mixture-of-experts architecture that calls four experts for roughly the compute cost of one, a trick meant to buy intelligence without the usual inference bill. Reported strengths include AIME 2025, TerminalBench, and SWE-Bench Verified, the benchmark suite researchers use to test whether a model can actually fix real GitHub issues rather than toy problems.

None of that comes from an audited leaderboard. OpenRouter's own model page carries no official benchmark scores, no intelligence index, nothing from Artificial Analysis or a similar independent tracker. What's circulating instead is developer chatter. One X user running the DeepSWE ten-task coding benchmark posted that Ox Alpha averaged 80%, clearing eight of ten tasks outright with two near misses, ahead of the pack it was tested against. Others on OpenCode's own thread reported it outperforming DeepSeek V4 Flash on multi-step refactors and bug fixes. That's useful signal. It is not a peer-reviewed result, and treating it as one would be a mistake.

[Stripe Agrees to Buy AI Router OpenRouter for More Than $7 Billion](https://startupfortune.com/stripe-agrees-to-buy-ai-router-openrouter-for-more-than-7-billion/)

Stripe has reportedly agreed to buy OpenRouter, the startup whose API routes developer requests across more than 400 AI models, for more than $7 billion, according to Bloomberg. That's over five times OpenRouter's $1.3 billion valuation from a funding round closed just months earlier, and the deal hands one of the world's largest payment... - [stripe acquires openrouter ai router company for billions](https://startupfortune.com/stripe-agrees-to-buy-ai-router-openrouter-for-more-than-7-billion/) - [openrouter seven billion dollar acquisition by stripe announced](https://startupfortune.com/stripe-agrees-to-buy-ai-router-openrouter-for-more-than-7-billion/)

## Who built it?

Nobody knows, and that's the actual story here, not the benchmark chart. Speculation is running toward a major lab quietly stress-testing a model ahead of a public launch, the same pattern seen before with other stealth releases routed through OpenRouter. OpenRouter's post announcing the model didn't hint at an identity, and neither did OpenCode's. Guessing games are already filling the gap: one Chinese-language X thread flagged that OpenCode hasn't disclosed the model's size, parameter count, or training data, calling it "more like an anonymous alpha test" than a launch. That's probably the most honest read available right now.

There's also a small tell that something real is running behind the curtain. At least one developer reported getting rate-limited on their very first prompt, despite the 100-trillion-token daily capacity OpenCode advertised. Free, high-demand previews strain infrastructure fast, and a company confident enough to promise near-unlimited usage still hit a wall within hours of launch. That's worth remembering before you build a workflow around this model.

Here's the practical part. This is free, and it won't stay that way. OpenCode's offer runs for a week from its August 20 announcement, and stealth models on OpenRouter have a history of vanishing or converting to paid tiers once the identity behind them becomes public or the testing window closes. If you want to run your own benchmarks against your own codebase rather than trust a screenshot from someone else's terminal, the clock is already running. Whoever built Ox Alpha is watching the usage data pour in for free, and that's the whole point of the exercise.

**Also read:** [How Vendor-Locked AI Coding Agents Are Quietly Raising Your Engineering Costs](https://startupfortune.com/how-vendor-locked-ai-coding-agents-are-quietly-raising-your-engineering-costs/) • [Micro1 Went From $7 Million to $500 Million in Revenue in a Year](https://startupfortune.com/micro1-went-from-7-million-to-500-million-in-revenue-in-a-year/) • [ChatGPT Can Now Read and Send Your iMessages on a Mac](https://startupfortune.com/chatgpt-can-now-read-and-send-your-imessages-on-a-mac/)
