# Ornith-1.5 Is a Free Open-Source Model That Claims to Rival Claude Opus 4.8

> Source: <https://startupfortune.com/ornith-15-is-a-free-open-source-model-that-claims-to-rival-claude-opus-48/>
> Published: 2026-08-20 16:43:38+00:00

*Ornith-1.5 is current because Ornith announced it on X on August 20, 2026, but the safest version of this story is narrower than the published draft. The release is real, the benchmark claims are vendor-reported, and the outside-attributed comparison to Claude Opus 4.8 needed to go.*

Ornith has turned a small open-model release into a much bigger claim: three new coding models, a self-improvement loop, and benchmark scores that the lab says put its largest system near Claude Opus 4.8. You should read that carefully. The models are open, the license is permissive, and the numbers are striking, but the numbers still come from Ornith's own announcement.

According to Ornith's X post on August 20, the Ornith-1.5 family includes a 9B dense model, a 35B mixture-of-experts model, and a 397B MoE flagship. The lab says all three are trained with self-improving strategies and released under the MIT license, including quantized versions in FP8, GGUF, MLX and NVFP4 formats. That part is the practical hook. If the model files are usable at the sizes Ornith is promising, developers don't have to wait for a closed API just to test an agentic coding system on their own stack.

That's the real story.

## The numbers need a skeptical read

Ornith says the 397B version scored 86.1 on Terminal-Bench 2.1, 86 on SWE-Bench Verified, 65.1 on SWE-Bench Pro and 79.6 on SWE-Bench Multilingual. It also reported 56 on DeepSWE, 44.6 on Humanity's Last Exam, 81.4 on ClawEval and 71.2 on Tool Decathlon. Those are not small claims. They put Ornith-1.5 in the same conversation as the strongest closed coding agents, at least on the benchmarks the lab chose to publish.

[Microsoft's MAI-Image-2.5-Pro Just Topped an AI Image Editing Leaderboard](https://startupfortune.com/microsofts-mai-image-25-pro-just-topped-an-ai-image-editing-leaderboard/)

Microsoft's MAI-Image-2.5-Pro debuted at number one on Artificial Analysis's Image Editing Leaderboard, beating GPT Image 2 and Reve 2.1 less than a month after its July 23 preview launch on Microsoft Foundry. The model prices at roughly $108.50 per thousand images, more than double its own mid-tier sibling. - [microsoft mai image 2.5 pro leaderboard ranking](https://startupfortune.com/microsofts-mai-image-25-pro-just-topped-an-ai-image-editing-leaderboard/) - [ai image editing model performance comparison 2026](https://startupfortune.com/microsofts-mai-image-25-pro-just-topped-an-ai-image-editing-leaderboard/)

But benchmark claims from a model creator are not the same as independent measurement. You don't need to dismiss them. You do need to separate the announcement from the proof. Terminal-Bench's public leaderboard currently lists Claude Code with Opus 4.8 at 78.9% on Terminal-Bench 2.1, while Ornith-1.0-397B sits at 77.5% in the same public view. Ornith-1.5's 86.1 score, if reproduced, would be a serious jump. Until it is reproduced, it is a vendor-reported result.

That caveat matters because Ornith-1.0 already showed how fast this field can turn into leaderboard theater. The earlier 397B model looked strong for an open release, with Hugging Face's model card showing 82.4 on SWE-Bench Verified and 77.5 on Terminal-Bench 2.1. It still trailed Claude Opus 4.8 on the main figures Ornith itself published for version 1.0. The new release is therefore not just an upgrade. It is a claim that the gap has closed.

## The training loop is the bet

Ornith's more interesting claim is not one benchmark row. It is the training method. In Ornith-1.0, the lab described self-scaffolding: the model learned to build the harness around a coding task, not only the answer. MarkTechPost covered that June release and noted the same core idea, a reinforcement-learning setup where the scaffold and solution are optimized together.

Ornith-1.5 pushes the loop further. The lab says the model now proposes new tasks, generates task-specific scaffolds and produces solution rollouts that feed back into reinforcement learning. If that works, the model is not simply consuming a fixed pile of human-made coding problems. It is making some of its own practice material.

That is a meaningful shift for anyone building with open models. Curated coding data is slow to collect, hard to clean and easy to overfit. A model that can create useful tasks for itself could cut one of the most awkward bottlenecks in agent training. Frankly, that is more important than whether one leaderboard line beats Claude by a point.

The open-model field is not waiting around. Alibaba's Qwen family, Google's Gemma line, DeepSeek, MiniMax and GLM models have all made coding and tool use a central fight. Ornith-1.0 was built on Gemma 4 and Qwen 3.5 foundations, according to its model cards and GitHub materials, so this is not a clean-room challenger appearing from nowhere. It is a post-training bet on top of strong existing base models.

For developers, the 9B release may end up being the most useful piece. A 397B MoE model is a statement, not a casual download. A 9B model with quantized builds is the one you can actually test on a local machine or a small server. If it handles repository work well enough, it gives small teams something more interesting than another chat model with a coding badge.

[AI Agents Autonomously Hacked Taiwan's Government for Four Days Straight](https://startupfortune.com/ai-agents-autonomously-hacked-taiwans-government-for-four-days-straight/)

Researchers at the Israeli-Austrian firm Dream say suspected China-linked hackers used open-source AI agents to autonomously breach Taiwan's government over four days in July, compromising 85 accounts and stealing 2,500 records with almost no human oversight. The campaign, built on the Hermes and OpenClaw frameworks, later expanded to Taiwan's... - [ai agents autonomous hacking government systems](https://startupfortune.com/ai-agents-autonomously-hacked-taiwans-government-for-four-days-straight/) - [taiwan government data breach four days](https://startupfortune.com/ai-agents-autonomously-hacked-taiwans-government-for-four-days-straight/)

**Also read:** [Microsoft's MAI-Image-2.5-Pro Just Topped an AI Image Editing Leaderboard](https://startupfortune.com/microsofts-mai-image-25-pro-just-topped-an-ai-image-editing-leaderboard/) • [Reddit Users Say OpenAI Quietly Cut Codex Usage Limits in Half](https://startupfortune.com/reddit-users-say-openai-quietly-cut-codex-usage-limits-in-half/) • [AI Agents Autonomously Hacked Taiwan's Government for Four Days Straight](https://startupfortune.com/ai-agents-autonomously-hacked-taiwans-government-for-four-days-straight/)
