consumer**~9B Dense** - the edge-deployable Ornith-1.5, released 2026-08-18. Compact enough for phones (a quantized Ornith-1.5-9B-Mobile variant targets iOS/Android) yet matches or exceeds much larger models like Gemma 4-31B and Qwen 3.6-35B on agentic coding. 262K context, MIT-licensed on HuggingFace at ornith-ai/Ornith-1.5-9B
(GGUF and MLX quantizations). #
**Coding (vendor self-reported):** Terminal-Bench 2.1 46.2, SWE-bench Verified 70.6, SWE-bench Pro 47.5, NL2Repo 32.4. -
**Reasoning:** HLE 20.2 (no tools) / 30.5 (with tools), GPQA-Diamond 86.4. -
Agentic: MCP-Atlas 54.2, Toolathlon-Verified 41.2, ClawEval 66.5.
Runs almost anywhere. Quantized builds run on phones, laptops, and small GPUs; the strongest open-weight coding model at this size. Vendor benchmarks are claims pending independent replication.
- 9.0B
- 262k
- mit
- 🇺🇸 USA
- Aug 2026
Scores #
Run it locally #
Per-quant memory needs and a static "can you run it?" reference - no rig entry required
Can you run it? - reference rigs
| Rig | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | BF16 |
|---|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 1274.7t/s | fast 1109.6t/s | fast 974.6t/s | fast 752.7t/s | fast 400.5t/s |
| NVIDIA DGX Station 748GB | fast 761.0t/s | fast 662.5t/s | fast 581.9t/s | fast 449.4t/s | fast 239.1t/s |
| 4x RTX 5090 (128GB) | fast 681.8t/s | fast 593.6t/s | fast 521.3t/s | fast 402.6t/s | fast 214.2t/s |
| 4x RTX 4090 (96GB) | fast 383.5t/s | fast 333.9t/s | fast 293.3t/s | fast 226.5t/s | fast 120.5t/s |
| 2x RTX 5090 (64GB) | fast 340.9t/s | fast 296.8t/s | fast 260.7t/s | fast 201.3t/s | fast 107.1t/s |
| 2x RTX 3090 (48GB) | fast 178.1t/s | fast 155.1t/s | fast 136.2t/s | fast 105.2t/s | fast 56.0t/s |
| Single RTX 5090 (32GB) | fast 170.5t/s | fast 148.4t/s | fast 130.3t/s | fast 100.7t/s | fast 53.6t/s |
| Mac Studio M4 Ultra 192GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s |
| Mac Studio M4 Ultra 512GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s |
| Single RTX 4090 (24GB) | fast 95.9t/s | fast 83.5t/s | fast 73.3t/s | fast 56.6t/s | fast 30.1t/s |
| MacBook Pro M5 Max 128GB | fast 63.7t/s | fast 55.5t/s | fast 48.7t/s | fast 37.6t/s | fast 20.0t/s |
| Single GTX 1080 Ti (11GB) | fast 46.0t/s | fast 40.1t/s | fast 35.2t/s | tight | |
no -> cloudFit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.
Download options #
Or run it in the cloud #
No per-token API provider pricing tracked for Ornith-1.5-9B yet.
For flagship list prices, see the
[calculator](/calculator).
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.