Ornith-1.5-9B Ornith AI released Ornith-1.5-9B, a 9-billion-parameter dense language model, on August 18, 2026, claiming it matches or exceeds larger models like Gemma 4-31B and Qwen 3.6-35B on agentic coding tasks. The model, available under an MIT license on HuggingFace, scores 70.6 on SWE-bench Verified and 86.4 on GPQA-Diamond, with a 262K context window and quantized builds for edge devices. Ornith-1.5-9B consumer ~9B Dense - the edge-deployable Ornith-1.5, released 2026-08-18. Compact enough for phones a quantized Ornith-1.5-9B-Mobile variant targets iOS/Android yet matches or exceeds much larger models like Gemma 4-31B and Qwen 3.6-35B on agentic coding. 262K context, MIT-licensed on HuggingFace at ornith-ai/Ornith-1.5-9B GGUF and MLX quantizations . - Coding vendor self-reported : Terminal-Bench 2.1 46.2, SWE-bench Verified 70.6, SWE-bench Pro 47.5, NL2Repo 32.4. - Reasoning: HLE 20.2 no tools / 30.5 with tools , GPQA-Diamond 86.4. - Agentic: MCP-Atlas 54.2, Toolathlon-Verified 41.2, ClawEval 66.5. Runs almost anywhere. Quantized builds run on phones, laptops, and small GPUs; the strongest open-weight coding model at this size. Vendor benchmarks are claims pending independent replication. - 9.0B - 262k - mit - 🇺🇸 USA - Aug 2026 Scores Run it locally Per-quant memory needs and a static "can you run it?" reference - no rig entry required Can you run it? - reference rigs | Rig | Q4 K M | Q5 K M | Q6 K | Q8 0 | BF16 | |---|---|---|---|---|---| | 4x H100 80GB 320GB | fast 1274.7t/s | fast 1109.6t/s | fast 974.6t/s | fast 752.7t/s | fast 400.5t/s | | NVIDIA DGX Station 748GB | fast 761.0t/s | fast 662.5t/s | fast 581.9t/s | fast 449.4t/s | fast 239.1t/s | | 4x RTX 5090 128GB | fast 681.8t/s | fast 593.6t/s | fast 521.3t/s | fast 402.6t/s | fast 214.2t/s | | 4x RTX 4090 96GB | fast 383.5t/s | fast 333.9t/s | fast 293.3t/s | fast 226.5t/s | fast 120.5t/s | | 2x RTX 5090 64GB | fast 340.9t/s | fast 296.8t/s | fast 260.7t/s | fast 201.3t/s | fast 107.1t/s | | 2x RTX 3090 48GB | fast 178.1t/s | fast 155.1t/s | fast 136.2t/s | fast 105.2t/s | fast 56.0t/s | | Single RTX 5090 32GB | fast 170.5t/s | fast 148.4t/s | fast 130.3t/s | fast 100.7t/s | fast 53.6t/s | | Mac Studio M4 Ultra 192GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s | | Mac Studio M4 Ultra 512GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s | | Single RTX 4090 24GB | fast 95.9t/s | fast 83.5t/s | fast 73.3t/s | fast 56.6t/s | fast 30.1t/s | | MacBook Pro M5 Max 128GB | fast 63.7t/s | fast 55.5t/s | fast 48.7t/s | fast 37.6t/s | fast 20.0t/s | | Single GTX 1080 Ti 11GB | fast 46.0t/s | fast 40.1t/s | fast 35.2t/s | tight | | no - cloud cloud-pricing Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast =20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark. Download options Or run it in the cloud No per-token API provider pricing tracked for Ornith-1.5-9B yet. For flagship list prices, see the calculator /calculator . Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.