# Ornith-1.5-9B

> Source: <https://tokenstead.ai/models/ornith-1-5-9b>
> Published: 2026-08-24 20:51:38+00:00

# Ornith-1.5-9B

consumer**~9B Dense** - the edge-deployable Ornith-1.5, released 2026-08-18. Compact enough for phones (a quantized Ornith-1.5-9B-Mobile variant targets iOS/Android) yet matches or exceeds much larger models like Gemma 4-31B and Qwen 3.6-35B on agentic coding. 262K context, MIT-licensed on HuggingFace at `ornith-ai/Ornith-1.5-9B`

(GGUF and MLX quantizations).

-
**Coding (vendor self-reported):** Terminal-Bench 2.1 46.2, SWE-bench Verified 70.6, SWE-bench Pro 47.5, NL2Repo 32.4. -
**Reasoning:** HLE 20.2 (no tools) / 30.5 (with tools), GPQA-Diamond 86.4. -
**Agentic:** MCP-Atlas 54.2, Toolathlon-Verified 41.2, ClawEval 66.5.

**Runs almost anywhere.** Quantized builds run on phones, laptops, and small GPUs; the strongest open-weight coding model at this size. Vendor benchmarks are claims pending independent replication.

- 9.0B
- 262k
- mit
- 🇺🇸 USA
- Aug 2026

## Scores

## Run it locally

Per-quant memory needs and a static "can you run it?" reference - no rig entry required

### Can you run it? - reference rigs

| Rig | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | BF16 |
|---|---|---|---|---|---|
| 4x H100 80GB (320GB) | fast 1274.7t/s | fast 1109.6t/s | fast 974.6t/s | fast 752.7t/s | fast 400.5t/s |
| NVIDIA DGX Station 748GB | fast 761.0t/s | fast 662.5t/s | fast 581.9t/s | fast 449.4t/s | fast 239.1t/s |
| 4x RTX 5090 (128GB) | fast 681.8t/s | fast 593.6t/s | fast 521.3t/s | fast 402.6t/s | fast 214.2t/s |
| 4x RTX 4090 (96GB) | fast 383.5t/s | fast 333.9t/s | fast 293.3t/s | fast 226.5t/s | fast 120.5t/s |
| 2x RTX 5090 (64GB) | fast 340.9t/s | fast 296.8t/s | fast 260.7t/s | fast 201.3t/s | fast 107.1t/s |
| 2x RTX 3090 (48GB) | fast 178.1t/s | fast 155.1t/s | fast 136.2t/s | fast 105.2t/s | fast 56.0t/s |
| Single RTX 5090 (32GB) | fast 170.5t/s | fast 148.4t/s | fast 130.3t/s | fast 100.7t/s | fast 53.6t/s |
| Mac Studio M4 Ultra 192GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s |
| Mac Studio M4 Ultra 512GB | fast 113.3t/s | fast 98.6t/s | fast 86.6t/s | fast 66.9t/s | fast 35.6t/s |
| Single RTX 4090 (24GB) | fast 95.9t/s | fast 83.5t/s | fast 73.3t/s | fast 56.6t/s | fast 30.1t/s |
| MacBook Pro M5 Max 128GB | fast 63.7t/s | fast 55.5t/s | fast 48.7t/s | fast 37.6t/s | fast 20.0t/s |
| Single GTX 1080 Ti (11GB) | fast 46.0t/s | fast 40.1t/s | fast 35.2t/s | tight |
|

[no -> cloud](#cloud-pricing)Fit tiers use the same will-it-run logic as the rig finder. For comfortable fits, the badge reflects decode speed: fast >=20 t/s, ok 8-20 t/s, slow <8 t/s. t/s is a bandwidth estimate, not a measured benchmark.

## Download options

## Or run it in the cloud

No per-token API provider pricing tracked for Ornith-1.5-9B yet.
For flagship list prices, see the
[calculator](/calculator).

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
