Published benchmark scores #
Published scores. Tinted cells mark the best score on each test, including ties.
| Benchmark | Pareto 26.9 | Fable 5.1 | GPT 6 Astra | DeepSeek 4.1 Flash |
|---|---|---|---|---|
| DeepSWE | 74 | 67 | 74 | 74 |
| Terminal-Bench 4.0 | 51 | 56 | 58 | 31 |
| MMMU-Pro | 78 | 81 | 87 | 77 |
| HLE (no tools) | 49 | 55 | 54 | 39 |
| ArXivMath | 88 | 72 | 91 | 28 |
Measured task costs and a composite score have not been published for this release. Scores do not establish cost per completed task.
Details and pricing #
Pareto accepts text and image inputs. Use pareto as the model identifier. The published comparison set is Fable 5.1, GPT 6 Astra, and DeepSeek 4.1 Flash.
Per million tokens: $2.50 input, $0.25 cached input, and $7.50 output. Cached input is billed at 0.1 times the input rate. Actual spend depends on token usage and cache hits.