premierAn anonymous “stealth” reasoning model served free on OpenRouter (stealth/ox-alpha
) starting Aug 20, 2026. The operator is undisclosed; community tokenizer + error-code fingerprinting strongly suggests Z.ai’s GLM-5.x family (an unreleased multimodal sibling of GLM-5.3), but no lab has claimed it.
Specs: 1M context, 131K max output, text/image/video in, text out, mandatory reasoning (low/high/max). Free during the preview (~$0 in/out).
Agents on Rails benchmark (Aug 2026, Le Mans round). 82.5% accuracy on 63 runs - tied with Grok 4.6 and behind only the Opus 5 / Kimi K3 / Fable 5 cluster. API recall 28.6%. Slow (19m 26s median, the third-slowest field) but strong. Treat identity and numbers as preliminary until the reveal; if the GLM fingerprint holds, it is a mid-tier frontier coder at preview price.
- 1000k
- proprietary
- Undisclosed
- Aug 2026
Scores #
Or run it in the cloud #
Live per-provider pricing, throughput and uptime - refreshed about 1 hour ago via OpenRouter. Click a column to sort.
| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
| API | 0.00 | 0.00 | - | - | - | - | cheapest |
Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.
Detailed API pricing page + JSON endpoint →
Inference cost over time #
Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.