Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC.
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.
Modalities
In / Out Price
$0.04 / $0.15per 1M
Context
260K
Released
Aug 31, 2026
80% off | $0.20$0.04 | $0.75$0.15 | $0.02$0.004 | 1.07s | 107 tps |
Throughput
107tok/s
P50, best across providers
Latency
1.07s
P50, best provider
100.00%
97.12%
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.