# Inception: Mercury 2.5 Preview on OpenRouter

> Source: <https://openrouter.ai/inception/mercury-2.5-preview>
> Published: 2026-09-01 22:38:00+00:00

Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC.

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.

Modalities

In / Out Price

$0.04 / $0.15per 1M

Context

260K

Released

Aug 31, 2026

80% off | $0.20$0.04 | $0.75$0.15 | $0.02$0.004 | 1.07s | 107 tps |

Throughput

107tok/s

P50, best across providers

Latency

1.07s

P50, best provider

100.00%

97.12%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the [Endpoints API](/docs/api/api-reference/endpoints/list-endpoints). [Learn more](/docs/provider-routing) about our load balancing and customization options.
