cd /news/large-language-models/inception-mercury-2-5-preview-on-ope… · home topics large-language-models article
[ARTICLE · art-118323] src=openrouter.ai ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Inception: Mercury 2.5 Preview on OpenRouter

Inception released Mercury 2.5, a diffusion large language model (dLLM) that generates tokens in parallel, achieving 1,107 tokens/sec on standard GPUs and a 10+ point intelligence jump over Mercury 2, with pricing at $0.04 per 1M input tokens and $0.15 per 1M output tokens on OpenRouter, available with an 80% discount through September 8, 2026. The model supports 260K context, tunable reasoning levels, parallel tool calls, and schema-aligned JSON output, targeting production workloads like search agents, voice pipelines, and coding subagents.

read1 min views1 publishedSep 1, 2026
Inception: Mercury 2.5 Preview on OpenRouter
Image: source

Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC.

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.

Modalities

In / Out Price

$0.04 / $0.15per 1M

Context

260K

Released

Aug 31, 2026

80% off | $0.20$0.04 | $0.75$0.15 | $0.02$0.004 | 1.07s | 107 tps |

Throughput

107tok/s

P50, best across providers

Latency

1.07s

P50, best provider

100.00%

97.12%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

── more in #large-language-models 4 stories · sorted by recency
── more on @inception 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/inception-mercury-2-…] indexed:0 read:1min 2026-09-01 ·