{"slug": "inception-mercury-2-5-preview-on-openrouter", "title": "Inception: Mercury 2.5 Preview on OpenRouter", "summary": "Inception released Mercury 2.5, a diffusion large language model (dLLM) that generates tokens in parallel, achieving 1,107 tokens/sec on standard GPUs and a 10+ point intelligence jump over Mercury 2, with pricing at $0.04 per 1M input tokens and $0.15 per 1M output tokens on OpenRouter, available with an 80% discount through September 8, 2026. The model supports 260K context, tunable reasoning levels, parallel tool calls, and schema-aligned JSON output, targeting production workloads like search agents, voice pipelines, and coding subagents.", "body_md": "Limited-time 80% discount via Inception through September 8, 2026 at 07:00 UTC.\n\nMercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents.\n\nModalities\n\nIn / Out Price\n\n$0.04 / $0.15per 1M\n\nContext\n\n260K\n\nReleased\n\nAug 31, 2026\n\n80% off | $0.20$0.04 | $0.75$0.15 | $0.02$0.004 | 1.07s | 107 tps |\n\nThroughput\n\n107tok/s\n\nP50, best across providers\n\nLatency\n\n1.07s\n\nP50, best provider\n\n100.00%\n\n97.12%\n\nWhen an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the [Endpoints API](/docs/api/api-reference/endpoints/list-endpoints). [Learn more](/docs/provider-routing) about our load balancing and customization options.", "url": "https://wpnews.pro/news/inception-mercury-2-5-preview-on-openrouter", "canonical_source": "https://openrouter.ai/inception/mercury-2.5-preview", "published_at": "2026-09-01 22:38:00+00:00", "updated_at": "2026-09-01 22:52:07.748870+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "ai-infrastructure"], "entities": ["Inception", "Mercury 2.5", "OpenRouter", "GPT-5.6 Luna", "Gemini 3.5 Flash-Lite", "Claude Haiku 4.5"], "alternates": {"html": "https://wpnews.pro/news/inception-mercury-2-5-preview-on-openrouter", "markdown": "https://wpnews.pro/news/inception-mercury-2-5-preview-on-openrouter.md", "text": "https://wpnews.pro/news/inception-mercury-2-5-preview-on-openrouter.txt", "jsonld": "https://wpnews.pro/news/inception-mercury-2-5-preview-on-openrouter.jsonld"}}