cd /news/artificial-intelligence/inception-launches-mercury-2-5-diffu… · home topics artificial-intelligence article
[ARTICLE · art-123643] src=cryptobriefing.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%

AI startup Inception launched Mercury 2.5, a diffusion-based large language model that delivers a 40% intelligence improvement over its predecessor while processing 1,107 tokens per second on standard NVIDIA GPUs. The model, priced at $0.20 per million input tokens and $0.75 per million output tokens (with an 80% launch discount), targets latency-sensitive production workloads and competes with GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite. Inception, which raised $50 million led by Menlo Ventures with backing from Andrew Ng and Andrej Karpathy, claims the model's diffusion architecture enables simultaneous token generation and a 260K token context window.

read2 min views3 publishedSep 8, 2026
Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%
Image: Cryptobriefing (auto-discovered)

The AI startup's latest diffusion-based language model processes over 1,100 tokens per second at a fraction of competitors' costs

Inception just dropped a new large language model that takes a fundamentally different approach to generating text. Mercury 2.5, the company’s latest diffusion-based LLM, delivers a 40% intelligence improvement over its predecessor while maintaining the kind of speed and cost profile that makes enterprise CFOs smile.

The model processes 1,107 tokens per second on standard NVIDIA GPUs.

What makes diffusion models different #

Most large language models you’ve interacted with, think GPT or Claude, generate text one token at a time in sequence. They’re autoregressive, meaning each word depends on the one before it. Diffusion models work differently. They generate multiple tokens simultaneously, more like how an image diffusion model creates a picture by gradually refining noise into something coherent.

Mercury 2.5 comes with a 260K token context window, which means it can process roughly the equivalent of a 500-page book in a single prompt. It also supports tunable reasoning levels, letting developers dial the model’s thinking depth up or down depending on whether they need deep analysis or quick responses. Parallel tool calls and structured JSON output round out the feature set.

Inception claims the model competes with GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite.

The pricing play #

Mercury 2.5’s standard pricing sits at $0.20 per million input tokens and $0.75 per million output tokens. Inception is offering an 80% discount during the launch phase that brings costs down to $0.04 per million input tokens and $0.15 per million output tokens.

The model ships via an OpenAI-compatible API, meaning any application already wired to talk to OpenAI’s endpoints can switch to Mercury 2.5 with minimal code changes.

Early adopters are already putting the model through its paces. OpenCall, a voice AI company, reported median latencies under 200 milliseconds for voice interactions using Mercury 2.5.

The company behind the model #

Inception raised $50 million in funding led by Menlo Ventures, with backing from Andrew Ng, the Stanford professor and former head of Google Brain, and Andrej Karpathy, who previously led AI at Tesla.

Mercury 2.5 is the company’s second major model release, building on the Mercury 2 foundation with the claimed 40% intelligence boost. The launch targets latency-sensitive production workloads specifically, including voice agents, coding assistants, and enterprise search.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our

Editorial Policy.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @inception 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/inception-launches-m…] indexed:0 read:2min 2026-09-08 ·