# Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%

> Source: <https://cryptobriefing.com/inception-mercury-2-5-diffusion-model-launch/>
> Published: 2026-09-08 17:18:17+00:00

# Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%

The AI startup's latest diffusion-based language model processes over 1,100 tokens per second at a fraction of competitors' costs

Inception just dropped a new large language model that takes a fundamentally different approach to generating text. Mercury 2.5, the company’s latest diffusion-based LLM, delivers a 40% intelligence improvement over its predecessor while maintaining the kind of speed and cost profile that makes enterprise CFOs smile.

The model processes 1,107 tokens per second on standard [NVIDIA](https://cryptobriefing.com/markets/nvidia/) GPUs.

## What makes diffusion models different

Most large language models you’ve interacted with, think GPT or Claude, generate text one token at a time in sequence. They’re autoregressive, meaning each word depends on the one before it. Diffusion models work differently. They generate multiple tokens simultaneously, more like how an image diffusion model creates a picture by gradually refining noise into something coherent.

Mercury 2.5 comes with a 260K token context window, which means it can process roughly the equivalent of a 500-page book in a single prompt. It also supports tunable reasoning levels, letting developers dial the model’s thinking depth up or down depending on whether they need deep analysis or quick responses. Parallel tool calls and structured JSON output round out the feature set.

Inception claims the model competes with GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite.

## The pricing play

Mercury 2.5’s standard pricing sits at $0.20 per million input tokens and $0.75 per million output tokens. Inception is offering an 80% discount during the launch phase that brings costs down to $0.04 per million input tokens and $0.15 per million output tokens.

The model ships via an OpenAI-compatible API, meaning any application already wired to talk to OpenAI’s endpoints can switch to Mercury 2.5 with minimal code changes.

Early adopters are already putting the model through its paces. OpenCall, a voice AI company, reported median latencies under 200 milliseconds for voice interactions using Mercury 2.5.

## The company behind the model

Inception raised $50 million in funding led by Menlo Ventures, with backing from Andrew Ng, the Stanford professor and former head of [Google](https://cryptobriefing.com/markets/alphabet/) Brain, and Andrej Karpathy, who previously led AI at [Tesla](https://cryptobriefing.com/markets/tesla/).

Mercury 2.5 is the company’s second major model release, building on the Mercury 2 foundation with the claimed 40% intelligence boost. The launch targets latency-sensitive production workloads specifically, including voice agents, coding assistants, and enterprise search.

**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our

[Editorial Policy](https://cryptobriefing.com/editorial-policy/).
