{"slug": "inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40", "title": "Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%", "summary": "AI startup Inception launched Mercury 2.5, a diffusion-based large language model that delivers a 40% intelligence improvement over its predecessor while processing 1,107 tokens per second on standard NVIDIA GPUs. The model, priced at $0.20 per million input tokens and $0.75 per million output tokens (with an 80% launch discount), targets latency-sensitive production workloads and competes with GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite. Inception, which raised $50 million led by Menlo Ventures with backing from Andrew Ng and Andrej Karpathy, claims the model's diffusion architecture enables simultaneous token generation and a 260K token context window.", "body_md": "# Inception launches Mercury 2.5 diffusion model, boosts intelligence by 40%\n\nThe AI startup's latest diffusion-based language model processes over 1,100 tokens per second at a fraction of competitors' costs\n\nInception just dropped a new large language model that takes a fundamentally different approach to generating text. Mercury 2.5, the company’s latest diffusion-based LLM, delivers a 40% intelligence improvement over its predecessor while maintaining the kind of speed and cost profile that makes enterprise CFOs smile.\n\nThe model processes 1,107 tokens per second on standard [NVIDIA](https://cryptobriefing.com/markets/nvidia/) GPUs.\n\n## What makes diffusion models different\n\nMost large language models you’ve interacted with, think GPT or Claude, generate text one token at a time in sequence. They’re autoregressive, meaning each word depends on the one before it. Diffusion models work differently. They generate multiple tokens simultaneously, more like how an image diffusion model creates a picture by gradually refining noise into something coherent.\n\nMercury 2.5 comes with a 260K token context window, which means it can process roughly the equivalent of a 500-page book in a single prompt. It also supports tunable reasoning levels, letting developers dial the model’s thinking depth up or down depending on whether they need deep analysis or quick responses. Parallel tool calls and structured JSON output round out the feature set.\n\nInception claims the model competes with GPT-5.6 Luna (Low) and Gemini 3.5 Flash-Lite.\n\n## The pricing play\n\nMercury 2.5’s standard pricing sits at $0.20 per million input tokens and $0.75 per million output tokens. Inception is offering an 80% discount during the launch phase that brings costs down to $0.04 per million input tokens and $0.15 per million output tokens.\n\nThe model ships via an OpenAI-compatible API, meaning any application already wired to talk to OpenAI’s endpoints can switch to Mercury 2.5 with minimal code changes.\n\nEarly adopters are already putting the model through its paces. OpenCall, a voice AI company, reported median latencies under 200 milliseconds for voice interactions using Mercury 2.5.\n\n## The company behind the model\n\nInception raised $50 million in funding led by Menlo Ventures, with backing from Andrew Ng, the Stanford professor and former head of [Google](https://cryptobriefing.com/markets/alphabet/) Brain, and Andrej Karpathy, who previously led AI at [Tesla](https://cryptobriefing.com/markets/tesla/).\n\nMercury 2.5 is the company’s second major model release, building on the Mercury 2 foundation with the claimed 40% intelligence boost. The launch targets latency-sensitive production workloads specifically, including voice agents, coding assistants, and enterprise search.\n\n**Disclosure:** This article was edited by Editorial Team. For more information on how we create and review content, see our\n\n[Editorial Policy](https://cryptobriefing.com/editorial-policy/).", "url": "https://wpnews.pro/news/inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40", "canonical_source": "https://cryptobriefing.com/inception-mercury-2-5-diffusion-model-launch/", "published_at": "2026-09-08 17:18:17+00:00", "updated_at": "2026-09-08 17:29:33.881138+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-startups", "ai-products"], "entities": ["Inception", "Mercury 2.5", "NVIDIA", "Menlo Ventures", "Andrew Ng", "Andrej Karpathy", "OpenCall", "GPT-5.6 Luna (Low)"], "alternates": {"html": "https://wpnews.pro/news/inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40", "markdown": "https://wpnews.pro/news/inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40.md", "text": "https://wpnews.pro/news/inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40.txt", "jsonld": "https://wpnews.pro/news/inception-launches-mercury-2-5-diffusion-model-boosts-intelligence-by-40.jsonld"}}