cd /news/large-language-models/mercury-2-5-diffusion-llm-1107-t-s-i… · home topics large-language-models article
[ARTICLE · art-123998] src=byteiota.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Mercury 2.5 Diffusion LLM: 1,107 t/s in Production Now

Inception Labs released Mercury 2.5, a diffusion-based large language model, on September 8, achieving 1,107 tokens per second on commodity NVIDIA GPUs at $0.04 per million input tokens. Unlike autoregressive models such as GPT-6, Claude Fable, and Llama, Mercury 2.5 generates tokens in parallel, breaking the traditional speed-quality tradeoff for budget LLM deployments.

by read1 min views3 publishedSep 9, 2026

Inception Labs shipped Mercury 2.5 on September 8 — a diffusion language model generating 1,107 tokens per second on commodity NVIDIA GPUs at $0.04 per million input tokens. For teams running voice agents, real-time search, or high-volume RAG pipelines, that combination breaks what developers have accepted as the speed-quality tradeoff in the budget LLM tier. What Mercury 2.5 Actually Does Differently Most LLMs — GPT-6, Claude Fable, Llama — are autoregressive: they generate one token at a time, left to right, each token dependent on the one before it. That sequential dependency is the fundamental throughput ceiling, regardless of how […]

The post

── more in #large-language-models 4 stories · sorted by recency
inceptionlabs.ai · · #large-language-models
Mercury 2.5
── more on @inception labs 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/mercury-2-5-diffusio…] indexed:0 read:1min 2026-09-09 ·