{"slug": "mercury-2-5-diffusion-llm-1107-t-s-in-production-now", "title": "Mercury 2.5 Diffusion LLM: 1,107 t/s in Production Now", "summary": "Inception Labs released Mercury 2.5, a diffusion-based large language model, on September 8, achieving 1,107 tokens per second on commodity NVIDIA GPUs at $0.04 per million input tokens. Unlike autoregressive models such as GPT-6, Claude Fable, and Llama, Mercury 2.5 generates tokens in parallel, breaking the traditional speed-quality tradeoff for budget LLM deployments.", "body_md": "Inception Labs shipped Mercury 2.5 on September 8 — a diffusion language model generating 1,107 tokens per second on commodity NVIDIA GPUs at $0.04 per million input tokens. For teams running voice agents, real-time search, or high-volume RAG pipelines, that combination breaks what developers have accepted as the speed-quality tradeoff in the budget LLM tier. What Mercury 2.5 Actually Does Differently Most LLMs — GPT-6, Claude Fable, Llama — are autoregressive: they generate one token at a time, left to right, each token dependent on the one before it. That sequential dependency is the fundamental throughput ceiling, regardless of how […]\n\nThe post", "url": "https://wpnews.pro/news/mercury-2-5-diffusion-llm-1107-t-s-in-production-now", "canonical_source": "https://byteiota.com/mercury-25-diffusion-llm-1107-tokens-per-second/", "published_at": "2026-09-09 00:10:56+00:00", "updated_at": "2026-09-09 00:15:58.878914+00:00", "lang": "en", "topics": ["large-language-models", "generative-ai", "ai-products", "ai-infrastructure"], "entities": ["Inception Labs", "Mercury 2.5", "NVIDIA", "GPT-6", "Claude Fable", "Llama"], "alternates": {"html": "https://wpnews.pro/news/mercury-2-5-diffusion-llm-1107-t-s-in-production-now", "markdown": "https://wpnews.pro/news/mercury-2-5-diffusion-llm-1107-t-s-in-production-now.md", "text": "https://wpnews.pro/news/mercury-2-5-diffusion-llm-1107-t-s-in-production-now.txt", "jsonld": "https://wpnews.pro/news/mercury-2-5-diffusion-llm-1107-t-s-in-production-now.jsonld"}}