cd /news/artificial-intelligence/inception-launches-mercury-voice-for… · home › topics › artificial-intelligence › article
[ARTICLE · art-142039] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Inception launches Mercury Voice for faster AI phone agents

Inception made Mercury Voice generally available to enterprise customers on September 29th, claiming a company-reported 320-millisecond median time to first answer token on production customer-service prompts, below the roughly 500-millisecond conversational target it uses for comparison. The diffusion-based language model supports tool calls and long system prompts, with a reported 95th-percentile time to first answer token of 750 milliseconds, and is priced at $0.20 per million input tokens at launch. Inception says Mercury Voice scored above GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash and Qwen3.5-397B on a composite of tau3-bench Telecom, Retail and Airline, IFBench and BFCL v4, though the benchmark results are company-reported and measure when the model begins producing an answer rather than the full voice round trip.

by read4 min views3 publishedSep 29, 2026
Inception launches Mercury Voice for faster AI phone agents
Image: Runtimewire (auto-discovered)

The diffusion model posts a company-reported 320-millisecond median time to first answer token; enterprise customers can access it at launch pricing of $0.20 per million input tokens.

        By [Ryan Merket](https://runtimewire.com/author/ryan-merket)
        · Published 

Primary source: [X](https://x.com/_inception_ai/status/2104974439314321660)

Why it matters #

Voice agents need to answer quickly while following instructions and using tools. Mercury Voice is Inception’s attempt to make parallel diffusion generation competitive in that tight latency window, though its published timings measure the model rather than the full voice call.

Inception made Mercury Voice generally available to enterprise customers on September 29th, pitching its diffusion-based language model as a way to give voice agents time to reason and call tools without leaving callers in silence. The company reports a 320-millisecond median time to first answer token on production customer-service prompts, below the roughly 500-millisecond conversational target it uses for comparison.

https://x.com/_inception_ai/status/2104974439314321660 That speed claim is central to the company’s bet. Inception co-founder and CEO Stefano Ermon is an associate professor of computer science at Stanford whose research in generative AI helped lead to the company’s diffusion approach. Rather than generating text one token at a time, as conventional autoregressive models do, diffusion language models refine multiple tokens in parallel. Inception’s argument is that this can leave time for a model to reason and use tools before a voice agent answers.

The company says Mercury Voice supports tool calls and long system prompts, capabilities that matter when a phone agent has to follow a workflow instead of simply respond to a basic question. Inception says the model’s 95th-percentile time to first answer token is 750 milliseconds. Both figures measure when the model begins producing an answer, not the complete round trip from a caller’s speech through transcription, model processing and synthesized audio. The overall will also depend on the rest of the voice pipeline.

Inception’s benchmark results are company-reported. The firm says Mercury Voice scored above models including GPT-6 Luna, Gemma 4 31B, GLM-5.3-Flash and Qwen3.5-397B on a composite of tau3-bench Telecom, Retail and Airline, IFBench and BFCL v4. In its comparison, Mercury Voice was the only model with a median below 500 milliseconds; the company also reports that its 750-millisecond p95 beat the median for all but one of the tested models. The company describes the result as more than twice the speed of comparison models, but the benchmark details and selected model settings determine how broadly that comparison applies.

The release builds on Inception’s previous focus on speed. The company introduced Mercury 2.5 earlier in September, saying it had grown usage across search, voice and coding workloads. Mercury Voice was previewed alongside that model before this enterprise availability announcement. Inception’s product explanation describes diffusion generation as an iterative process that revises an output in parallel, in contrast to the sequential token generation used by autoregressive systems.

Inception named three customers in its launch post. Audivi AI uses Mercury Voice for automated drive-through ordering, including order changes and upsells. Altur says it uses the model for financial-institution calls involving payment-plan negotiations. OpenCall says the model brought median response latency close to 170 milliseconds on its production workload. Those customer results are distinct from Inception’s benchmark figures: they reflect particular deployments, and the cited figures do not establish how other customers’ systems will perform.

The model is priced at $0.40 per million input tokens and $1.50 per million output tokens, with launch pricing at half those rates. Inception estimates a typical voice-agent workload would cost about $0.009 per conversation minute. That is a model-layer estimate; speech recognition, text-to-speech and telephony add costs. Enterprise customers can contact Inception for access.

The company has raised a $50 million seed round, announced in November 2025 and led by Menlo Ventures, with participation from Mayfield, Innovation Endeavors, Microsoft’s M12, Snowflake Ventures, Databricks Investment and NVentures, Nvidia’s venture arm. Andrew Ng and Andrej Karpathy also participated as angel investors, according to TechCrunch. That financing backed a broader effort to commercialize diffusion language models; Mercury Voice puts the approach against a specific operational hurdle for voice products: a model must respond quickly while still following instructions and using tools.

For customers, the practical test will be performance on their own prompts and full voice stacks. Inception’s published median and p95 figures describe model response timing on its test set, while a deployed phone agent must also handle transcription, speech generation and network delays. The launch makes Mercury Voice available to enterprise buyers, but the company’s reported benchmark advantage is the starting point for that evaluation, not a guarantee of end-to-end response time.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @inception 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/inception-launches-m…] indexed:0 read:4min 2026-09-29 · —