Inception's Mercury 2.5: Diffusion LLM Hits 1,100 Tokens/sec, Claims Frontier-Level Quality
Inception Labs releases Mercury 2.5, a diffusion-based LLM boasting 40% quality gains over Mercury 2, 1,107 tokens/sec throughput, and pricing that undercuts mainstream models. The update targets latency-critical workloads like voice agents and coding subagents.

Inception Labs has shipped Mercury 2.5, the follow-up to its diffusion-based language model. The company claims a 40% jump in intelligence over Mercury 2 while keeping the same low-latency, low-cost serving profile. If the numbers hold, this is the largest diffusion LLM ever trained, and it's aimed squarely at production workloads where speed is the bottleneck.
What's new
Mercury 2.5 is built on the same diffusion architecture as its predecessor, but the training loop has been sharpened using production feedback. Inception says customer failures and eval improvements drove the gains, not just benchmark chasing. Key specs:
- Quality: 40% improvement over Mercury 2, comparable to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
- Speed: 1,107 tokens per second on widely available NVIDIA GPUs.
- Context: 260K tokens.
- Price: $0.20 per million input, $0.75 per million output. Launch promo: 80% off, dropping to $0.04 and $0.15 respectively.
- Capabilities: Tunable reasoning, parallel tool calls, and schema-aligned JSON output.
That price point is aggressive. Even at full price, it undercuts most frontier APIs, and the launch discount makes it a no-brainer for high-volume workloads.
Production proof points
Inception isn't just shipping a spec sheet. They've got real deployments to back it up.
Voice agents
OpenCall, which builds AI phone agents, reports median model response latency of ~170ms after switching to Mercury. Their P99 dropped from minutes to one second, and P50 from 0.4s to under 0.2s. For voice, latency is the product—that's the difference between a natural conversation and an awkward pause.
Coding subagents
Augment Code uses Mercury for context compaction and model routing. They cut latency by 82% (150s to 27s) and cost by 90% while maintaining quality. Tool-search summaries now return in under a second.
These are the workloads where latency compounds—dozens of model calls per user request, each one adding to the total. Mercury's speed turns that from a liability into a feature.
New tools: Mercury Voice and Router
Inception also previewed two new products:
- Mercury Voice: A dLLM optimized for voice agents with sub-170ms time-to-first-token.
- Mercury Router: Uses a dLLM to understand prompts and route them to the best model (open or closed) based on quality, speed, and cost tradeoffs.
Both are in preview and signal Inception's push beyond raw model sales into the inference orchestration layer.
Availability
Mercury 2.5 is live via the Inception API, Baseten, and OpenRouter. Enterprise customers get dedicated capacity, autoscaling, compliance controls, and configurable data retention.
The diffusion approach is still unconventional—most LLMs are autoregressive—but Inception is betting that speed wins. With numbers like these, it's hard to argue.
In voice, latency isn’t an infrastructure detail. It is the pause a caller hears.
| Feature | Mercury 2 | Mercury 2.5 |
|---|---|---|
| Intelligence | Baseline | +40% |
| Speed (tokens/sec) | Not specified | 1,107 |
| Context window | Not specified | 260K |
| Price per M input | Not specified | $0.20 ($0.04 launch) |
| Price per M output | Not specified | $0.75 ($0.15 launch) |
Discussion
0 Comments
Be the first to start the discussion.