Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
Healthtech & Biotech iconHealthtech & BiotechSeptember 9, 2026

Cloverleaf Bio raises $33M to develop RNA cancer therapies

Cloverleaf Bio raises $33M to develop RNA cancer therapies
Rna TherapeuticsOncology+3
Healthtech & Biotech iconHealthtech & BiotechSeptember 9, 2026

BrainChild Bio raises $116M to advance DIPG brain cancer therapy

BrainChild Bio raises $116M to advance DIPG brain cancer therapy
BiotechOncology+3

Founders Mentioned

Stefano Ermon

Inception

saas icon
SaaS

Aditya Grover

Inception Labs

saas icon
SaaS

Volodymyr Kuleshov

Inception Labs

saas icon
SaaS

Oliver Silverstein

OpenCall

saas icon
SaaS

Stefano Ermon

Inception

saas icon
SaaS

Aditya Grover

Inception Labs

saas icon
SaaS

Volodymyr Kuleshov

Inception Labs

saas icon
SaaS
SaaS iconSaaS
September 9, 2026
Large Language ModelsDiffusion ModelsAi InfrastructureAi Price War

Inception launches Mercury 2.5 at 1,107 tokens/second

Diffusion-based language model promises 5-10x faster inference than autoregressive competitors, with launch pricing at $0.04 per million tokens for AI-powered apps.

Inception launches Mercury 2.5 at 1,107 tokens/second

Inception unveiled Mercury 2.5 on September 8, claiming its diffusion-based language model delivers 1,107 output tokens per second on standard NVIDIA GPUs and runs five to ten times faster than leading autoregressive rivals. Launch pricing came in at $0.04 per million input tokens and $0.15 per million output tokens—a limited-time promotional rate down from standard list prices of $0.20 input and $0.75 output—as the Palo Alto company looks to convert developers frustrated by inference costs and latency.

The pitch arrives at a moment when cost-per-token has displaced raw capability as the defining constraint for companies deploying conversational agents and coding assistants at scale. Whether Mercury 2.5 can deliver on its speed claims under production load remains an open question. BenchLM noted the day after launch that independent runtime measurements had not yet surfaced, though earlier Mercury releases consistently ranked near the top of speed leaderboards tracked by Artificial Analysis.

Language model inference has become a brutal efficiency contest. Gartner projected AI-optimized infrastructure spending would grow by 96 percent in 2026, driven overwhelmingly by inference workloads. McKinsey wrote in July that enterprises scaling agentic loops now measure everything in energy-per-token and cost-per-token, metrics that favor any architecture willing to break from the autoregressive default.

Autoregressive models—the family tree that includes GPT, Claude, and Gemini—generate text one token at a time, marching left to right in an inherently serial process. Diffusion language models borrow a trick from image generation: they produce and refine multiple tokens in parallel through iterative denoising. The approach trades serial bottlenecks for computational parallelism. Google DeepMind shipped experimental diffusion models in 2025 and followed with DiffusionGemma, a 26-billion-parameter mixture-of-experts release, in mid-2026. Inception's Mercury 2 debuted in February of that year and was noted near the top of speed leaderboards through mid-2026, according to aggregators including LMSpeed and Easy Benchmarks.

Mercury 2.5 ships with a 260,000-token context window, up to 65,536 output tokens, tool calling, structured JSON mode, and what the company describes as tunable reasoning. The model routes through Inception's own API—which includes 100 million free tokens for new users—plus third-party gateways like Baseten, OpenRouter, Vercel AI Gateway, and OminiGate.

The Academic Roots

Three professors co-founded Inception: Stefano Ermon at Stanford, Aditya Grover at UCLA, and Volodymyr Kuleshov at Cornell. The trio raised $50 million in seed funding led by Menlo Ventures in November 2025, with participation from Mayfield, Innovation Endeavors, Microsoft's M12, Snowflake Ventures, Databricks Ventures, NVIDIA's NVentures, and individual investors including Andrew Ng and Andrej Karpathy, according to TechCrunch.

Tim Tully at Menlo Ventures wrote in a mid-2026 perspective piece that the firm believes "all LLMs will eventually be built on diffusion" because autoregressive decoding's sequential nature cannot meet enterprise latency and throughput requirements at scale. He described Mercury 2 as "the world's first—and fastest—large-scale reasoning model built on diffusion."

Shruti Koparkar, senior manager of product in NVIDIA's accelerated computing group, said in Inception's launch materials that Mercury 2.5's combination of intelligence, sustained speed, and low cost "reflects how quickly new architectures can mature into production-ready systems on the NVIDIA platform."

NVIDIA's own August blog posts positioned its Blackwell GPU family around tokens per second per watt, claiming throughput and cost-per-token advantages over the prior Hopper generation on inference tasks. The vLLM project, a widely used open-source inference server, added native support for DiffusionGemma in June, incorporating Triton Attention and FlashAttention-4 backends. That move signals mainstream inference tooling is beginning to accommodate diffusion-specific optimizations, though adoption remains nascent compared to autoregressive stacks.

Regulatory shifts are also reshaping the landscape. The European Union's AI Act obligations for general-purpose AI model providers—documentation requirements, authorized EU representatives, copyright policies, training-data summaries—moved from voluntary codes of practice to enforceable fines after August 2, 2026, according to European Commission guidance. Compliance costs may tilt smaller model providers toward lower-overhead architectures or force consolidation among regional players.

Early Adopters Report Gains

Digital illustration for article section "Early Adopters Report Gains" in "Inception launches Mercury 2.5 at 1,107 tokens/second" - A conceptual and minimalist surreal digital collage representing a dramatic decrease in voice-platfo...

OpenCall, a voice-agent platform, reported its P99 response time dropped from several minutes to one second after switching to Mercury, while P50 latency fell from 0.4 seconds to under 0.2 seconds, according to co-founder and CEO Oliver Silverstein. "That's significantly faster than any other provider we've seen, and that's including reasoning," Silverstein said in the launch post.

Augment Code, which builds coding assistants and subagents, moved a compaction workload to Mercury and saw latency drop 82 percent—from roughly 150 seconds to 27 seconds—while cutting cost by 90 percent. Inception reported that several large search-infrastructure companies now run Mercury in production but declined to name them.

Microsoft's NLWeb project, described by the company in May 2025 as "an open project designed to simplify the creation of natural language interfaces for websites," lists Inception as its founding LLM partner, according to Inception's product blog. An August 11 post titled "Mercury 2 for Search" published reproducible pipeline benchmarks comparing Mercury 2 to other models on search datasets including WideSearch, FRAMES, and DeepSearchQA.

Competitors occupy different price bands, which complicates direct comparisons. Launch-promo Mercury 2.5 undercuts most speed-tier alternatives on input costs. xAI's Grok release from August charged $2.00 per million input tokens and $6.00 per million output tokens. Google's Gemini Flash-Lite pricing tracked around $0.30 input and $2.50 output through mid-2026, according to multiple aggregators. OpenAI's smaller GPT models generally charged $0.20 input and $1.25 output for their nano-tier offerings.

Liquid AI's LFM model claimed 15,000 tokens per second at high concurrency on a single H100 in August, but that figure reflects a 2.6-billion-parameter model optimized for edge deployment rather than cloud reasoning workloads. IBM's Granite models tracked between 100 and 400 tokens per second depending on provider and task, according to independent snapshots from aggregators. Mercury's 1,107-token claim would place it well above autoregressive models in the same capability bracket, assuming the numbers hold under independent review.

What Comes Next

Digital illustration for article section "What Comes Next" in "Inception launches Mercury 2.5 at 1,107 tokens/second" - A clean, minimal, and highly conceptual surreal digital collage representing a leap forward in capab...

Inception CEO Stefano Ermon wrote in the launch post that the company has already begun training its next model, describing it as "a leap in capability without giving up diffusion's speed and token-efficiency." Release is expected in the coming months, though no specifics were shared.

The broader research pipeline for diffusion-in-text remains active. Cornell researchers published FlashDLM in 2026 on KV caching and guided diffusion. Predict-then-Diffuse appeared in May on compute-budgeted inference. Speculative Correction followed in July. Google's June release of DiffusionGemma with vLLM support lowered switching costs for developers evaluating diffusion alternatives, though most production systems remain anchored to autoregressive defaults.

Founders building voice agents, coding subagents, or search pipelines should watch for independent speed benchmarks from Artificial Analysis or similar third parties, since day-one measurements have not yet materialized. The launch promo's expiration timeline remains unannounced. Developers evaluating cost models should confirm current pricing with their chosen provider, as promotional rates can shift without warning. Azure AI Foundry, Baseten, OpenRouter, Vercel AI Gateway, and OminiGate all listed Mercury 2.5 between September 8 and 9, offering broad routing options for A/B tests against incumbent models.

If independent benchmarks confirm Inception's speed claims and quality holds under production load, Mercury 2.5 could accelerate adoption of diffusion architectures in latency-sensitive enterprise workloads. Voice interfaces demand sub-200-millisecond first-token times. Multi-turn coding agents burn tokens on context compaction. Search systems stack dozens of LLM calls per query. The economics push in that direction: at $0.04 input and $0.15 output, a 10,000-token search pipeline costs $1.90 on Mercury versus $5.30 on smaller OpenAI models or $27.50 on Google's Flash-Lite tier, assuming list prices and no volume discounts.

Whether diffusion becomes the new default or remains a niche optimization for speed-obsessed applications will depend on how well these models perform when the benchmarks move beyond synthetic tests and into the messy reality of production traffic. For now, Inception has given developers a reason to revisit assumptions about how language models get built.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • Cloverleaf Bio raises $33M to develop RNA cancer therapies
  • BrainChild Bio raises $116M to advance DIPG brain cancer therapy
  • ChatGPT referrals to B2B sites jumped 303% in one year
  • Limetax raises €36M to build AI-powered accounting firm group
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.