Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechFebruary 26, 2026

Electric Kilns Race to Decarbonize Cement and Lime Production

Electric Kilns Race to Decarbonize Cement and Lime Production
Clean TechCarbon Management+2
Climate / Social Tech iconClimate / Social TechFebruary 26, 2026

CuspAI Raises $100M to Cut Materials Discovery From Years to Months

CuspAI Raises $100M to Cut Materials Discovery From Years to Months
Materials ScienceArtificial Intelligence+3

Founders Mentioned

Aditya Grover

Inception Labs

saas icon
SaaS

Volodymyr Kuleshov

Inception Labs

saas icon
SaaS

Aditya Grover

Inception Labs

saas icon
SaaS

Volodymyr Kuleshov

Inception Labs

saas icon
SaaS
SaaS iconSaaS
February 26, 2026
Diffusion ModelsEnterprise AiModel OptimizationStartup FundingB2b Saas

Inception Labs' Mercury 2: Diffusion LLM Hits 1,000 Tokens/Sec

Backed by $50M from Menlo and Microsoft, the startup's diffusion-based reasoning model challenges autoregressive orthodoxy with 5x speed gains and enterprise cloud availability.

Inception Labs' Mercury 2: Diffusion LLM Hits 1,000 Tokens/Sec

The first thing you notice about Mercury 2 isn't what it does—it's what it doesn't do.

While OpenAI, Anthropic, and Google chase reasoning supremacy with models that pause, deliberate, and burn tokens on internal monologue, a Stanford-born startup called Inception Labs is placing a different wager. The company launched Mercury 2 on February 24, 2026, with a provocative claim: that the real bottleneck in artificial intelligence isn't chip architecture or cluster size, but something more fundamental. It's the way nearly every large language model in production actually generates text.

Mercury 2 doesn't predict words left-to-right like GPT-4 or Claude. It doesn't build sentences sequentially, token by token, the way autoregressive models have since the transformer architecture emerged. Instead, it uses diffusion—the probabilistic framework that powers image generators like Stable Diffusion—to refine entire sequences in parallel. Think of it less as writing and more as sculpting: iteratively shaping a complete response rather than laying down one word after another.

The measured result? 1,009 tokens per second on NVIDIA's Blackwell GPUs. That's not just fast. It's fast enough to change what kinds of AI applications become feasible.

The technical gamble is backed by $50 million from Menlo Ventures and Microsoft's M12 fund, announced last November. Co-founders Stefano Ermon (Stanford), Aditya Grover (UCLA), and Volodymyr Kuleshov (Cornell Tech) aren't newcomers to probabilistic AI. Their pitch, refined over months of investor meetings, boils down to a single insight: autoregressive decoding works fine for chatbots. But it breaks down in the emerging ecosystem of AI agents, real-time voice interfaces, and multi-step retrieval pipelines where latency doesn't just add up—it multiplies.

"Every model call compounds," Ermon explained in an early presentation deck obtained by sources familiar with the fundraising process. For a debugging agent running 20 iterations or a voice assistant that needs sub-500ms response times to feel natural, sequential generation becomes the bottleneck.

Whether the market agrees remains to be seen.

A Market Splitting Apart

The LLM inference landscape is fracturing. On one side: reasoning depth. OpenAI's o-series models—o3, o3-mini, o4-mini—push test-time compute to extremes, burning tokens on deliberation to solve hard math and coding problems. DeepSeek R1, released with open weights earlier this year, proved reinforcement learning could deliver similar reasoning at a fraction of the cost. Sub-$1 per million tokens in some deployments, versus OpenAI's premium pricing.

But reasoning models share a liability. They're slow. The mechanism that makes them smart—extended "thinking" through chain-of-thought or search—adds seconds per query. For a customer support chatbot, tolerable. For an AI agent iterating through debugging cycles? A dealbreaker.

On the other side: raw speed. Hardware vendors have their own answers. Cerebras's wafer-scale chips clock over 2,500 tokens per second on Llama 4 Maverick, according to Artificial Analysis benchmarks. Groq's LPU architecture targets sub-200ms time-to-first-token. NVIDIA's Blackwell platform broke 1,000 tokens per second per user in early tests. Microsoft, betting on scale, deployed what it calls the world's first supercomputer-grade GB300 NVL72 cluster—4,608 GPUs linked into a single accelerator delivering 1.44 petaflops of inference capacity.

Mercury 2 enters this arms race with a counterintuitive thesis: that parallelism should happen inside the model, not just in the datacenter.

Inception's diffusion approach generates tokens through iterative refinement rather than sequential prediction. Different portions of a response can emerge simultaneously. The company claims this yields a 5x speed advantage over "leading speed-optimized LLMs" while maintaining competitive quality on reasoning benchmarks like AIME 2025 (91.1) and GPQA (73.6).

The skeptics will want to see those numbers hold up under adversarial inputs and edge cases not covered by standardized tests. But Inception's distribution strategy—Azure AI Foundry, AWS Bedrock and SageMaker JumpStart, OpenRouter—makes experimentation low-friction. The OpenAI-compatible API means developers can swap Mercury in without rewriting pipelines.

Three Converging Pressures

Three forces are conspiring to make inference speed a strategic problem, not just an operational nuisance.

First, agentic AI is moving from research demos to production deployments. Gartner projects 40% of enterprise applications will include task-specific AI agents by 2026, up from less than 5% in 2024. These systems don't make single API calls. They orchestrate dozens or hundreds of LLM invocations—planning, acting, observing, refining. A debugging agent, for instance, might query the model for code analysis, error diagnosis, and a proposed fix in each iteration. If each step takes three seconds, a 10-step loop burns half a minute. Multiply that across concurrent users, and compounding latency becomes the dominant cost driver.

Second, voice interfaces demand sub-500ms end-to-end latency to feel conversational. That budget gets carved into speech-to-text, LLM inference, and text-to-speech segments. Industry sources peg the LLM inference slice at 150–250ms to leave room for the rest of the pipeline. Autoregressive models, even fast ones, struggle to hit that threshold consistently under load.

Inception's customer roster reflects this pressure. Wispr Flow uses Mercury for real-time transcript cleanup. OpenCall built voice agents around it. Happyverse AI leverages it for voice avatars. These aren't research projects—they're production systems where perceptible lag breaks the user experience.

Third, the economics are shifting. Gartner forecasts $644 billion in global generative AI spending for 2025, up 76% year-over-year. IDC predicts AI spending will hit $632 billion by 2028, with GenAI capturing $202 billion. McKinsey pegs GenAI's total economic potential at $2.6 to $4.4 trillion annually across functions.

As deployment scales, the marginal cost of inference—tokens per dollar, compute per query—becomes a competitive moat. Mercury 2's pricing ($0.25 per million input tokens, $0.75 per million output) undercuts many reasoning models while promising faster throughput. The press release claims "dramatically lower inference cost" compared to competitors, though direct comparisons are murky given differences in model size and capability.

The regulatory environment adds another wrinkle. The EU AI Act entered force in August 2024, with obligations for general-purpose AI models kicking in this past August and broader high-risk system rules coming next August. In the U.S., Executive Order 14110 was rescinded in January 2025, but NIST's AI Risk Management Framework and Generative AI Profile—published last July—remain the de facto standard for enterprise buyers. Copyright litigation, exemplified by the New York Times v. OpenAI case allowed to proceed in March 2025, looms over training practices.

Model providers differentiating on speed and deployment flexibility, rather than just scale, may find easier paths through this legal thicket. Or so the theory goes.

Production Deployments

Digital illustration for article section "Production Deployments" in "Inception Labs' Mercury 2: Diffusion LLM Hits 1,000 Tokens/Sec" - A hyper-realistic 3D visualization illustrating the speed and efficiency of production deployments i...

Inception's go-to-market centers on verticals where latency stacks up: enterprise search, coding tools, real-time interaction.

SearchBlox, a Georgia-based search platform, published a case study in January touting "sub-second RAG" with Mercury. The company's CTO cited lower inference costs and deployment options across Azure and AWS as deciding factors. RAG workloads are particularly latency-sensitive. Each multi-hop query triggers multiple retrieval-and-generation cycles. Research papers from early 2025 highlight TTFT (time-to-first-token) and throughput improvements as critical to making iterative RAG strategies viable at scale.

Viant, which runs agentic customer engagement loops, called out Mercury 2's speed for handling what it describes as "agentic loops"—systems that iterate on plans and actions in real time. Each LLM call adds cumulative delay.

Inception's earlier Mercury Coder models showed traction in developer tools. Independent evaluations by Artificial Analysis ranked Mercury Coder as a speed leader last year. User preference data from Copilot Arena placed "Mercury Coder Mini" near the top while maintaining fastest-in-class throughput. Microsoft's NLWeb integration, announced at Build 2025, demonstrated Mercury powering natural-language interfaces for semi-structured websites. Partners included TripAdvisor, Shopify, and Snowflake—enterprises with high-volume, low-latency query patterns.

The technical foundation traces back to academic work by Inception's founders. Ermon's Stanford lab contributed to discrete diffusion frameworks (Multinomial Diffusion, D3PM). Collaborators published on mask-based refinement and non-autoregressive decoding dating to 2019's Mask-Predict. Inception's own research output includes papers on remasking (inference-time scaling for diffusion models), guidance mechanisms for discrete diffusion, and block diffusion—a hybrid approach interpolating between autoregressive and diffusion paradigms.

A June 2025 technical report detailed Mercury's transformer-parameterized diffusion architecture, reporting 1,100 tokens per second on NVIDIA H100s for the Mini variant and 700-plus for Small.

Recent academic work validates the reasoning potential of diffusion LMs. "Diffusion-of-Thought" (February 2024) applied the framework to multi-step reasoning tasks. "LaDiR" (October 2025) introduced latent diffusion for text reasoning. "DiffThinker" (December 2025) extended diffusion reasoning to multimodal settings. "d1" (April-June 2025) demonstrated that reinforcement learning could scale diffusion LMs for reasoning—an approach Inception appears to be commercializing with Mercury 2's "tunable reasoning" feature.

The Unresolved Question

The diffusion-versus-autoregressive debate won't be settled by one launch or one benchmark.

Autoregressive models benefit from nearly a decade of optimization across training, serving, and hardware stacks. Diffusion LMs are newer, with fewer practitioners and less battle-tested infrastructure. Skeptics will rightly ask whether Mercury 2's speed claims hold under diverse workloads, adversarial inputs, or edge-case reasoning tasks not captured by AIME and GPQA benchmarks.

But there are mitigating factors. NVIDIA's Shruti Koparkar, quoted in Inception's launch materials, positioned the >1,000 tokens per second result as ecosystem validation, not just a startup claim. Investor enthusiasm suggests institutional belief in the approach. Menlo Ventures and Mayfield, alongside Microsoft, Snowflake Ventures, Databricks Ventures, and angels Andrew Ng and Andrej Karpathy, backed Inception's $50 million round last November. Press release quotes from investors assert that a "diffusion-first approach can reset expectations for scalable reasoning."

Whether that proves true depends on production deployments, independent evaluations, and—crucially—customer retention over the next 12 months.

The real test will be whether diffusion models can match or exceed autoregressive quality as reasoning complexity increases. OpenAI's o3 and Anthropic's "hybrid reasoning" features in Claude represent entrenched competition with massive resources. DeepSeek R1's open weights and sub-$1 pricing, widely cited in benchmarks published in January, set a high bar for cost-efficiency. If Mercury 2 can deliver comparable reasoning at 5x the speed and lower cost, it reshapes the build-versus-buy calculus for enterprises deploying agentic systems.

The broader trajectory points toward tighter latency budgets and more architectural diversity. As voice agents and multi-step workflows proliferate, the sub-500ms threshold will frame not just model selection but application design. Research on test-time scaling—methods like "First-Finish Search" and value-guided search—aims to make reasoning more efficient without sacrificing quality. Hardware will continue its arms race: NVIDIA's Blackwell and Rubin roadmaps, Cerebras's wafer-scale chips, Groq's LPUs all target lower latency and higher throughput.

Mercury 2's claims position algorithmic parallelism as orthogonal and composable with faster hardware. A diffusion model on Blackwell, in theory, should outpace an autoregressive model on the same chip. In theory.

Regulatory dynamics may accelerate the shift. EU AI Act obligations for general-purpose models took effect last August, with broader high-risk system enforcement coming this August. Model providers that can demonstrate lower compute footprints, faster inference, and simpler deployment may find easier compliance paths and lower scrutiny.

A Stake in the Ground

Digital illustration for article section "A Stake in the Ground" in "Inception Labs' Mercury 2: Diffusion LLM Hits 1,000 Tokens/Sec" - A dramatic, conceptual visualization of a high-tech metallic stake driven firmly into a digital foun...

What's certain is that 1,009 tokens per second is no longer just a benchmark. It's a stake in the ground.

Whether diffusion LMs become a serious alternative to autoregressive reasoning models or remain a niche tool for latency-critical workloads will depend on the next year of deployments, evaluations, and customer retention. Inception has placed its bet—that parallel refinement can dethrone sequential prediction, that speed is a feature worth trading hardware compatibility for, that the inference bottleneck is algorithmic rather than just a matter of throwing more GPUs at the problem.

The market, as always, will have the final word. But for now, at least, the autoregressive consensus has company.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Electric Kilns Race to Decarbonize Cement and Lime Production
  • CuspAI Raises $100M to Cut Materials Discovery From Years to Months
  • Zelostech Raises $400M Series B for Autonomous Logistics Fleet
  • Gushwork Raises $9M to Automate AI Search Optimization
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.