Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSFebruary 9, 2026

PRESAGE Raises $1.4M to Predict Cloud Chaos Before It Happens

PRESAGE Raises $1.4M to Predict Cloud Chaos Before It Happens
Cloud InfrastructureAi Agents+3
Fintech iconFintechFebruary 9, 2026

Advance Raises $8.55M Seed to Modernize Insurance Payments

Advance Raises $8.55M Seed to Modernize Insurance Payments
InsurtechFintech+3

Founders Mentioned

Andrew Feldman

Cerebras Systems

saas icon
SaaS

Jensen Huang

Nvidia

saas icon
SaaS

Simon Willison

N/A

saas icon
SaaS

Andrew Feldman

Cerebras Systems

saas icon
SaaS

Jensen Huang

Nvidia

saas icon
SaaS

Simon Willison

N/A

saas icon
SaaS
SaaS iconSaaS
February 9, 2026
Semiconductor TechArtificial IntelligenceEnterprise AiModel Optimization

Inside Cerebras: The Wafer-Scale Chip Beating NVIDIA on Speed

Cerebras Systems' WSE-3 delivers 2,500+ tokens/second on frontier AI models—more than 2x faster than NVIDIA's Blackwell—as inference becomes AI's next battleground.

Inside Cerebras: The Wafer-Scale Chip Beating NVIDIA on Speed

When Meta released its Llama 4 Maverick 400B model in late May, the usual suspects lined up to run it through their paces. Independent benchmarking firm Artificial Analysis put the new model on servers from all the major players. Cerebras Systems, a Silicon Valley startup few mainstream investors had heard of three years ago, clocked 2,522 tokens per second per user.

NVIDIA's much-anticipated Blackwell DGX B200—the system the entire industry had been waiting for—delivered 1,038 tokens per second on the same workload. A respectable figure, certainly. But less than half Cerebras' output.

For those keeping score in AI infrastructure, margins like that don't happen by accident. They signal something fundamental shifting beneath the surface.

What's unfolding is a quiet but consequential realignment in how artificial intelligence gets built and deployed. Inference, the unglamorous work of actually running AI models for users, is becoming the industry's next battleground. And Cerebras—a company that bet its entire existence on a single, wildly unconventional chip design—finds itself leading a race that suddenly matters more than almost anyone predicted two years ago.

When the Math Changes, Everything Changes

The numbers tell the story clearly enough. AI spending is barreling toward $2.5 trillion by 2026, according to Gartner's latest projections. But here's what's different: more than 55 percent of AI-optimized infrastructure spending this year will go toward inference, not the model training that dominated budgets in 2023 and 2024. AI-optimized infrastructure-as-a-service alone should hit $37.5 billion this year, Gartner says, with inference claiming the majority.

Why the shift? Simple adoption curves meeting more demanding technology. GenAI use among U.S. adults jumped roughly 10 percentage points over 12 months, reaching 54.6 percent by August 2025, according to Federal Reserve Bank of St. Louis survey data. McKinsey's State of AI 2025 report found nearly universal experimentation among enterprises, with 62 percent actively testing agentic AI workflows—those multi-step reasoning systems that require AI to think, plan, and act rather than just respond.

Those agents don't just consume more compute per task. They need it faster. NVIDIA CEO Jensen Huang said publicly that reasoning models like DeepSeek R1 can demand up to 100 times more compute than straightforward chatbots. The economics suddenly look very different when your AI has to "think" through dozens of reasoning steps instead of firing back a single answer.

Cerebras launched its inference service in August 2024 making claims that seemed, at the time, borderline fantastical: 1,800 tokens per second on Llama 3.1 8B. On the massive 70B parameter model, 450 tokens per second, all at 16-bit precision. Industry observers were skeptical. By November, the company posted 969 output tokens per second on Llama 3.1-405B with time-to-first-token under 240 milliseconds. Artificial Analysis ran the benchmarks independently. The numbers held up. They were real, and they were substantially faster than anything GPU cloud providers could match at scale.

A Chip the Size of a Dinner Plate

Digital illustration for article section "A Chip the Size of a Dinner Plate" in "Inside Cerebras: The Wafer-Scale Chip Beating NVIDIA on Speed" - A conceptual visualization of a massive, singular silicon wafer chip dominating the composition to r...

The performance advantage traces back to an architectural decision that most chip designers would consider reckless. Where NVIDIA, AMD, and others build discrete GPUs and then link them together with high-speed interconnects—an approach that's dominated computing for decades—Cerebras took a different path entirely. They built a single chip that spans an entire 300-millimeter silicon wafer.

Think about that for a moment. Most chips occupy a few hundred square millimeters of a wafer. Cerebras uses the whole thing.

The WSE-3, fabricated on TSMC's 5-nanometer process, measures roughly 46,225 square millimeters. It packs 4 trillion transistors, 900,000 AI cores, and 44 gigabytes of on-chip SRAM—more memory than most laptops carried a decade ago. Memory bandwidth inside the chip hits 21 petabytes per second. The on-wafer interconnect fabric delivers over 200 petabits per second.

Those numbers are almost absurdly large. But they produce a concrete result: models run with minimal data movement. Standard GPUs rely on external high-bandwidth memory, which means weights and activations have to cross the chip boundary constantly. Each crossing costs time. Cerebras keeps everything on-wafer. For the kind of sequential inference computations that dominate today's frontier models, that difference matters enormously.

A single CS-3 system can hold models up to 20 billion parameters entirely on-chip. Larger models—Llama 3.3 70B, for instance—split across as few as four CS-3 units, still running at full 16-bit precision without compromise.

Speed translates into user experience in ways that commercial deployments care about deeply. Classic human-computer interaction research, the kind that's informed interface design for decades, sets 0.1 seconds as the threshold for "instantaneous" response. One second maintains flow. Anything beyond 10 seconds breaks attention. Cerebras consistently delivers time-to-first-token under 0.3 seconds on 70-billion-parameter models, which puts genuine real-time conversational AI within practical reach.

Groq and Google Vertex actually edge slightly faster on TTFT—0.22 and 0.18 seconds respectively on Llama 3.3 70B—but Cerebras leads decisively on throughput. Nearly 2,000 tokens per second versus Groq's roughly 340 tokens per second on the same model. When you're running agentic workflows, those multi-step reasoning loops that might require 20 to 100 times more tokens than a simple chat completion, throughput compounds. High tokens per second plus low latency together mitigate the end-to-end delays that otherwise make AI agents feel sluggish and frustrating to use.

The Customers Who Matter

Perplexity picked Cerebras to power its Sonar search product. The system runs Llama 3.3 70B at 1,200 tokens per second. Perplexity's public statements emphasized the "blazing fast" experience, and in practice the speed shows up tangibly—Sonar churns through multi-source summarization tasks in near real-time. Mistral's Le Chat assistant, running Mistral Large 2 at 123 billion parameters, delivers more than 1,100 tokens per second on Cerebras infrastructure. Developer Simon Willison flagged the Cerebras hosting in a February blog post, noting the responsiveness felt noticeably better than comparable GPU-backed services.

Beyond consumer-facing applications, Cerebras is picking up the kind of enterprise and scientific workloads that need throughput at serious scale. Mayo Clinic announced a genomic foundation model in January 2025, trained on Cerebras hardware. The model tackles applications in rheumatoid arthritis drug response, cancer predisposition, and cardiovascular phenotypes. Training benefited from Cerebras' large-scale data parallelism capabilities, but inference deployment—where the model actually assists clinical decisions—is where speed becomes critical.

The most significant commercial validation arrived in January 2026. The Financial Times reported that OpenAI signed a multiyear, $10 billion agreement with Cerebras for 750 megawatts of inference capacity running through 2028. The deal positions Cerebras as both a hedge against persistent GPU shortages and a way for OpenAI to diversify away from dependence on NVIDIA. It also signals confidence that Cerebras can scale beyond its current six announced datacenters, which collectively should deliver more than 40 million Llama-70B tokens per second at full capacity.

Pricing reveals the underlying economics. Cerebras launched at $0.10 per million input tokens for Llama 3.1 8B and $0.60 per million for Llama 3.1 70B, claiming 100-times-better price-performance than GPU clouds. Those rates compete well with hyperscale providers on a per-token basis, but the real value proposition is total cost of ownership when you factor in lower latency and higher user concurrency. A system that can serve twice as many simultaneous users effectively costs half as much per user, even if the raw hardware price stays constant.

The Bet Gets Bigger

Digital illustration for article section "The Bet Gets Bigger" in "Inside Cerebras: The Wafer-Scale Chip Beating NVIDIA on Speed" - A conceptual visualization of massive financial scaling and high-stakes corporate valuation, depicte...

Cerebras raised $1.1 billion at an $8.1 billion valuation in September 2025. Five months later, in February 2026, another $1 billion Series H round valued the company at roughly $23 billion post-money. Investors include Tiger Global, Benchmark, Fidelity, AMD, Coatue—the kind of firms that typically don't make mistakes about infrastructure trends. They're betting inference will justify a second massive wave of AI infrastructure buildout, perhaps as large as the training-focused boom that already happened.

CEO Andrew Feldman has argued publicly that inference will "dwarf training" in economic terms and described the current moment as AI's "broadband moment," when falling latency unlocks use cases nobody could previously imagine. There's precedent for that kind of transformation. The jump from dial-up to broadband didn't just make existing websites load faster. It enabled YouTube, Netflix, real-time gaming, video conferencing—entirely new categories of application.

The narrative is plausible. But Cerebras faces structural challenges that capital alone won't solve.

NVIDIA's ecosystem lock-in remains formidable, particularly for training workloads where CUDA libraries and tooling have accumulated two decades of momentum and optimization. NVIDIA's roadmap includes Rubin, scheduled for the second half of 2026, which promises up to five times greater inference performance and 10 times lower cost per token than Blackwell. Those are projections, of course. NVIDIA has a history of delivering on ambitious roadmaps, but also of slipping schedules. If those gains materialize on time, though, Cerebras will need to stay ahead on absolute speed while expanding capacity fast enough to meet surging demand.

Supply chain constraints loom large and getting larger. TSMC CEO C.C. Wei said publicly that advanced-node manufacturing capacity falls roughly three times short of AI demand. Wafer-scale chips are inherently yield-sensitive—a single defect can potentially ruin an entire wafer. Cerebras has demonstrated it can manufacture WSE-3 at commercial scale, but ramping to meet OpenAI's 750-megawatt commitment while serving other major customers will test execution capabilities in ways the company hasn't faced before.

Power and permitting have become equally critical bottlenecks. The International Energy Agency projects data center electricity consumption will nearly double to 945 terawatt-hours by 2030, with the U.S. and China accounting for 80 percent of that growth. Siting datacenters with reliable, affordable power is increasingly difficult. Cerebras' strategy of owning and operating facilities in places like Oklahoma City and Montreal gives it some control over that variable, but also concentrates risk. A single permitting delay or power grid issue could cascade across commitments.

Regulatory headwinds add another layer of uncertainty. U.S. export controls on advanced chips tightened in early 2025, then were partially revised in January 2026, with case-by-case licensing now required for H200 and MI325X-class hardware going to China. Flows to the Middle East, particularly the UAE, remain under close government scrutiny. Cerebras hasn't publicly disclosed customer concentration by geography. Any significant exposure to restricted markets could complicate growth trajectories.

The competitive landscape is getting crowded as well. Groq, d-Matrix, and others are targeting inference with specialized silicon architectures. Google's TPU v5e claims 2.5 times better performance per dollar than TPU v4 for inference workloads. AWS keeps expanding Inferentia and Trainium instance types. None of these alternatives match Cerebras on raw tokens per second for frontier models today. But the race is far from over, and history is littered with companies that led a market briefly before getting outflanked.

The Falling Cost of Intelligence

Digital illustration for article section "The Falling Cost of Intelligence" in "Inside Cerebras: The Wafer-Scale Chip Beating NVIDIA on Speed" - Create a professional and conceptual illustration representing the rapid decline in artificial intel...

What's becoming clear is that inference economics are shifting rapidly, perhaps faster than training economics did. Sam Altman has said publicly that inference costs are falling roughly 10 times per year, driven by both hardware improvements and aggressive software optimization. Token prices will continue dropping. But total tokens consumed will surge as agents and reasoning models proliferate across consumer and enterprise applications. That dynamic favors platforms capable of delivering both speed and cost efficiency at genuine scale.

Cerebras is well-positioned for this particular moment—arguably better positioned than any competitor not named NVIDIA. The wafer-scale architecture delivers a measurable, reproducible speed advantage on the models enterprises and consumers care about most right now. The OpenAI deal provides both substantial capital and enormous credibility. And broader market tailwinds, from agentic AI adoption to enterprise deployment accelerating, all point toward dramatically higher demand for fast, scalable inference infrastructure.

Whether Cerebras can sustain that lead as NVIDIA iterates its roadmap, supply chains tighten further, and capital requirements balloon into tens of billions will define the company's next chapter. The technical challenges are formidable. The execution risks are real.

But for now, the numbers speak clearly enough. 2,500 tokens per second is fast enough that users can feel the difference. In consumer technology, feeling matters as much as any benchmark. Perhaps more.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • PRESAGE Raises $1.4M to Predict Cloud Chaos Before It Happens
  • Advance Raises $8.55M Seed to Modernize Insurance Payments
  • Omega-3 Research Sparks $2B Market Shift Toward Brain Health
  • Claude's Fast Mode: 2.5x Speed at 6x the Price for Developers
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.