Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
eCommerce iconeCommerceJuly 4, 2026

Dunzo's Kabeer Biswas Raises $12M for AI Concierge Startup M

Dunzo's Kabeer Biswas Raises $12M for AI Concierge Startup M
Ai AgentsConversational Ai+3
SaaS iconSaaSJuly 4, 2026

PhysicsX Hits $2.4B Valuation Building AI for Chip Manufacturing

PhysicsX Hits $2.4B Valuation Building AI for Chip Manufacturing
AiSemiconductor Tech+3
SaaS iconSaaS
July 4, 2026
YcAi HardwareAi InfrastructureSemiconductor TechDeeptech

Inside Baud Labs' Plan to Eliminate Multipliers from AI Training

A three-person YC startup claims its novel chip architecture can train AI models orders of magnitude more efficiently than NVIDIA—by removing multiplications entirely.

Inside Baud Labs' Plan to Eliminate Multipliers from AI Training

Three engineers walk into Y Combinator with a pitch that sounds, at first blush, like technical fantasy: we can train large language models orders of magnitude more efficiently than NVIDIA. The trick? Rip out the multipliers.

Baud Labs—operating under the less memorable legal name Cerelyze, Inc.—emerged from a recent YC batch with a thesis that feels almost heretical in an industry built entirely on matrix multiplication. Swap those power-hungry multipliers for adders, the founders argue, and you can pack more compute and memory onto the same sliver of silicon. Faster, cooler, cheaper.

It's the kind of audacious claim that could mark the beginning of something or simply become another cautionary tale in the graveyard of would-be NVIDIA killers. But the fact that they're trying at all reveals something about where the constraints are starting to bind.

When the Infrastructure Starts Groaning

NVIDIA's Blackwell platform represents the current high-water mark. The GB200 NVL72 rack—72 GPUs lashed together via NVLink at 900 GB/s—claims 4× faster training than the previous H100 generation and up to 30× better inference performance, according to NVIDIA's benchmarks. Recent MLPerf submissions scaled to 8,192 Blackwell GPUs, a flex that reinforced NVIDIA's chokehold on both performance and mindshare.

Google Cloud rolled out its eighth-generation TPUs in 2026, split between training (8t) and inference (8i) variants. AMD keeps shipping on its annual cadence—the MI325X arrived with 288GB of HBM3E, the MI350X is on deck—but market share remains stubbornly tilted toward NVIDIA. AWS brought Trainium2 to general availability last fall. Microsoft unveiled the Maia 200 inference accelerator in early 2026.

Yet beneath those headline numbers, the cracks are showing.

High-bandwidth memory—HBM, the stuff that sits stacked atop AI accelerators—has been effectively sold out, with supply constraints bleeding into procurement timelines and spawning the occasional class-action grumble around DRAM pricing. TSMC's CoWoS advanced packaging capacity, critical for stacking HBM onto compute dies, has been expanding but not fast enough. Industry estimates suggest capacity might reach somewhere between 120,000 and 140,000 wafers per month, narrowing the supply-demand gap but hardly eliminating it.

Then there's power. A Deloitte survey from last year found that 72% of respondents cited grid capacity as a serious or extreme obstacle for data center buildouts. FERC Order 1920, finalized in mid-2024, aimed to reform regional transmission planning, but reforms take years. Industry reports have documented local pushback and grid constraints slowing projects across multiple states.

It's a market that IDC projects will hit roughly $487 billion in AI infrastructure spending this year, climbing past $1 trillion by decade's end. That kind of money draws competition.

The Economics of Elimination

Training a single frontier model now costs tens to hundreds of millions of dollars in compute alone—perhaps more than the founders expected when they first scribbled napkin math. The Stanford AI Index's 2024 report estimated GPT-4's training at roughly $78 million and Gemini Ultra at around $191 million in 2023, figures that have only climbed as model sizes balloon. Meta reportedly trained Llama 4 on more than 100,000 NVIDIA H100 GPUs, according to statements from Mark Zuckerberg, a cluster scale that few organizations can replicate.

Multipliers consume disproportionate die area and power. A multiplier's area scales quadratically with bit width; an adder scales linearly. Baud's claim is that by eliminating multipliers and using what they describe as a losslessly compressed arithmetic representation, they can fit more compute and memory onto the same silicon footprint at the same process node and memory bandwidth.

The underlying idea isn't entirely novel. Academic work over the past several years has explored multiplication-free transformer training, AdderNet architectures, logarithmic number systems. What's different—if Baud can pull it off—is productizing it at scale, with a compiler that supports PyTorch nn.Module, torch.distributed, FSDP2, and tensor parallelism from day one. That software layer is critical. No one wants to rewrite their training loops for exotic hardware, no matter how fast it runs.

Baud's public demo runs a 50-million-parameter model trained on 5 million tokens at over 1,000 tokens per second on a single Xilinx U200 FPGA clocked at 125 MHz. It's a proof of concept, not a production deployment. The company has opened early access to its first cluster, supporting workloads up to 32 billion parameters for pretraining, fine-tuning, reinforcement learning post-training, and inference. The positioning is explicit: a "multiplier-free ASIC" experience, delivered via cloud.

Still. An FPGA demo is not an ASIC. And 32 billion parameters is well short of the frontier models that justify the "orders of magnitude" rhetoric.

The Rush for Custom Silicon

Digital illustration for article section "The Rush for Custom Silicon" in "Inside Baud Labs' Plan to Eliminate Multipliers from AI Training" - A conceptual, minimalist illustration representing the rush for custom silicon, featuring a single, ...

Baud isn't alone in attempting an end run around NVIDIA's ecosystem. The past year has seen a wave of startups and hyperscalers bet on custom silicon, each targeting a different slice of the problem.

Groq, focused on inference, reportedly raised $650 million to expand its inference cloud. The company announced a partnership with Meta for the official Llama API and positioned its LPU (Language Processing Unit) architecture within NVIDIA's Rubin platform at a recent GTC event. The pitch: deterministic, low-latency inference at a fraction of the cost.

Etched emerged from stealth mid-2026 with reports of $500 million to $800 million raised and over $1 billion in contracts. Their chip—an algorithm-specific transformer ASIC they call "Sohu," built on an advanced process node—targets 10× efficiency gains over GPUs on selective workloads. It's a narrow bet: transformers only, no general-purpose flexibility. If the architecture shifts, so does their moat.

d-Matrix entered full production in June 2026 with its Corsair inference platform, built on TSMC's N6 process. The design is SRAM-heavy, using LPDDR5 instead of HBM to sidestep CoWoS bottlenecks. It's an acknowledgment that the supply chain, as much as the architecture, shapes what's viable.

The hyperscalers are hedging their own bets. Meta announced its MTIA roadmap, targeting inference with a multi-year chip series through Broadcom. OpenAI's collaboration with Broadcom, announced late last year, reportedly targets 10 GW of custom accelerators; press coverage has named "Jalapeño" as OpenAI's first in-house chip. Anthropic signed a deal with Google Cloud for multi-gigawatt capacity and access to up to a million TPUs, according to late-2025 reports.

Broadcom CEO Hock Tan framed the trend bluntly in recent coverage: firms "need their own chips." It's a reflection of both technical differentiation and supply chain anxiety. Relying on a single vendor, even one as dominant as NVIDIA, introduces risk that boards and CFOs are no longer willing to tolerate.

The Team and the Gamble

Baud's founders bring relevant pedigrees, if not a long track record in silicon. CEO Sarang Zambare is a multi-time founder with four patents and over seven years in deep learning and AI hardware. He led ML for Peloton Guide from concept to what the company says was 100,000+ devices and was a founding ML engineer at Caper, later acquired by Instacart.

Chief Hardware Architect Eric Taylor has more than a decade of ASIC experience, with multiple tape-outs, major IP releases, and prior stints at NVIDIA, Freescale/NXP, Arteris IP, and Enfabrica.

The Enfabrica connection is worth noting. The networking startup raised $115 million in late 2024 and reportedly signed licensing and hiring deals with NVIDIA worth around $900 million. Enfabrica's ACF SuperNIC and Ethernet-attached memory products address the same scale-up and memory challenges that Baud's architecture aims to solve from a different angle—albeit without reinventing arithmetic.

But here's the rub: Baud's technical claims rest, for now, on unverified assertions. No silicon benchmarks have been disclosed. No power-performance-area data. No MLPerf submissions. The early access cluster supports models up to 32 billion parameters—well short of the frontier models that justify the "orders of magnitude" language the company uses. The FPGA demo is encouraging, sure, but FPGAs are not ASICs. Emulation speed and production speed are different things, sometimes wildly so.

If the approach works—and that's a meaningful if—the bottlenecks don't disappear. They shift. Eliminating multipliers might reduce die area and power, but memory remains a constraint. HBM4 reportedly entered mass production in the first half of this year, per TrendForce, but supply remains tight. CoWoS capacity is expanding, but not fast enough to saturate demand. And PyTorch/FSDP2 compiler support is necessary but not sufficient; the ecosystem needs to trust the stack before committing training runs that cost millions.

The U.S. export controls introduced in 2022 and expanded through 2024 add another layer of complexity. Advanced computing ICs face restrictions on shipments to China and certain other destinations, shaping both market access and competitive dynamics. The CHIPS Act—with awards to Intel, TSMC, Samsung, and Micron—aims to bring advanced nodes and packaging capacity onshore, but those fabs won't reach full production for years.

The Next Twelve Months

Digital illustration for article section "The Next Twelve Months" in "Inside Baud Labs' Plan to Eliminate Multipliers from AI Training" - A conceptual, minimalist hand-drawn illustration representing a radical breakthrough against establi...

Baud's bet is that the current infrastructure is inefficient enough that a radically different approach can break through. Perhaps they're right—stranger things have happened in Silicon Valley. Or perhaps the industry's inertia around NVIDIA's CUDA ecosystem, PyTorch integrations, and proven benchmarks will prove too heavy to overcome.

The next twelve months will clarify which. For now, Baud is a three-person team with a bold claim, a working FPGA demo, and a summer spent convincing early customers to take a chance on silicon that doesn't multiply.

In an industry where the infrastructure is straining, the power bills are climbing, and the supply chains are stretched thin, that might be just enough of an opening. Or it might not. We'll know soon enough.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Dunzo's Kabeer Biswas Raises $12M for AI Concierge Startup M
  • PhysicsX Hits $2.4B Valuation Building AI for Chip Manufacturing
  • Nearly Half of YC's Spring 2025 Batch Built AI Agents
  • AI Startups Challenge $7B Cat Modeling Giants After Record Losses
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.