Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
SaaS iconSaaSJuly 19, 2026

Toyota Spinout Walden Robotics Lands $300M Seed at Unicorn Valuation

Toyota Spinout Walden Robotics Lands $300M Seed at Unicorn Valuation
RoboticsAutomotive Tech+3
Healthtech & Biotech iconHealthtech & BiotechJuly 19, 2026

Auxilium Health Raises $3.4M Seed for Antibiotic-Free Wound Tech

Auxilium Health Raises $3.4M Seed for Antibiotic-Free Wound Tech
BiomaterialsAntimicrobial Resistance+3

Founders Mentioned

Sarang Zambare

Baud Labs

saas icon
SaaS

Sarang Zambare

Baud Labs

saas icon
SaaS
SaaS iconSaaS
July 19, 2026
YcAi HardwareSemiconductor TechAi Infrastructure

The Race to Reinvent AI Chips: Startups Challenge GPU Dominance

As AI semiconductor demand hits $400B annually, startups like YC-backed Baud Labs are pioneering radical architectures—from multiplier-free chips to transformer ASICs—to solve power, memory, and cost bottlenecks that GPUs can't.

The Race to Reinvent AI Chips: Startups Challenge GPU Dominance

In a cramped San Francisco office, three people are trying to rewrite the basic arithmetic of artificial intelligence. Baud Labs—young even by startup standards—has built a prototype chip that does something counterintuitive: it eliminates multiplication, the bedrock operation that's powered every GPU and AI accelerator for a decade. Instead, simpler math. The team claims they can compress model weights by more than 10x. Their demo system, a field-programmable gate array clocking a modest 125MHz, generates over 1,000 tokens per second.

It sounds almost absurd. It's also characteristic of a moment in the semiconductor industry where the old playbook no longer pencils out.

Research projections for 2026 indicate increasing challenges in memory bandwidth and infrastructure expenditures, with bottlenecks growing too acute, capital outlays too staggering, and inefficiencies too glaring for the status quo to survive uncontested. The result is a Cambrian explosion of architectures—some incremental, others radical—each probing for advantage in a market where margins matter and constraints multiply.

The Trillion-Dollar Scaffold

Start with the numbers, which frame everything else. AI semiconductors are projected to account for roughly 30% of the global semiconductor market in 2026, according to Gartner—a market projected to exceed $1.3 trillion. That puts AI chip revenue in the neighborhood of $400 billion annually. Hyperscaler capital expenditure on AI infrastructure is projected to climb more than 50% in 2026 alone, a figure that would have seemed reckless just two years ago but now reads as table stakes.

NVIDIA remains the sun around which much of the ecosystem orbits. The company's Blackwell architecture—embodied in the GB200 NVL72 rack-scale system with 36 Grace CPUs and 72 Blackwell GPUs—dominated MLPerf Training v6.0 benchmarks in June. Liquid-cooled, laced together with NVLink, the GB200 represents the current apex of what billions in R&D and a decade-deep moat can produce.

But the landscape is diversifying, perhaps faster than most observers expected. AMD's Instinct MI325X, planned based on earlier roadmaps, features 288GB of HBM3E memory and 6 TB/s of bandwidth; the MI350 series is already queuing up behind it. Google's Trillium TPU v6, announced in 2024 and now training Gemini models, delivers 4.7x the per-chip compute of its predecessor. AWS Trainium3 UltraServers, which reached general availability in December 2025, pack up to 20.7 terabytes of HBM3e and handle production inference for customers like Anthropic. Microsoft's Maia 200, built on TSMC's 3nm node and optimized for FP4 and FP8 precision, powers GPT-5.2 and Copilot.

Each hyperscaler now designs custom silicon tailored to its workloads. The shift is both strategic and economic. When you're deploying hundreds of billions of dollars in infrastructure, even small gains in cost-per-token or power-per-token compound into real money—fast.

Three Constraints, Tightening

Memory bandwidth, power availability, and manufacturing capacity. These three forces are reshaping the market, and they're tangled together in ways that resist easy fixes.

Memory is arguably the tightest bottleneck. IDC flagged high-bandwidth memory—HBM—as the primary constraint choking the semiconductor market's surge past $1 trillion this year. TrendForce projected in 2025 that HBM shipments will surpass 30 billion gigabits in 2026, with HBM4 expected to overtake HBM3e as the mainstream standard by the second half of the year. Samsung entered HBM4 mass production early in 2026, yet supply-demand imbalances remain sharp. Negotiations for 2027 supply began by mid-2026 amid high demand, which may lead to price hikes due to the squeeze.

Advanced packaging faces similar pressure. TSMC's CoWoS technology—chip-on-wafer-on-substrate, essential for integrating logic dies with HBM stacks—has seen demand explode. The company reported a supply-demand gap of around 20% that's expected to narrow to roughly 10% by year-end as new capacity comes online. TSMC's CoWoS is experiencing high demand with expected substantial growth, according to forecasts from 2022. On July 16, the company announced an additional $100 billion investment in Arizona—four or more 2nm fabs plus advanced packaging facilities—bringing total U.S. commitments to approximately $265 billion. TSMC executives cited "strong multi-year AI demand" and raised their 2026 revenue growth guidance to slightly above 40%.

Power is the third constraint, and it's not abstract. Data center electricity demand jumped 17% in 2025, driven largely by AI workloads, according to the International Energy Agency. Gartner forecasts data center power demand will climb another 27% this year, reaching roughly 132 gigawatts. The Electric Power Research Institute's Powering Intelligence 2026 study suggests U.S. data center load could hit 9% of total electricity generation by 2030 in high-growth scenarios. Jensen Huang has used the "AI factories" framing repeatedly over several years, including in 2026, describing data centers that produce tokens as a commodity—underscoring the shift: tokens per watt and tokens per dollar are the metrics that matter now.

These forces favor specialization. If general-purpose GPUs are overkill for your workload—or if memory and power constraints make them prohibitively expensive—domain-specific accelerators start to look less like a gamble and more like a rational hedge.

The Custom Silicon Wave

Digital illustration for article section "The Custom Silicon Wave" in "The Race to Reinvent AI Chips: Startups Challenge GPU Dominance" - A clean, minimalist conceptual representation of the "Jalapeño" custom inference silicon chip, actin...

OpenAI's collaboration with Broadcom offers perhaps the clearest signal that custom inference chips have moved from theory to practice. In late June, the two companies unveiled "Jalapeño," an LLM-optimized inference ASIC designed from scratch and taped out in just nine months. Engineering samples are running ML workloads in labs now; initial deployment is targeted for later this year. Broadcom brings ASIC design chops, Celestica handles board and rack integration, and OpenAI contributes workload expertise. Nine months from concept to tape-out is fast by semiconductor standards—and suggests a new playbook for companies with scale and specific needs.

Etched took a more aggressive approach. The startup emerged from stealth on June 30 with working silicon on TSMC's N4P node, $800 million in total funding, and over $1 billion in signed customer contracts. The company's valuation reportedly hit $5 billion in its December 2025 round. Etched's chip is a transformer-specific ASIC—hardwired for attention mechanisms and nothing else. It sacrifices flexibility entirely for raw performance on a narrow set of models. The company is now validating rack-scale deployments with customers, betting that efficiency gains justify architectural rigidity.

Groq raised $650 million in June to expand its Language Processing Unit inference cloud. The company's claim to fame is ultra-low latency; it powers the official Meta Llama API and recently struck a licensing deal with NVIDIA to integrate Groq's inference technology into the LPX platform announced at GTC 2026. Groq's architecture is deterministic and synchronous, eliminating the scheduling overhead that plagues GPUs at inference time. It's a niche. But one that commands a premium when milliseconds matter.

d-Matrix shipped its Corsair inference accelerator into full production in June, built on TSMC's N6 node with Alchip. The chip employs digital in-memory compute—a technique that performs matrix operations directly in SRAM to reduce data movement. The company positions Corsair for "agentic AI inference workloads," where multi-turn interactions and long-context windows stress memory subsystems. Like Groq and Etched, d-Matrix is targeting a wedge where incumbents are suboptimal.

The Multiplication Heresy

Which brings us back to Baud Labs. Founded this year and currently in Y Combinator's Summer 2026 batch, the team consists of CEO Sarang Zambare—who led ML for Peloton Guide and was a founding ML engineer at Caper, later acquired by Instacart—and Chief Hardware Architect Eric Taylor, a veteran of NVIDIA, Freescale/NXP, and Enfabrica with over a decade of ASIC design and four tape-outs under his belt. Their technical bet is radical: replace multiplication with addition-based arithmetic in neural network computations.

The idea isn't entirely new. Academic work like AdderNet, published between 2019 and 2021, explored replacing multiplications with additions in convolutional neural networks, though training stability required special techniques. Baud claims to have developed "a new arithmetic representation of neural networks that does not contain multiplications," co-designed silicon to execute that representation efficiently, and built a PyTorch-compatible compiler to convert standard models into their format. They've validated the approach on GlobalFoundries' 12nm process and are targeting tape-out by year's end. Their live FPGA cluster already handles training and inference workloads, at least at a proof-of-concept scale.

The value proposition, if it works, is compelling. Multipliers are among the most area- and power-hungry components in a chip. Eliminating them could yield ASICs with smaller, simpler cores than GPU tensor units or TPU processing elements. Baud claims their design uses commodity DRAM, standard interconnects, and air cooling—avoiding the exotic memory and liquid cooling that rack-scale GPU clusters now require. The company reports achieving more than 10x weight compression "without loss in intelligence," which would ease both memory bandwidth and capacity constraints.

It's early, of course. The team is three people. The demo is a 50-million-parameter model trained on 5 million tokens, running at 125MHz on a Xilinx U200 FPGA—a proof of concept, not a production system. But the approach aligns with broader industry trends: low-precision arithmetic (FP8, FP4, even lower), algorithm-hardware co-design, and a relentless focus on reducing the cost and energy per token. NVIDIA introduced NVFP4 for both training and inference; AMD and others are pushing similar formats. MLPerf v6.0 benchmarks this year emphasized Mixture-of-Experts models, which dominate frontier systems, and low-precision workloads.

Cerebras continues shipping its wafer-scale CS-3 systems; the Condor Galaxy 3 supercomputer uses 64 of them. Lightmatter is advancing photonic interconnects—sampling 1.6 terabits-per-second-per-fiber co-packaged optics and joining NVIDIA's NVLink Fusion ecosystem—to address the data-movement bottleneck at rack and cluster scale. Each company is probing a different angle.

An Industry Gearing Up

Digital illustration for article section "An Industry Gearing Up" in "The Race to Reinvent AI Chips: Startups Challenge GPU Dominance" - A clean, minimalist composition featuring a central, stylized architectural structure representing a...

The AI chip market in 2026 looks less like a winner-take-all race and more like an ecosystem under rapid diversification. Specialization is accelerating because the economics demand it and because manufacturing capacity exists to support heterogeneity at scale. TSMC's Arizona expansion, Micron's CHIPS Act-backed HBM packaging investments (up to $6.165 billion in incentives supporting a roughly $200 billion U.S. expansion vision), and Samsung's HBM4 ramp all point to an industry preparing for sustained, multi-year AI infrastructure buildouts.

What happens next depends on which bottlenecks tighten first. SemiAnalysis noted in mid-2026 that AI's share of TSMC's N3 wafer demand is roughly 60% this year and projected to hit around 86% in 2027. Packaging constraints are easing as CoWoS capacity expands, but HBM remains the choke point. If memory supply loosens, power becomes the binding constraint; Gartner already frames "power security" as the next economic battleground.

For founders and CTOs evaluating next-generation infrastructure, the strategic calculus is shifting. Do you lock in GPU capacity and ride the NVIDIA/AMD/Google roadmap, accepting the memory and power premiums? Or do you bet on domain-specific accelerators—inference ASICs, multiplier-free architectures, in-memory compute—that promise better efficiency at the cost of flexibility and ecosystem risk?

The answer may not be binary. Inference is already fragmenting across latency-optimized (Groq), transformer-specific (Etched), and hyperscaler-custom (Maia 200, Jalapeño) designs. Training workloads are adopting FP4/FP8 and Mixture-of-Experts, which favors architectures with fast interconnects and efficient data movement over raw FLOPS. vLLM's rapid evolution through 2026—adding FP8 KV-cache, tiered offload, and MoE handling—signals consolidation around high-throughput inference runtimes that can abstract over heterogeneous hardware.

Regulation and geopolitics add another layer of complexity. The U.S. Bureau of Industry and Security revised export licensing policy for advanced semiconductors to China on January 13, allowing case-by-case review for chips like NVIDIA's H200 and AMD's MI325X to approved customers—a shift from blanket restrictions. The EU AI Act entered full application in August, with transparency rules now in effect and member states issuing implementation guidance. The proposed Cloud and AI Development Act, introduced in June, aims to expand European cloud capacity and data sovereignty. These regulatory frameworks will shape where training happens, which architectures gain traction, and how supply chains evolve.

The Industrialization of Intelligence

Digital illustration for article section "The Industrialization of Intelligence" in "The Race to Reinvent AI Chips: Startups Challenge GPU Dominance" - A clean, minimalist conceptual illustration of a single, elegant, stylized processor block symbolizi...

The chips that succeed in 2027 and beyond will be those that solve real bottlenecks, not hypothetical ones. Memory bandwidth, power density, cost per token—these are measurable. Startups like Baud Labs are making technical bets that, if validated, could carve out niches or redefine categories. Larger players are hedging with custom silicon programs that give them control over roadmaps and economics.

Jensen Huang's "AI factories" metaphor captures something essential about this moment: intelligence is industrializing. The question is no longer whether alternatives to general-purpose GPUs will emerge, but which alternatives will scale and at what cost. Raw performance still matters, sure. But the race has shifted. It's about efficiency now, specialization, and the ability to manufacture at volume under constraints that aren't going away.

In that three-person office in San Francisco, Baud Labs is betting it can build a chip without multipliers that works. Whether they're right matters less, in some ways, than the fact that the bet is even plausible. The ground has shifted enough that heretical ideas get funding and airtime. The economics have shifted enough that conventional wisdom—GPUs everywhere, forever—no longer holds by default.

That's the real story. Not any single architecture or company, but the fracturing of a market that, until recently, looked settled. The constraints are real. The solutions are proliferating. And the winners, whoever they turn out to be, will be the ones who solved for what actually bottlenecks, not what bottlenecked five years ago.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • Toyota Spinout Walden Robotics Lands $300M Seed at Unicorn Valuation
  • Auxilium Health Raises $3.4M Seed for Antibiotic-Free Wound Tech
  • The Race to Build Pesticides That Save Bees, Not Kill Them
  • Markov Studios Launches Data Platform for Computer-Use AI Training
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.