The pitch sounds almost comically ambitious. A three-person team, freshly graduated from Y Combinator's summer program, claims they've figured out how to train AI models without multiplication—yes, that multiplication—while compressing model weights by more than 10× and preserving full accuracy. If you've spent any time around semiconductor startups, you know this is the kind of promise that usually precedes either a spectacular flameout or a quiet pivot to something more mundane.
But Baud Labs isn't just waving around a whitepaper. The company has completed early validation of its chip design on GlobalFoundries' 12nm manufacturing process. A working FPGA cluster is running live workloads. There's an online demo showing a 50-million-parameter model cranking out over 1,000 tokens per second on a single Xilinx U200 FPGA running at 125 MHz. And they're planning to tape out their first custom silicon by year-end, which in the chip world means they're past the point of pure speculation.
Whether any of this scales to the frontier models that actually matter—the ones consuming thousands of GPUs and millions of dollars per training run—remains very much an open question. But the timing is intriguing. NVIDIA's grip on AI training infrastructure has never looked more dominant or more vulnerable at the same time.
The Technical Gambit
Strip away the marketing language, and here's what Baud is proposing: a fundamentally different arithmetic representation that eliminates multiplication operations from both the forward and backward passes of neural network training. Not approximate multiplication. Not quantized multiplication. Just... gone.
This isn't a drop-in replacement. Models have to be trained in Baud's representation from scratch, which immediately creates an adoption hurdle. But the architectural payoff, if it works as advertised, is significant. Without multiplication units—which gobble up die area and power in conventional GPUs and TPUs—you can theoretically pack far more compute and SRAM onto each chip. Baud's design runs on commodity DRAM and standard interconnects, avoiding the high-bandwidth memory bottlenecks and exotic packaging that currently plague the industry. Air cooling instead of the liquid cooling rigs now proliferating in AI data centers.
The company describes its approach as "lossless," a word that does some heavy lifting here. Most compression and quantization schemes trade accuracy for efficiency. Baud insists there's no trade-off, at least for models trained natively in their system. Their PyTorch-compatible compiler supposedly produces "bit-exact results in most cases" for architectures including GLM, Qwen, Gemma, DeepSeek, and others that have become industry benchmarks.
That phrase—"in most cases"—is where things get interesting. The exceptions matter more than the company's technical documentation lets on.
A Market Straining at the Seams
The broader context makes Baud's timing look either prescient or opportunistic, depending on how charitable you're feeling. Gartner projects global AI spending will hit $2.59 trillion in 2026, up 47% year-over-year, with infrastructure accounting for more than 45% of that. The hyperscalers are responding with capital deployment that looks less like normal corporate investment and more like an arms race.
S&P Global Ratings estimated in March that Alphabet, Amazon, and Microsoft would collectively guide around $495 billion in 2026 capex, up 61% from the prior year. Goldman Sachs floated scenarios approaching $700 billion when you factor in the rest of the industry. The majority of those dollars are flowing into AI data centers, and specifically into the GPUs, memory, networking, and power infrastructure required to train increasingly enormous models.
But the supply chain is cracking under the pressure. TrendForce reported in June that high-bandwidth memory now consumes roughly 22% of total DRAM wafer production, with projections pointing toward around 30% in 2027. HBM4 contract prices are heading up as suppliers cite cost and supply imbalances. TSMC's advanced packaging roadmap extends toward 14-reticle interposers that can accommodate 10 compute dies and 20 HBM stacks by 2028, yet packaging capacity remains constrained right now, when customers actually need it.
Then there's power. The U.S. Energy Information Administration noted in January that the country is experiencing its strongest four-year electricity demand surge since 2000, with data centers as a primary driver. The International Energy Agency clocked 17% data center power growth in 2025 alone and expects AI workloads to keep pushing demand higher through the decade. Uptime Institute surveys show liquid cooling adoption accelerating as rack densities climb past what traditional air systems can handle economically.
These aren't abstract constraints. They translate into months-long GPU wait times, eye-watering rental rates, and frantic negotiations over power allocations in markets where data center construction has outpaced grid capacity. Which is to say: if Baud's architecture actually delivers on its efficiency claims, there's a market ready to listen.
Who's Behind This

Startups attempting to dislodge NVIDIA typically bring one of two things: breakthrough technology or experienced leadership. Ideally both. Baud has a decent pedigree on the leadership side, though the team is admittedly small.
Sarang Zambare, one of the co-founders, previously led machine learning for Peloton Guide, the fitness company's camera-based strength training product that shipped to over 100,000 devices. Before that, he worked at Caper, the smart shopping cart startup that Instacart acquired. He holds four patents. Eric Taylor, coming from the semiconductor side, has done time at NVIDIA, Freescale, NXP, Arteris IP, and Enfabrica—a solid roster of chip design houses. Four tape-outs and two major IP releases give him credibility in a domain where most ML researchers have never touched silicon.
The pairing makes sense. Building an AI accelerator startup requires both deep knowledge of how models actually train and the gritty expertise of getting chips manufactured and working. But a three-person team is still a three-person team, and they're going up against organizations with thousands of engineers, decades of institutional knowledge, and billion-dollar R&D budgets.
The Incumbent Countermoves
NVIDIA, unsurprisingly, hasn't been sitting idle. The Blackwell platform, unveiled at GTC in March 2024, is the company's latest bid to extend its dominance through sheer performance scaling. The GB200 Grace Blackwell Superchip links two B200 GPUs with a Grace CPU via NVLink-C2C. The NVL72 configuration connects 72 Blackwell GPUs and 36 Grace CPUs in a single NVLink domain with 1.8 TB/s per-GPU bandwidth.
NVIDIA's marketing materials claim up to 30× inference speedup and roughly 4× training performance gains versus the H100 generation on certain large language models. Technical blogs from late last year and this year highlight NVFP4 and FP8 precision formats that allegedly deliver up to 3.2× better performance for Llama 3.1 405B training compared to Hopper-generation chips at the same GPU count. Whether those numbers hold in practice is always debatable, but the direction is clear: NVIDIA is betting it can stay ahead through process technology, architectural refinement, and ecosystem lock-in.
AMD, Intel, and Google are running their own playbooks. AMD's MI325X shipped in late 2024 with up to 288 GB of HBM3E memory and 8 TB/s bandwidth. The MI350 series, built on CDNA4 architecture, is slated for this year, with AMD projecting 35× inference gains versus the prior generation—a claim that, like most vendor performance projections, deserves skepticism until independently verified. Intel's Gaudi 3 expanded availability last year, landing IBM Cloud as its first service provider deployment in April. Google's seventh-generation Ironwood TPUs reportedly delivered a 3.7× improvement in Compute Carbon Intensity versus TPU v5p as of January.
The startup ecosystem is also crowded. Cerebras announced 8 exaFLOPs of additional AI compute capacity in its Condor Galaxy network back in March 2024. SambaNova unveiled its SN50 chip for agentic inference in February alongside a $350 million-plus financing round that brought in Intel Capital and SoftBank as customers. Demos showed heterogeneous setups using NVIDIA H200 GPUs for the prefill phase and SN50 chips for decode, a configuration that hints at where the market might be heading—specialized chips for specific workload segments rather than one-size-fits-all solutions.
d-Matrix announced full production of its Corsair in-memory compute platform in June after a $275 million Series C last November at a $2 billion valuation. Tenstorrent brought its Galaxy Blackhole servers to general availability in April, delivering 23 petaFLOPS of FP8 performance per 6U server starting at $110,000. The field is getting crowded, in other words, and most of these companies have raised orders of magnitude more capital than Baud appears to have access to.
The Validation Gauntlet

Here's where Baud's narrative faces its sternest test. A 50-million-parameter model running on an FPGA is a proof of concept, not a product. The models that matter today—the ones companies actually care about—contain tens or hundreds of billions of parameters and require training runs across thousands of accelerators. NVIDIA's Nemotron-H, for instance, trained on 6,144 H100 GPUs. Meta scaled its Llama 3 training from 4,000 to 24,000 to eventually 129,000 GPUs, according to a September 2025 disclosure. OpenAI published research in May on its custom networking protocols for massive training clusters.
Can Baud's architecture scale to that regime while maintaining its efficiency advantages? That's not a rhetorical question. Many chip architectures that look promising at small scale hit walls when you try to distribute workloads across thousands of nodes. Communication overhead, synchronization latencies, fault tolerance—these issues become exponentially harder as cluster size grows.
Then there's the compiler compatibility claim. Baud says any PyTorch-exportable model can be converted with bit-exact results "in most cases." What about the cases where it can't? If certain model architectures, training techniques, or optimization methods aren't supported—or if accuracy degrades in ways that only become apparent during long training runs—adoption will stall. The ecosystem inertia around CUDA and PyTorch is immense. Customers evaluating alternatives to NVIDIA typically demand either drop-in compatibility or ironclad proof that accuracy and convergence characteristics remain identical.
Ecosystem support poses another challenge. Meta's revelations about scaling Ethernet and RoCE clusters to 129,000 GPUs without network bottlenecks required extensive co-design work. OpenAI developed its own Multipath Reliable Connection protocol atop RoCE with SRv6. NVIDIA's Spectrum-X Ethernet platform has been adopted by CoreWeave, Lambda, and GMO. Baud claims to use "standard interconnects," which sounds good until you realize that "standard" at hyperscale still requires deep integration work, custom drivers, and extensive testing.
Export controls add yet another variable. U.S. restrictions continue to limit advanced AI chip exports to China and other jurisdictions. In November, the Commerce Department authorized exports of up to 35,000 NVIDIA Blackwell-class chips to UAE's G42 and Saudi Arabia's Humain under specific conditions. Baud's 12nm process on GlobalFoundries sits well below the cutting-edge nodes that trigger the most stringent scrutiny, which could be an advantage in certain markets. Or it could signal a performance ceiling that limits competitiveness against chips built on TSMC's 3nm or 2nm processes.
What Comes Next

The case for Baud, if there is one, rests on a bet that the industry is approaching a breaking point. HBM shortages aren't getting better. Power constraints are tightening in key markets. Packaging capacity remains a bottleneck. The hyperscalers are spending historic sums but still facing allocation shortages and delivery delays. Meanwhile, the cost per training run keeps climbing as models grow larger and require more compute.
If Baud's technical approach genuinely eliminates multiplications in a lossless manner and scales to production clusters, it would represent a rare instance of architectural innovation that sidesteps rather than merely optimizes away the current constraints. No HBM dependency means no exposure to memory pricing volatility. Smaller die area and lower power consumption could translate to dramatically better economics at volume. Air cooling instead of liquid cooling reduces infrastructure complexity and cost.
That's a lot of ifs. The company hasn't disclosed funding beyond Y Combinator's standard batch investment, which typically ranges from $500,000 to a few million dollars depending on the program and the deal structure. A three-person team, a 2026 tape-out timeline, and an early-access program opening now all suggest this is still very early stage. Customer feedback loops will start before silicon ships, which is smart—better to learn about show-stopping issues in the FPGA phase than after committing to an expensive production run.
The broader question hovering over Baud and every other AI accelerator startup is whether the industry is actually ready to diversify beyond the GPU monoculture. NVIDIA's installed base, software ecosystem, and developer mindshare create switching costs that go beyond raw performance metrics. But trillion-dollar capex cycles, export control pressures, and supply chain fragility are all symptoms of a system under stress, perhaps more stress than it can bear indefinitely.
Baud is proposing to eliminate an operation that has defined AI accelerators since the field's inception. Whether that proposition scales from an FPGA cluster to a production system training frontier models will determine whether this is a footnote or the beginning of something larger. The company is taking early-access customers now, which means we'll start getting real-world signal relatively soon.
For an industry accustomed to incremental improvements and evolutionary roadmaps, watching a three-person team try to rewrite the rules has a certain appeal. Even if—maybe especially if—the odds are long.
