Training Meta's Llama 3.1 405B required approximately 24,576 H100 GPUs running continuously. The energy consumed in that single training run rivals what a small city burns through in a year. Now multiply that—ironically—by every frontier model being trained at OpenAI, Anthropic, Google, and dozens of well-funded startups. The industry's power problem comes into focus quickly.
According to Epoch AI, training power per run has been growing at more than double each year. Multiple-gigawatt-class training runs are likely by 2030 unless someone figures out a radically different approach.
Enter Baud Labs. Three people, YC-backed, operating out of San Francisco under the legal entity Cerelyze, Inc. Their claim seems almost absurd: a chip architecture that eliminates multiplication operations entirely from neural network training and inference. Not reduces them. Eliminates them.
They say they can deliver "orders of magnitude better performance than incumbents" through what they call a novel arithmetic representation. Their first chip, validated on GlobalFoundries' 12nm process, is on schedule for tape-out by late this year.
The timing, perhaps, is less coincidence than inevitability. Data center AI capital expenditures are racing toward projected figures in the trillions by decade's end. Power forecasts for U.S. data centers have nearly doubled in some recent estimates. The semiconductor industry is scrambling for alternatives to the GPU-centric paradigm that has dominated AI compute for the past decade.
Whether Baud has actually found one remains an open question.
The Energy Wall
The numbers tell a stark story, though the exact trajectory depends on whose forecast you trust. Goldman Sachs has estimated that global data center power demand tied to AI could increase 165% by 2030 compared to recent baselines, requiring approximately $720 billion in grid infrastructure spending. Other forecasts suggest that data centers could consume nearly 20% of American electricity within the next decade, with roughly half dedicated to AI training and inference.
That may sound alarmist. It's worth considering that the compute itself keeps scaling faster than hardware efficiency can compensate. Epoch AI's data shows training compute for frontier large language models has grown approximately five-fold per year since 2020, doubling roughly every 5.2 months. The global AI compute "stock" has been doubling every seven months.
Even as NVIDIA's Blackwell platform promises up to 25 times less cost and energy versus Hopper for trillion-parameter models, and Google's Trillium TPU claims nearly five times the compute per chip versus its predecessor, absolute energy consumption continues its relentless climb.
CoreWeave, the specialized GPU cloud provider, has surpassed 1 gigawatt of active data center power and is targeting more than 8 gigawatts by 2030. That's one company. McKinsey analysis suggests global data center demand could more than triple by decade's end, with the GPU-as-a-service market alone reaching tens of billions of dollars annually, excluding the hyperscalers' internal deployments.
The constraint is becoming less about chip availability and more about where to plug them in.
Utilities cannot approve interconnects fast enough. Municipal power grids designed for incremental growth are facing requests for hundreds of megawatts in single projects. The worldwide server market has been growing at rates above 30% year-over-year recently, driven by GPU servers, but supply constraints in non-accelerated segments signal broader infrastructure bottlenecks.
The Architecture Insurgency
Against this backdrop, a wave of specialized AI chip startups is attacking the problem from multiple angles.
Some, like Groq with its LPU inference engine, focus on deterministic execution and extreme throughput. Groq has raised significant capital and claims to deliver hundreds of tokens per second on standard Llama models. d-Matrix pursues in-memory compute for inference through its Corsair platform, which has entered production following a Series C that valued the company at $2 billion. Cerebras, with its wafer-scale approach, has raised over a billion dollars and has been positioning for an eventual public offering. SambaNova has closed funding at an $11 billion valuation.
Each takes a different architectural bet. Most, though, still rely on conventional arithmetic operations at their core.
Multiplication remains the fundamental bottleneck. In standard neural network math, matrix multiplications consume the majority of compute cycles and energy. Every forward pass, every backward propagation gradient, every parameter update involves countless multiply-accumulate operations.
This is where Baud's claim becomes intriguing—audacious, even.
The company states it has developed an arithmetic representation that eliminates multiplications and compresses neural networks in both forward and backward passes. The technical details remain sparse, as one might expect from a startup protecting potential intellectual property. What Baud has made public is that it built an ASIC "architected specifically for this representation" and created a compiler that can convert PyTorch-exportable models into its multiplication-free format. These claims remain self-reported without independent third-party validation.
The academic foundations exist, even if production implementations at scale do not. Microsoft's BitNet research, published in recent years, demonstrated that 1.58-bit large language models using ternary weights—values constrained to -1, 0, and 1—could achieve competitive performance while drastically reducing arithmetic complexity. Earlier work on binary neural networks, including the XNOR-Net architecture from 2016, showed that multiplications could be replaced with bit operations under certain constraints. Recent computer vision research has explored "AddBit-Operation-Only" binary neural networks designed explicitly for hardware efficiency.
Whether Baud has cracked the code on applying these principles to frontier-scale model training without accuracy loss remains unproven in the public domain.
The company operates an FPGA cluster emulating its ASIC and claims to be "running inference at 1000+ tokens per second on a single FPGA" on a small trained model. For context, Groq's production LPU systems were demonstrating similar token rates on 7-billion-parameter models in early stages of deployment. Baud's early-access partner program offers pretraining, fine-tuning, reinforcement learning post-training, and inference on the FPGA-emulated cluster, suggesting the platform is further along than vaporware.
At least in theory.
The GlobalFoundries 12nm validation is worth noting. While trailing-edge by semiconductor standards—NVIDIA's Blackwell uses TSMC 4nm, and Microsoft's Maia inference accelerator uses TSMC 3nm—it signals a pragmatic choice. Mature process nodes are cheaper, more available, and if Baud's efficiency gains come primarily from algorithmic innovation rather than transistor density, 12nm could be perfectly adequate. TSMC's advanced nodes are supply-constrained and expensive; GlobalFoundries' capacity is more accessible for a bootstrapped startup burning through seed funding.
The Incumbent Response

The incumbents are not standing still.
NVIDIA's Blackwell GB200 NVL72 rack-scale system connects 72 GPUs in a single NVLink domain with 130 terabytes per second of intra-rack bandwidth and 30 terabytes of HBM memory per rack. The company claims 4 times faster training versus the H100 generation and 30 times the inference throughput. Software innovations like the NVFP4 data type have delivered significant speedups on Llama 3.1 405B training versus Hopper FP8 baselines.
Microsoft's Maia 200 targets inference workloads with native FP8 and FP4 support and more than 200 gigabytes of HBM3e memory running at roughly 7 terabytes per second of bandwidth. It's designed to run alongside NVIDIA and AMD silicon in Azure's heterogeneous fleets, powering models including forthcoming GPT iterations. AMD's MI350 series promises a 35-fold generational increase in AI inference performance over prior generations, with rack-scale Helios systems aimed at large mixture-of-experts workloads.
Google continues refining its Trillium TPU, now in general availability. AWS iterates on Trainium2, though public technical details remain sparse compared to NVIDIA's disclosure cadence. The hyperscalers have collectively invested tens of billions in custom silicon precisely because the economics of renting NVIDIA GPUs at scale become prohibitive even for companies with their margins.
What none of these solutions fundamentally challenges, though, is the multiply-accumulate paradigm. They make it faster, more power-efficient, more tightly integrated. They reduce precision from FP16 to FP8 to FP4.
But the core arithmetic remains multiplication-heavy.
The Verification Gap
Here's the tension: Baud's claims are extraordinary, and extraordinary claims demand extraordinary evidence.
As of this writing, no independent third-party benchmarks have surfaced validating the "orders of magnitude" performance advantages the company asserts. The YC company page and Baud's own website are the primary sources. The demo is live, but trained on what the company describes as "a small model"—not a frontier 405-billion-parameter beast.
The history of semiconductor startups is littered with compelling pitch decks that failed to survive contact with production realities. Compilation complexity. Model conversion accuracy. Debugging toolchains, firmware stability, thermal management, yield rates at tape-out, ecosystem integration. The graveyard of good ideas is vast.
Even Cerebras, with its billion-dollar-plus funding and years of production experience, remains a niche player compared to NVIDIA's market dominance. Groq's significant capital raise and impressive inference benchmarks have not yet translated into widespread hyperscaler adoption at the scale required to move market share.
Baud's compiler, which converts PyTorch models to its format, will be critical. PyTorch has become the de facto standard for research and increasingly for production. If Baud requires extensive model retraining or architecture changes to achieve its claimed efficiency gains, adoption friction rises steeply. If conversion is truly lossless and the performance delta is as large as claimed, the value proposition becomes compelling very quickly—particularly for inference, where cost-per-token economics drive margin.
The early-access partner program offers some signal. If Baud can sign credible AI labs or model builders willing to invest engineering time in evaluation, that provides market validation separate from technical verification. The fact that the company is live with an FPGA cluster, rather than purely in simulation, suggests progress beyond PowerPoint.
But it's a long way from an FPGA cluster to production silicon at scale.
The Window of Opportunity

The broader industry context creates a peculiar window for challengers. Industry forecasts suggest AI server shipments could grow more than 28% annually in coming years, with a rising share of ASIC-based systems. That phrase—"ASIC-based systems"—is doing a lot of work. It encompasses everything from Google's TPUs and Microsoft's Maia to startups like Baud, d-Matrix, and SambaNova.
Export controls on advanced AI chips to China, tightened in recent years, have created segmented markets and design-to-spec optimization problems. A chip that delivers extreme efficiency at a mature process node could sidestep some control thresholds while still offering competitive performance. The CHIPS Act incentives, including GlobalFoundries' substantial award, aim to rebuild U.S. semiconductor manufacturing capacity. A startup choosing GlobalFoundries for its first tape-out aligns with that strategic direction.
Power and grid constraints are real and worsening. Recent industry analysis emphasizes hybrid and edge deployment to control per-token costs, stressing the importance of optimizing for power envelopes and total cost of ownership, not just peak throughput. If Baud's multiplication-free approach genuinely reduces energy consumption by the magnitude claimed, it addresses the industry's most acute physical constraint.
The challenge is that "orders of magnitude" is vague.
Is that 10x? 100x? At what model size, batch size, precision, and utilization? Against what baseline? NVIDIA's own Blackwell-versus-Hopper claims of "up to 25x" improvement come with context-specific asterisks around trillion-parameter models and specific workloads. Vendor-stated performance numbers are marketing as much as engineering.
The Path Forward
For Baud to matter, several things must be true.
The arithmetic representation must preserve model quality across a range of architectures—transformers, mixture-of-experts, multimodal models. The compiler must handle not just toy examples but production-scale code bases with minimal developer friction. The ASIC, when it tapes out later this year, must hit its power and performance targets at acceptable yield and cost. And the company must navigate the chasm between interesting technology and production deployment at the scale where infrastructure buyers care.
The founders—Sarang Zambare as CEO and Eric Taylor as Chief Hardware Architect—have assembled a credible technical thesis rooted in legitimate computer science. Whether that thesis survives the crucible of silicon implementation and market adoption is the billion-dollar question. Perhaps literally, if the technology proves out and the company scales.
The AI infrastructure market is large enough and desperate enough for efficiency gains that there is room for multiple winners. NVIDIA will remain dominant; their software moat and installed base are formidable. But just as AWS, Google, and Microsoft have all invested in custom silicon to optimize specific workloads, AI labs and enterprises will increasingly look for architectural diversity to manage cost and power.
If Baud's chip works as advertised, it becomes a potential hedge against both NVIDIA's pricing power and the industry's unsustainable energy trajectory.
If it falls short, it joins the long list of semiconductor startups that discovered building chips is easier in slide decks than in fabs.
By the time the company expects to complete its tape-out, the industry will have a clearer answer. Until then, the idea that you can train frontier models without multiplications remains tantalizingly possible and thoroughly unproven—a Schrödinger's chip, neither transformative nor vaporware until someone opens the box and measures what's inside.
