Baud Labs, a three-person outfit backed by Y Combinator, surfaced in August with an unusual pitch: an AI training chip that does away with hardware multipliers entirely. The San Francisco startup says its approach compresses neural network weights by more than tenfold without sacrificing accuracy, and it has already validated an initial ASIC design on GlobalFoundries' 12-nanometer manufacturing process.
The claim sounds radical, but it builds on a thread of academic work stretching back years. Neural networks have traditionally leaned hard on multiplication, an operation that consumes outsize amounts of silicon real estate and power. Baud's founders argue they've found a way around that bottleneck by recasting the arithmetic itself.
The company, which participated in Y Combinator's Summer 2026 cohort, intends to tape out its first chip before year's end. For now, workloads are running on an FPGA-based emulation cluster, a stopgap that lets Baud demonstrate the concept before committing to silicon.
Rethinking the Math
At the heart of Baud's technology sits what the company calls a lossless representation that eliminates multiplications in both the forward and backward passes of neural network training. Traditional multipliers in silicon grow roughly with the square of their bit width, a scaling problem that eats into chip area and energy budgets. Adders, by contrast, grow linearly. Baud's architecture exploits that gap by sidestepping multiplies altogether.
The startup has built a compiler that converts models exported from PyTorch into its proprietary format. According to the company, the tool produces bit-exact results in most cases and already supports a range of architectures including GLM, Qwen, Gemma, DeepSeek, Flux, and Wan. A demo hosted on Baud's website shows a 50-million-parameter model trained on 5 million tokens generating inference at better than 1,000 tokens per second on a single Xilinx U200 FPGA running at 125 megahertz.
Whether that proof of concept translates to competitive silicon remains the open question. The FPGA emulation offers a glimpse, not a guarantee.
The Team Behind It
Sarang Zambare, Baud's chief executive, previously ran machine learning efforts for Peloton Guide and was among the founding ML engineers at Caper, a startup Instacart later acquired. Eric Taylor, the chief hardware architect, has four chip tape-outs to his name and has worked at NVIDIA, Freescale (later NXP), Arteris IP, and Enfabrica.
Baud operates under the legal name Cerelyze, Inc., according to trademark filings. It's a lean operation by design, perhaps betting that a tightly focused technical approach can compete in a market increasingly crowded with well-funded alternatives.
Academic Roots
The notion of neural networks without multiplication isn't new. A 2018 paper demonstrated inference and training on constrained devices without floating-point operations or multiplies. More recently, Microsoft Research published work on its BitNet architecture, particularly a 1.58-bit variant that uses ternary weights of negative one, zero, and positive one. In that scheme, matrix multiplications collapse into simpler operations: addition, subtraction, and the occasional skip. Those papers appeared on arXiv in 2024 and into early 2025.

Baud takes the concept further by embedding the arithmetic changes directly into custom silicon, targeting the energy and area costs at the hardware level rather than relying solely on algorithmic shortcuts. That distinction matters in data centers where power budgets and chip footprints dictate economics at scale.
A Crowded Field
The timing of Baud's emergence is both opportune and challenging. AI chip demand continues to strain manufacturing capacity and advanced packaging supply chains. Data center AI capital expenditures could reach $1.7 trillion by 2030, according to projections from Dell'Oro Group released in February. Gartner, in an August report, forecast AI processing semiconductors growing at a compound annual rate of nearly 27 percent through the end of the decade.
But the supply side remains tight. TSMC's CoWoS advanced packaging capacity, critical for integrating high-bandwidth memory, stayed constrained through much of 2026 despite expansion plans aiming for a 60-percent boost by 2027, according to TrendForce in April. SK hynix broke ground on a U.S.-based HBM production facility in Indiana in late August, though near-term bottlenecks persist.
Baud enters a landscape dominated by NVIDIA's Blackwell platform, AMD's Instinct MI400 line, Intel's Gaudi 3, Google's TPU v5p, and Microsoft's Maia 200. AWS Trainium2 competes on the hyperscaler side, while a cohort of venture-backed challengers jostles for position: Cerebras with its wafer-scale engines, d-Matrix (which reached full production with its Corsair inference platform in June), SambaNova, Tenstorrent, and Etched, a specialist in transformer ASICs.
OpenAI and Broadcom revealed "Jalapeño," a co-designed inference chip, in July, with engineering samples already running workloads tied to what the companies called GPT-5.3-Codex-Spark. Microsoft's Maia 200, unveiled in January on TSMC's 3-nanometer process with 216 gigabytes of HBM3e, targets inference at FP8 and FP4 precision.
Baud distinguishes itself not on raw inference speed but on training capability and the underlying arithmetic architecture. The company's compiler already handles pretraining, supervised fine-tuning, reinforcement fine-tuning, and inference on the FPGA cluster. When silicon arrives, Baud claims it will deliver performance improvements measured in orders of magnitude compared to incumbents at equivalent process nodes and memory bandwidth. That's the promise, anyway.

Silicon Still to Come
The GlobalFoundries validation was a first checkpoint, confirming the design could be manufactured at 12 nanometers. Baud has opened an early-access program that allows partners to test workloads on FPGA hardware before the ASIC tape-out.
Independent verification will depend on actual chips coming back from the fab and running production workloads. The company has not disclosed funding amounts, investor names beyond Y Combinator, or revenue projections. For a team of three, that opacity is perhaps understandable. The real test lies ahead, when Baud's novel arithmetic meets the unforgiving benchmarks of the data center.
