Three engineers working out of a cramped San Francisco office think they've found a shortcut around one of computing's most fundamental operations. Their claim sounds almost too bold to be credible: train and run neural networks without ever multiplying two numbers together.
Baud Labs—the commercial face of a legal entity called Cerelyze, Inc.—emerged from Y Combinator's Summer 2026 batch with hardware already running on field-programmable gate arrays and a custom chip design supposedly headed for fabrication before the year ends. The pitch is deceptively simple. Replace multiplication with addition, and suddenly you can cram far more compute into the same sliver of silicon.
It's the kind of proposition that makes veteran chip architects either sit up straight or roll their eyes, depending on how many times they've heard a version of this story before.
Why Multiplication Matters (and Why It's Expensive)
The technical argument hinges on a trade-off that's been baked into chip design for decades. Multipliers scale quadratically as you increase bit-width. Adders scale linearly. If you're designing a neural network accelerator, that quadratic growth becomes punishingly expensive as models demand higher precision and larger parameter counts.
Baud's founders argue they can sidestep the problem entirely. By co-designing hardware around an arithmetic representation that eliminates multiplications—compressing neural networks during both forward and backward passes—they claim to pack significantly more compute and memory into the same die area, at the same process node, with the same memory bandwidth.
Whether that works in practice is another question entirely.
The company says it has built a compiler that converts PyTorch exportable models to its proprietary format, claiming bit-exact results in most cases. The homepage lists compatibility with models including Qwen, DeepSeek, GLM, Flux, and Wan, though it stops short of explaining exactly which PyTorch export paths are supported or what "bit-exact results in most cases" actually means in production.
That vagueness may be strategic. Or it may reflect the reality of a three-person team still working out the details.
A Different Kind of Hardware Bet

What sets Baud apart from the current wave of AI accelerators isn't just the math. It's the commitment required to use their platform.
In its Y Combinator launch materials, the company drew a sharp line between itself and specialized inference players like Cerebras, Groq, Etched, and d-Matrix. Those companies, Baud argues, "focus on inference" and "leave the math side untouched" because they need to run existing pretrained weights without modification. Drop-in compatibility is the selling point.
Baud is making the opposite wager. If you're willing to train—or retrain—in their format from the ground up, they believe the performance gains will justify the friction. That's a harder sell in a market where model labs are already stretched thin and retraining frontier models can cost millions of dollars.
The company is targeting pretraining, supervised fine-tuning, reinforcement fine-tuning, and inference, claiming order-of-magnitude speedups for teams willing to absorb the conversion costs—though those performance claims have yet to be independently verified.
Perhaps more than the founders expected, that bet depends not just on performance, but on convincing engineers to trust an unproven arithmetic representation with their most valuable asset: the models themselves.
From FPGA to ASIC (Maybe)

According to the company's July announcement, Baud's first chip has been validated on GlobalFoundries' 12nm process and is on track for tape-out by year-end, though no independent verification of the design or timeline has surfaced yet—typical for pre-silicon startups but leaving room for skepticism.
For now, the company is running early access on FPGAs. A public demo showcases a 50-million-parameter model trained on 5 million tokens from the SimpleStories dataset, running on a Xilinx Alveo U200 FPGA at 125 MHz. Baud claims inference speeds exceeding 1,000 tokens per second on a single FPGA, though no external benchmarks have been published. The demo page calls the setup a "proof of concept" and warns that "the model will make mistakes."
Fair enough. But context matters.
Groq's LPU has clocked 241+ tokens per second on Llama 2 70B in benchmarks published over the past year or so. Baud's demo model is far smaller—orders of magnitude smaller—making direct comparisons tricky at best. The real test will come when the ASIC ships and independent teams can run their own workloads.
If it ships.
Early Access, With Caveats
Baud has opened early access to what it describes as its "first cluster," which runs on FPGAs emulating the planned ASIC. According to the design partner intake form, the cluster supports pretraining, supervised fine-tuning, reinforcement fine-tuning, and inference "up to a certain model size." The company hasn't disclosed capacity limits, system specs, or pricing.
The early access program reads like a practical move for a hardware startup attempting to introduce a new arithmetic representation. Getting design partners on board before the ASIC ships can surface edge cases in the compiler and reveal which model architectures translate cleanly—or catastrophically—to Baud's format.
It's also a way to gather testimonials and real-world validation before asking investors or customers to bet bigger on unproven silicon.
The People Behind the Platform
Baud was founded by Sarang Zambare and Eric Taylor, two engineers with complementary but distinct backgrounds.
Zambare, the CEO, brings over seven years in deep learning and AI hardware. He led machine learning for Peloton's Guide product from concept through shipping 100,000+ devices and was a founding ML engineer at both Peloton and Caper, which Instacart later acquired. He holds four patents and is, somewhat improbably, a certified pilot.
Taylor, the chief hardware architect, has four tape-outs and two major IP releases to his name, along with two patents and a publication. His resume includes stints at NVIDIA, Freescale/NXP, Arteris IP, and Enfabrica. That kind of ASIC architecture experience is critical when you're attempting to bring a novel chip design from concept to production in under a year.
The team of three is working out of San Francisco. Interestingly, the company's legal entity, Cerelyze, Inc., previously appeared in Y Combinator's Summer 2023 batch with an entirely different product focused on converting research papers to code. That company's website now displays a placeholder message: "Building something cool, brb.."
The nature of that pivot—whether strategic shift or complete restart—remains unclear, though it suggests Baud's current incarnation is newer than it might appear at first glance.
A Fragmented Market, a Narrow Opening

Baud is entering a hardware landscape that has splintered sharply in recent years. NVIDIA's Vera Rubin platform, announced at GTC earlier this year, continues to dominate training infrastructure for frontier labs including Anthropic, Meta, Mistral, and OpenAI. Meanwhile, specialized inference accelerators have proliferated at a dizzying pace.
d-Matrix's Corsair platform entered full production in June on TSMC's N6 node. Etched is building transformer-specific ASICs and investing in frontier inference clusters. OpenAI partnered with Broadcom on the Jalapeño accelerator, unveiled in June, signaling a broader trend of hyperscalers and model labs pursuing custom silicon co-designed with their software stacks.
The market has essentially bifurcated: massive distributed training platforms on one side, hyper-optimized inference accelerators on the other.
Baud is positioning itself in both categories. But there's a catch, and it's not a small one. You need to adopt their arithmetic representation. That's a fundamentally harder sell than drop-in acceleration. If the performance claims hold up, though, it could be worth it for teams willing to retrain or fine-tune in Baud's format from the outset.
The question is whether enough teams will be willing to take that leap.
Proof in the Silicon
Whether Baud can deliver on its order-of-magnitude performance claims will become clearer once the ASIC ships and independent benchmarks emerge. For now, the company has a working FPGA demo, a design partner program, and a tape-out date on the calendar.
Tape-out dates, of course, have a way of slipping. And even when chips come back from the fab on schedule, they don't always work as intended. The history of semiconductor startups is littered with promising architectures that never made it past the prototype stage.
Still, Baud's bet is an interesting one. If they're right, they've found a genuine shortcut around one of computing's most expensive operations. If they're wrong—or if the trade-offs don't pencil out in real-world workloads—they'll have joined a long list of startups that discovered the hard way that hardware is unforgiving.
The proof, as always, will be in the silicon.
