By the time most engineering teams review code, the damage is already done—or at least, that's the argument a Boston startup is making with fresh venture capital behind it.
Baz, which emerged from the minds behind Bridgecrew (sold to Palo Alto Networks for roughly $156 million), just raised a $9 million seed extension and launched what it's calling Planner: an AI system designed to catch bugs, security holes, and architectural blunders before developers write a single line of code. The extension, co-led by Battery Ventures and boldstart ventures, pushes Baz's total seed funding to $17 million. AFG Partners and Disruptive VC came in as new investors.
The timing matters. Baz unveiled Planner on June 29, 2026, at the AI Engineer World's Fair in San Francisco. This isn't just another code review tool. It's a bet that the entire concept of code review needs to shift earlier in the development cycle, especially now that AI systems are churning out production code at scale.
"Baz Planner intervenes at the planning stage to help eliminate entire classes of bugs before code is even authored," said Ed Sim of boldstart ventures. It's a framing that positions the product less as a safety net and more as a gatekeeper.
Four Agents Walk Into a Code Plan
Here's how it works: Planner deploys four specialized AI agents that operate in what the company describes as "dynamic loops," scrutinizing proposed changes against a risk matrix before implementation begins. When they spot trouble—security vulnerabilities, architectural inconsistencies, potential bugs—they don't simply raise a flag. They collaborate to diagnose root causes, propose fixes, validate those patches, and sometimes rewrite entire plans.
One of those agents, the Spec Reviewer, runs on Amazon Bedrock AgentCore, using browser tools and Model Context Protocol servers to test how proposed changes would interact with live specifications. Baz detailed that architecture with AWS engineers in a June 2 blog post, lending technical credibility to what could otherwise sound like marketing speak.
The core thesis is straightforward: catch problems at the planning stage, and you eliminate entire categories of downstream issues rather than playing whack-a-mole with individual bugs after merge.
The Pedigree and the Pitch

Baz's founders know this space intimately. Guy Eisenkot (CEO) and Nimrod Kor (CTO) previously built Bridgecrew, an infrastructure-as-code security platform that Palo Alto Networks acquired in March 2021 for about $156 million in cash. That exit—and the focus on pre-emptive security—shows up clearly in how Baz approaches the problem.
Barak Schoster, a partner at Battery Ventures, described Baz as a "super harness" for fleets of coding agents, emphasizing security, quality, reliability, and now planning. The investment logic seems to be this: as AI-generated code becomes ubiquitous, the governance layer grows more valuable than the generation layer itself. Perhaps it's a hedge—if models can write code faster than humans can safely review it, the bottleneck shifts to the layer that prevents bad code from ever entering the pipeline.
Whether that thesis plays out remains to be seen. For now, Baz claims to have signed more than 100 customers since launching its core agentic code review platform in 2025, though this figure comes from the company without independent verification. LinkedIn data from early June 2026 places the team at 11 to 50 employees, with offices in Boston, Tel Aviv, and Hamburg. An undated Datadog case study notes the company processes roughly one million LLM operations daily—a scale that suggests real usage, though it doesn't necessarily translate to retention or expansion.
What Early Users Are Saying

The most detailed case study available comes from LSports, a sports data provider, which tested Baz's predecessor system in May 2026. Across 241 repositories, LSports reported deployment frequency jumped 5.3 times, merge rates improved by 7.1 percentage points, and P90 lead time dropped 26 percent—from 96.9 hours down to 71.7 hours. The company credited Baz with catching 226 production-impacting bugs and three security vulnerabilities during review.
Baz also claims the top precision score on the Code Review Bench, an independent benchmark run by Martian in March 2026. That matters, though the competitive landscape here is metric-sensitive and increasingly crowded. CodeRabbit claims the highest F1 score on the same benchmark. Qodo says it leads on "toughest bugs." Kilo insists it's the top open-source reviewer "across all beta values." Different vendors optimize for different measures, and evaluation methodologies are still evolving. It's a bit like watching startups argue over leaderboard rankings—valid, but context-dependent.
Early adopters of Planner have reported a reduction of more than 65 percent in downstream rework, as measured by the frequency of revert and hotfix pull requests after merge, according to company-provided data.
Pricing runs around $20 to $50 per active developer per month. A Pro tier sits at $30, with a credit-based option at $0.01 per credit. A typical session with Baz's Fixer agent costs $1 to $4.30; the Advanced Security Reviewer runs $1.75 to $2.20. Not cheap, but not outrageous for enterprise dev tools, either.
A Crowded, Fast-Moving Market

Planner arrives as engineering teams confront a new normal: AI doesn't just assist with coding anymore—it writes significant chunks of production code. Cursor launched a web app in June 2025 to manage background coding agents outside the IDE. Academic papers in 2026 have explored multi-agent, human-gated workflows for code review. The agentic coding wave has moved from theoretical to operational.
What Baz is wagering on is a shift in where the bottleneck sits. If models can generate code faster than humans can safely review it, the value migrates to the layer that stops bad code from entering the pipeline in the first place. Planning-stage intervention becomes a leverage point—theoretically, one that eliminates classes of problems rather than individual instances.
But whether that thesis holds will depend on whether engineering teams see Planner as essential infrastructure or just another AI tool in an already bloated stack. The market for developer tooling is notoriously fickle; tools that seem indispensable one quarter can feel redundant the next as workflows evolve.
For now, Baz has $17 million in the bank and a product it believes addresses a problem that intensifies as AI writes more code. The next year or so will reveal whether customers agree—and whether "shift left" is more than just the latest DevOps catchphrase.
