Fifteen percent. That's the share of AI-authored code commits that introduce at least one bug into software repositories, according to an analysis of more than 302,000 such commits. For Charles Pan (Stanford CS '22) and Dean Stratakos (Stanford CS), who'd built trading floor algorithms and watched AI agents churn out code at breakneck speed, that figure crystallized a concern they'd been nursing for months: the machines were writing faster than people could think.
Their response? A startup called Stage, which emerged from Y Combinator's Spring 2026 cohort with a deliberately contrarian pitch. While much of Silicon Valley races to automate code review itself—flooding the market with bots that comment, approve, and merge with minimal human oversight—Pan and Stratakos are building tools to make human reviewers faster and sharper. Not to replace them.
"We're not building a code review bot," Pan said flatly during Stage's April launch on Hacker News, a post that drew 130 upvotes and 111 comments, many from engineers skeptical they needed yet another tool cluttering their workflows.
The product itself launched that same month: a platform that breaks pull requests into what the founders call "chapters," organizing related code changes into a logical sequence rather than the typical file-by-file slog. When a developer opens a PR, Stage analyzes the diff and groups edits into small, ordered units—each with a summary and a list of "things to double check." Reviewers can navigate through the story of what changed, in the order it makes sense, not the arbitrary structure Git happens to spit out.
It's a modest idea, perhaps. But the timing matters.
The Accumulation Problem
Research published in March found that human reviewers exchange nearly 11.8% more comment rounds when scrutinizing AI-generated code compared to code written by other humans. And when AI suggestions do get merged? They tend to increase complexity more than contributions from flesh-and-blood engineers. Another study, updated in April, reported that nearly 23% of bugs introduced by AI commits persist all the way to the latest code revision. Almost 90% fall into the category researchers politely term "code smells"—the kind of rot that doesn't break things immediately but makes future work harder.
There's a term floating around for this: cognitive debt. The idea that AI tools let developers ship features they don't fully understand, deferring the reckoning to some future quarter when the bill comes due. A January paper documented developers explicitly acknowledging uncertainty around code they'd accepted from AI assistants, admitting they'd postponed testing or simply didn't grasp what they'd just merged.
Stage's founders know this firsthand. Stratakos, a Division I tennis player turned software engineer, led the AI push at Five Rings, a trading firm, and built an in-house coding agent. The experience left him convinced that the bottleneck wasn't code generation—it was comprehension. Pan, an early engineer at Yuzu Health after his own stint at Five Rings, saw similar patterns. Fast-forward to April 2026, and the two-person team, working out of San Francisco with YC partner Pete Koomen, had a product.
Against the Automation Wave

Stage's human-first stance puts it at odds with much of the current market. Anthropic launched Claude Code Review on March 9, positioning its AI as a tireless reviewer that can spot issues and approve changes autonomously. Tools like Greptile and CodeRabbit have sprung up with similar promises: automated agents that handle the grunt work of review so humans don't have to. Graphite bundles AI review with stacked PRs and merge queues. DeepSource pairs static analysis with AI oversight.
The pitch is seductive—who wouldn't want a bot doing the tedious parts? But Stage's founders argue the premise is backwards. The problem isn't that review takes human time. It's that AI-generated code is harder for humans to understand in the first place. Automation just papers over the gap.
Instead, Stage syncs comments and approvals back to GitHub, slotting into existing workflows rather than trying to own them. Public examples of the chapter interface are visible at stagereview.app/explore, though the company's own pricing page proved elusive during the April rollout—a detail that sparked frustration in early Hacker News threads. External reports put the hosted platform at $30 per seat monthly, with a 14-day trial. The company uses a mix of Google Gemini, Anthropic's Claude, and OpenAI models via an API gateway, according to security documentation dated late April.
Opening the CLI

On May 12, Stage open-sourced a command-line tool that brings the chapter concept to local development. Available through npm as "stagereview," the CLI works with any coding agent and opens a browser UI to display structured changes from the current branch. MIT-licensed, it includes an agent skill command—/stage-chapters—for direct integration with AI assistants.
"The CLI is completely free," Stratakos confirmed in response to pricing questions. It's a move that sidesteps the monetization headaches while seeding adoption. The GitHub repository has drawn interest, though community feedback has centered on practical concerns: How does the tool handle monster diffs spanning dozens of files? Does generating chapters create its own cognitive load? And what happens when the AI gets the narrative structure wrong?
The founders indicated that a terminal-native UI and deeper hooks into tools like Linear and GitHub Issues are on the roadmap. But those are plans, not products. For now, Stage is a two-person operation testing a thesis.
A Bet on Narrative

Stage enters a noisy, fast-moving space. Code review tools aren't new—GitHub's native interface has been the default for over a decade, and countless startups have tried to improve on it. What's new is the pressure AI-generated code puts on that system. When an agent can write 500 lines in seconds, the old rhythms break down.
The question is whether organizing PRs into chapters—essentially imposing narrative structure on what's usually a technical dump—resonates enough to overcome inertia and the gravitational pull of established tools. Stage's security posture looks standard for a developer tool: TLS in transit, AES-256 at rest, token encryption for GitHub OAuth. Chapter narratives get persisted; diffs are fetched on demand. Terms of service, effective late March, outline subscription billing through Stripe. All reassuring, if table stakes.
But the core bet isn't about encryption or pricing tiers. It's about whether the code review crisis is fundamentally a speed problem or a comprehension problem. Pan and Stratakos clearly believe it's the latter. As AI agents generate more code, the bottleneck shifts from writing to understanding—and understanding, they'd argue, is a distinctly human domain. One that benefits from structure and narrative, not just more automation.
Whether that thesis holds depends on something harder to measure than bug rates or review times: whether engineers, already juggling Slack and Jira and a dozen other tools, see enough value in chapters to add one more layer to their stack. If 15% of AI commits really do introduce issues, and those issues accumulate faster than teams can address them, someone probably needs to rethink how humans parse machine-generated code.
Stage is wagering that narrative—not another bot—is the fix. The next few months will show whether engineering teams buy it. Or whether, like so many well-intentioned devtools before it, the idea just gets lost in the noise.
