Picture the software engineer interviewing for a job. Odds are good they have Claude or GitHub's Copilot humming in a browser tab somewhere, maybe tucked behind the assessment window. That habit—now nearly universal—has created a small crisis in technical recruiting. Most companies either ban AI tools outright during coding tests, or they simply have no idea how candidates are deploying them in real time.
Saffron, a three-person startup that emerged from Y Combinator earlier this year, is wagering on a third path. Instead of trying to police the use of AI, the company's software embraces it, then forensically analyzes how a candidate actually wields the tool. The underlying premise is stark: "Good code is easy now. Good engineers aren't," reads the startup's homepage. It's a blunt distillation of a worldview—that in an age when AI can autocomplete entire functions, the interview process itself needs rewiring.
The founders aren't exactly shouting into a void. Recent industry surveys report that around 90 percent of software developers now use AI tools in their work. If nine out of ten engineers are already using these tools on the job, the logic goes, why structure interviews as though they don't exist?
A Problem Born at Jane Street
Robert Chondro, Jerry Yao, and Kazuma Choji came to the problem from different angles but arrived at the same conclusion. Chondro and Yao both spent time at Jane Street, the quantitative trading firm known for rigorous technical screens. Choji, a Harvey Mudd graduate with machine learning papers at NeurIPS and ICML, brought a research lens. Together, they'd experienced the awkwardness firsthand: hiring managers who wanted to assess real-world engineering chops but had no way to measure how candidates interacted with the AI tools they'd inevitably use on the job.
The traditional options felt insufficient. Lock down the assessment environment entirely—no internet, no assistants—and you're testing a skill set that barely resembles modern software work. Allow AI freely but with no visibility, and you're left guessing whether the candidate leaned on autocomplete for every line or just used it to scaffold boilerplate.
Saffron's answer is a platform that does more than permit AI—it captures the entire interaction. Candidates work in a browser-based IDE with Anthropic's Claude Code integrated directly. They write prompts, accept or reject suggestions, and navigate the codebase while Saffron's backend logs every keystroke, every AI prompt, every acceptance or rejection, and every code diff.
What the Platform Actually Does

Here's how it unfolds in practice. A company connects its own GitHub repository—or uses a template scenario—then defines what it cares about: code quality, system design thinking, problem decomposition, whatever the hiring rubric demands. The candidate receives a link and works through the problem without needing to install anything locally. Claude Code integration is built in.
What Saffron surfaces on the other end is where things get interesting, or invasive, depending on your perspective. A session replay feature lets reviewers scrub through the entire assessment like a director watching dailies. More than that, the platform deploys what it describes as over ten independent AI agents to evaluate the work against the company's custom criteria. The result is a deterministic score—at least in theory—with line-by-line attribution showing which chunks of code the human wrote versus what the AI suggested or modified.
Each test comes with a five-dollar Claude Code budget baked in. (Whether that's generous or stingy probably depends on how chatty the candidate is with the AI.)
The approach, with its emphasis on quantitative measures like AI reliance percentage and originality scores, reflects the founders' quantitative trading roots. Instead of a hiring manager's vague sense that "this person seems overly reliant on AI," Saffron aims to surface something closer to metrics. Whether engineering leaders actually want that level of granularity—whether they'll trust it—is another question entirely.
Pricing That Signals Early-Stage Hustle
Saffron has rolled out tiered pricing that feels both startup-scrappy and surprisingly structured. A Basic plan runs $199 monthly for five assessments, including AI-generated debrief questions and session replay. Premium, at $499, bumps the cap to 15 assessments and adds priority support. Enterprise customers get unlimited tests, single sign-on, team management, and a dedicated account manager, though the company hasn't publicly disclosed what that tier costs.
For companies wanting to dip a toe in, individual assessments go for $49 each. It's a pricing model that suggests Saffron is still figuring out its ideal customer—enterprise talent teams with volume needs, or smaller startups willing to experiment on a handful of critical hires.
What's notably absent: customer logos, case studies, or public testimonials. That's not shocking for a company that only recently came out of YC's latest cohort, with David Lieb as its primary backer at the accelerator. Afore Capital is also in the cap table. But the lack of public validation means Saffron is still in show-don't-tell mode.
A Suddenly Crowded Arena

Saffron isn't alone in sensing the moment. CodeSignal, one of the incumbent assessment platforms, launched what it called "agentic coding assessments" this spring, explicitly positioning them for the AI era. HackerRank has updated its platform with AI-assisted interview features. A cluster of newer entrants—Talia.ai, The Cognitive, ScreenDesk, RoundOne AI—are attacking the same problem from different angles. Codility published analysis on how AI reshapes technical screening. And a startup called Rounds is marketing itself with the tagline that its interviews are ones "AI can't do," which is either contrarian or wishful thinking depending on who you ask.
Saffron's comparison materials emphasize two differentiators: candidates work on the company's actual codebase, not sanitized sandbox exercises, and the platform offers visibility into AI usage that competitors allegedly don't match. Whether that's defensible differentiation or just early positioning remains to be seen. The truth is, the market is moving fast enough that today's unique feature becomes tomorrow's table stakes.
The Open Question

For now, Saffron represents a specific hypothesis about where technical hiring is headed. Not away from AI, but deeper into understanding how engineers use it. The founders have correctly identified a real tension: if the vast majority of working engineers are using AI daily, excluding it from interviews starts to feel less like rigor and more like theater.
But there's a counterargument, too. Some engineering leaders worry that over-indexing on AI fluency risks hiring for prompt crafting rather than systems thinking. Others question whether line-by-line attribution really captures the collaborative, iterative nature of writing software with an assistant. And plenty of hiring managers, frankly, might not want that much data—they just want to know if the person can solve the problem.
Saffron is betting they do. That companies will pay for forensic visibility into how candidates think with AI, not just in spite of it. It's a bet that feels both inevitable and slightly unsettling, which might be exactly the right read on where hiring is headed.
