The coding interview—that ritual of whiteboard algorithms and contrived puzzles—has been dying a slow death for years. What's replacing it, though, remains up for grabs.
Litmus, a four-person New York City startup, thinks it has an answer. And it's simpler than you might expect: just let candidates work the way actual engineers work. That means their own IDE, their preferred terminal setup, and yes, full access to Claude and Copilot and whatever other AI assistant they've wired into their workflow.
The company, which came through Y Combinator's latest cohort, doesn't ask engineers to prove they can invert a binary tree. Instead, it generates technical assessments pulled directly from a hiring company's own repositories, internal tickets, and job descriptions. Candidates build end-to-end features, not isolated functions. Litmus captures the prompts they send to AI coding tools, runs the submissions in a sandbox, and grades them on process as much as outcome.
It's a bet that technical hiring should look less like a SAT for programmers and more like, well, programming.
The company reports it has already processed over 3,000 assessments and reached $60,000 in annual recurring revenue as of July 2026. Early customers include Neo Scholars and a handful of portfolio companies—though independent case studies haven't surfaced publicly yet.
Your Codebase as the Test
Most coding platforms rely on standardized problem banks. Litmus flips that model. Engineering teams connect their repositories, link up internal tickets, and feed in job descriptions. The platform analyzes the codebase and suggests a multi-stage interview pipeline tailored to how that specific team actually ships software.
The resulting assessments are scoped as feature work, not theoretical exercises. A candidate might fix a bug in a real module or extend an existing API endpoint—tasks that mirror what they'd encounter on day one.
Litmus integrates with applicant tracking systems like Ashby, Greenhouse, and Lever, and offers a REST API for teams that want programmatic control. Documentation suggests the system handles cohorts of five to ten candidates daily, with polling recommended to check submission status. Webhooks aren't available yet according to the documentation—a notable gap for teams running high-volume pipelines.
The Messy Reality of How Code Gets Written
Here's where Litmus diverges from the browser-based assessment tools that still dominate the market. Candidates work in their own environment. Not a stripped-down web IDE with limited tooling. Their actual setup.
And crucially, AI usage isn't just permitted—it's encouraged. The platform explicitly captures the prompts engineers send to tools like Claude Code and Copilot CLI. After submitting their code (in an interface that mimics opening a pull request), candidates record a walkthrough explaining their design decisions. Every submission runs in a sandboxed environment, graded on tool fluency, process, and results based on what the hiring company cares about.
It's an acknowledgment of something the industry has been slow to articulate: coding in 2026 looks nothing like it did five years ago.
Co-founder Elena Zhao, a software engineer who spent time at Two Sigma and Meta, announced the launch on LinkedIn in what appears to be early July. Her co-founder Shaivi Rau, the CEO, came to Litmus from Morgan Stanley, a few early-stage startups, and a stint in venture capital. Rau's background—computer science and film at Columbia—suggests someone comfortable bridging technical depth and narrative clarity, perhaps an edge in explaining why assessments should evolve.
A Crowded Field, Shifting Fast

Litmus is hardly alone in rethinking technical hiring. The market has been in flux for a while now, and the acceleration is hard to miss.
CodeSignal introduced what it branded "agentic coding assessments" this past April, designed explicitly to measure how engineers collaborate with AI rather than solve problems in isolation. HackerRank rolled out AI add-ons and desktop applications in January, adding identity-matching features to flag discrepancies between take-home assignments and live interview performance—a tacit admission that integrity concerns have spiked alongside AI adoption.
The pattern is unmistakable: platforms are moving away from classic algorithm screens and toward evaluations that reflect the messy, tool-augmented reality of modern software development. CodeSignal reported that roughly a third of its customers adopted AI-assisted formats last year. It also flagged a noticeable uptick in cheating attempts caught by detection systems—an unsurprising byproduct of the shift.
Litmus's angle is different, at least in emphasis. Rather than adapting generic tasks to accommodate AI usage, it generates assessments from the hiring company's actual codebase. The hypothesis: engineers perform differently on authentic work samples from a company's stack than on standardized scenarios, even when both permit AI assistance.
Whether that thesis holds up remains an empirical question. Early customer logos on the site include August, Spur, Autumn AI, Pine Labs, and Delphi—names that suggest traction in the startup ecosystem, if not yet proof of broader adoption.
The startup also exposes a Model Context Protocol server, a technical detail that matters more than it might sound. Teams can connect Claude Desktop, Claude Code, Cursor, or custom agents to read pipeline state and pull submissions. Full assessment generation through the MCP endpoint is on the roadmap but not yet live. An audit log in the dashboard tracks all tool calls triggered through MCP, a nod to security-conscious engineering leaders who want visibility into what's happening under the hood.
Pricing isn't published. Litmus operates on a demo-led go-to-market strategy, which typically means customized contracts and enterprise-style sales cycles, even for relatively small teams.
The Wager

In a hiring market where AI coding assistants have become table stakes, Litmus is making a straightforward wager: companies want to see how candidates actually build in their environment, with their stack, using the tools the team uses every day.
It's a hypothesis that feels intuitively right—perhaps more intuitively right than it is empirically validated. But the early traction, modest as it is, suggests enough engineering leaders are willing to run the experiment. Whether Litmus proves more predictive than established alternatives will depend on data most startups in this space are reluctant to share: correlation between assessment performance and on-the-job success.
For now, the company is betting that authenticity beats standardization. In technical hiring, that's not a settled question. But it's increasingly the right one to ask.
