The promise of AI coding agents sounded almost too good: software that writes itself, debugs itself, fixes itself. Then engineering teams started using them in earnest, and a pattern emerged. On straightforward tasks, the agents performed admirably. But hand them a gnarly, multi-threaded memory corruption bug—the kind that haunts senior engineers at 3 a.m.—and watch the wheels come off.
The agents hallucinate. They chase phantom root causes. They burn through API tokens while proposing fixes that, somehow, make everything worse.
Undo, a Cambridge-based company with roots in academic computer science, thinks the problem isn't the agents themselves. It's what they're being asked to work with: incomplete information, fragmented logs, source code that tells you what should happen but not what actually did.
On June 15, 2026, the company announced a $37 million growth investment led by Elsewhere Partners. The pitch? That deterministic runtime recording—technology that captures every instruction, variable mutation, and system call during program execution—is the missing ingredient for AI-assisted debugging that doesn't spiral into hallucination. Rod Favaron from Elsewhere Partners has joined as Executive Chair, a signal that investors see potential in marrying old-school debugging rigor with new-school AI tooling.
What Complete Context Actually Means
Most AI coding agents today work a bit like detectives arriving at a crime scene after everyone's gone home. They get the source code. Maybe some logs. Perhaps a stack trace if they're lucky. Then they're expected to reconstruct what happened.
Undo's LiveRecorder takes a different approach: it records the crime as it unfolds. During program execution, the software captures everything—every instruction, variable change, thread event, system call. The result is a deterministic recording that can be replayed identically, rewound, fast-forwarded, queried at any point in time.
Where things get interesting is the integration with AI agents. The recordings plug into Anthropic's Model Context Protocol, an open standard that works across Claude Code, GitHub Copilot, Cursor, and other MCP-compatible tools. Instead of an agent spinning its wheels trying to reproduce a bug through trial and error, it can interrogate the recording directly. "When did this variable last change?" "What was the call stack 200 milliseconds before the crash?"
The technical implementation relies on dynamic JIT binary translation to capture non-determinism and precise execution timing. Recordings are portable, self-contained, analyzable offline. The company supports Linux on x86, x86_64, and ARM64. GPU code execution itself can't be recorded—physics, apparently, still imposes limits—though CPU-side GPU interactions can be.
The Performance Question

There's no free lunch in systems programming. Runtime recording means overhead.
Undo's own benchmarks suggest a 1.5× to 5× per-thread slowdown depending on workload characteristics. Sqlite clocks in at 2.9× overall. Gzip hits 1.5× per-thread. Ffmpeg shows 5.6× overall but only 1.4× per-thread, a disparity that speaks to how differently concurrent applications behave under recording. Memory footprint typically doubles compared to the recorded process, leveraging OS copy-on-write semantics to minimize the damage.
For enterprise teams, though, the calculus isn't purely about CPU cycles. It's about time to resolution. Undo claims—based on its own benchmarks, worth noting—that agents solve 92% of complex bugs when given runtime context, up from 38% without it. The company also reports customers seeing up to 100× faster root-cause analysis, though independent validation of those figures remains limited.
In January 2026, Undo published an experiment debugging a real GDB crash. Using Claude Code and Codex CLI, the agent nailed it on the first attempt when working with an Undo recording. Attempts using only source code and compile-run cycles failed repeatedly and consumed more tokens. There's an economic argument buried in there: grounded agents that actually understand what happened solve problems more efficiently than agents that hallucinate their way through the problem space.
Who's Actually Using This

The customer roster tilts heavily toward companies building extraordinarily complex systems. Enterprise users named by the company include SAP, Siemens EDA, Synopsys, Cadence, AMD, and Cisco—organizations where bugs don't just annoy users but cost millions in engineering time and delayed product launches.
Palo Alto Networks' SVP of Engineering appears in the funding announcement discussing automated root-cause analysis and speed improvements. What the company hasn't published, at least not yet, are detailed third-party case studies showing AI agent deployments with Undo in production environments beyond their own examples.
The company recently released version 10.0 of its UDB debugger with MCP integration baked in. The Undo AI feature is currently in tech preview—a polite way of saying it's not yet battle-tested at scale. A VS Code extension can auto-configure GitHub Copilot to use UDB as an MCP server, though you'll need VS Code 1.101 or later, extension version 2.1.12 or higher, UDB 10.0 or above, and MCP enabled in Copilot. For Java and Kotlin developers, there's a separate lr4j_mcp server handling mixed Java/native debugging scenarios.
The Timing Is Everything

Undo's funding arrives during a year when AI agent reliability has become something of an obsession in engineering circles, with recent studies documenting the challenges of long-context agent debugging. Not always for good reasons.
In April and May, stories made the rounds about AI agents wiping databases, purging production code, generally wreaking havoc when left unsupervised. The incidents felt almost designed to validate what academic researchers have been documenting: long-context agent debugging without structured runtime information is, charitably, unreliable.
The deterministic debugging space is getting busier, though competitors address different layers of the problem. Replay.io focuses on browser and web stack time-travel debugging—JavaScript land, essentially. Antithesis builds deterministic verification environments where agents can iteratively correct their own code. Agent observability frameworks like Laminar, which raised a $3 million seed round in March, track LLM and tool-call traces but don't capture full native runtime execution.
Undo's approach is narrower in one sense. It's Linux-only, on-premise software aimed squarely at teams debugging compiled code. No cloud service, no browser support, no grand ambitions to solve every debugging problem for every platform.
But it goes deeper on execution fidelity than most alternatives. For engineering leaders evaluating AI coding agents on sprawling C++, Rust, or Go codebases—the kind where a single memory leak can take days to isolate—the value proposition is almost brutally simple.
Give the agent the complete picture, or watch it hallucinate. The choice, perhaps, was always that stark.
