It's 2 AM. An AI coding agent has just spun up a new button layout, run the tests, and merged the changes into your main branch. Everything compiles. The unit tests pass. And yet—somewhere in that automated workflow—the checkout flow has shifted three pixels to the left, just enough to break mobile payments.
You won't know until a customer complains.
This is the paradox at the heart of AI-assisted development: the tools writing our code can't actually see what they're building. GitHub Copilot can autocomplete an entire React component. Cursor can refactor your CSS. Claude can debug your styles. But none of them render a browser window to check their work. They're flying blind, at least when it comes to the visual layer.
Now, a small but growing number of developers think they've found an answer—or at least the beginning of one.
The Thing Agents Can't Do
The problem isn't that AI writes bad code, exactly. Industry observers have noted that coding agents "struggle with front-end" work, but the issue is more specific than that: they lack visual feedback loops. An agent can generate pixel-perfect markup. It just can't verify that the button it styled actually renders the way a human would expect, especially across different screen sizes or when dynamic content starts rearranging the DOM.
Conversations in developer communities over recent months paint a familiar picture. Agents excel at the mechanical parts—generating markup, writing CSS classes, even handling state logic. Then they fail, quietly, when an animation causes a timing issue or a media query doesn't trigger as expected. One developer put it plainly in an online discussion: agents need "mandatory verification after every action" to catch misclicks and UI drift before code ships.
The phrase "trust but verify" gets used a lot in software. With agents, though, the verification step has been missing. At least for the visual stuff.
Two Projects, One Problem

In recent weeks—perhaps sensing the same gap—two open-source projects have landed, both tackling visual regression testing for agent-generated code. Both released under MIT licenses. Both designed to run inside CI pipelines, no cloud dependencies required.
SnapDrift arrived first, in early March. Version 0.1.0 does what the name suggests: it captures full-page screenshots during test runs, compares them pixel-by-pixel against stored baselines, and surfaces any drift directly in pull request comments. The tool is branch-aware, meaning teams can scope baselines to feature work without cluttering the main branch. Notably, screenshots never leave the build environment—a design choice that matters for teams in healthcare, finance, or anywhere else data residency isn't negotiable.
lasTest, from a developer going by Dexilion, takes a slightly different tack. It's a command-line tool that generates Playwright-based visual tests on the fly, then compares live environments to development branches across multiple viewports. The results show up in tabbed HTML reports. There's an AI-assisted diffing mode alongside traditional pixel comparison, and the tool includes stabilization features—freezing timestamps and animations before capture—to cut down on false positives.
Both tools share a philosophy, if not an implementation. Local-first. PR-integrated. Built for the workflows developers already use, assuming you're running Playwright or something similar for browser automation.
Neither has cracked 1,000 GitHub stars yet. But that might not be the point.
Why This, Why Now
The timing isn't random. GitHub unveiled its coding agent features and "Agent HQ" dashboard late last year, and by early this year had published formal documentation on agentic workflows—complete with sections on security guardrails and human approval gates. The infrastructure for autonomous agents inside CI exists now. What's been missing is the last-mile verification for anything visual.
To be fair, visual regression testing isn't new. Percy, Applitools Eyes, Chromatic—these are mature, commercial products with sophisticated diffing algorithms, some powered by their own AI models. But they're SaaS platforms. Screenshots get uploaded to third-party servers, analyzed in the cloud, stored in someone else's database. For a regulated startup or a security-conscious enterprise, that's often a dealbreaker.
The new wave of open-source tools emerged specifically to sidestep that constraint. Full verification capability, zero cloud storage, MIT-licensed code you can audit yourself if you're so inclined.
Some developers aren't stopping there. A handful have started building "skills" for coding assistants—plugins that hook visual verification directly into Cursor or Claude-based IDEs. One example, listed on a community skill repository, lets agents request screenshot comparisons mid-task, before they even propose a code change. It's early, mostly experimental, maintained by volunteers. But it hints at where this could go: agents that check their own UI work, in real time, as part of the coding loop.
Whether that's reassuring or unsettling probably depends on how much you trust the agent in the first place.
What's Not Solved Yet

Neither SnapDrift nor lasTest has hit mainstream adoption—both are weeks old, not months. Repository activity looks promising in that scrappy, early-stage way, but we're talking dozens of stars and a few contributors, not thousands. The tools solve the technical problem: local diffing, PR integration, CI compatibility. What they haven't tackled yet is the workflow problem.
How do teams set useful baselines without someone manually reviewing every screen? How do you distinguish between an intentional design change and an actual regression? When does a three-pixel shift matter, and when is it noise?
Those questions get thornier as agents write more of the codebase. For the moment, the tooling exists. Whether it scales beyond early adopters depends on two things: how quickly agent-native development becomes standard practice, and whether teams start trusting visual diffs enough to actually gate merges on them.
That second part—the trust part—might take a while.
