Ben Cochran has spent the better part of two decades coaxing maximum performance out of silicon—first at NVIDIA, then AMD, eventually earning the title of Distinguished Engineer. So when he decided AI coding agents had a reliability problem, he approached it the way a chip architect would: not by tweaking the inputs, but by redesigning the system itself.
The issue, as Cochran saw it, wasn't that agents were incompetent. They were unpredictable. And no amount of clever prompting would fix that.
On May 12, 2026, he posted Statewright to Hacker News with a characteristically direct pitch: visual state machines that enforce tool access at the protocol layer, cutting out prompt engineering entirely. Within 24 hours, the post had pulled 112 upvotes and 51 comments, many from engineering leads who wanted to know exactly how the licensing worked and whether this thing could actually deliver.
Hard Constraints, Not Gentle Suggestions
Here's the fundamental architecture: Statewright sits between an AI agent and whatever tools it wants to use—file editors, bash commands, API calls—as an MCP gateway. When the agent reaches for a tool, the gateway checks the current state. Is that tool allowed right now? If not, the call gets blocked before execution. The agent receives a structured error. Simple, unglamorous, effective.
"Hard means tool calls blocked at protocol layer before the model sees them," Cochran wrote in the project's README. "Advisory means rules injected in context without enforcement."
It's the difference between asking someone politely not to touch the stove and removing the knobs.
The visual workflow builder lets developers drag states around a canvas, draw transitions between them, and specify which tools an agent can access during each phase. A typical bugfix workflow might look like this: read-only file browsing while planning, capped edits during implementation, test-only bash commands during validation. If tests fail, guards automatically loop the agent back to the implementation state. No negotiation.
Even when bash access is permitted, Statewright layers on additional filtering. File redirects (>, >>), destructive operations (rm, shred), in-place edits via sed -i—all blocked by default. Developers can provide an allowlist of command prefixes, but baseline safety holds regardless of what the agent tries.
The Benchmarks (and the Asterisks)

Cochran published early results from a five-task subset of SWE-bench showing a 5x improvement on coding tasks compared to baseline agents. But there's context: those numbers don't come from the full 2,294-instance gauntlet. The experiment harness isn't public yet, though Cochran promised in the Hacker News thread that it's on the way.
The platform also claims to prevent what Cochran calls "read-loop death spirals"—situations where an agent gets stuck endlessly reading the same files without making progress. By scoping the available tool space and enforcing max_iterations limits at state boundaries, Statewright forces decision points. Whether that holds up across diverse codebases and agent architectures remains an open question.
Installation Is Fast, Licensing Got Messy
Getting started takes about two minutes. Add the plugin from the marketplace, generate an API key (they're prefixed with sw_live_), run /statewright list to verify. The MCP gateway endpoint defaults to https://mcp.statewright.ai. Statewright works with Claude Code, Codex, opencode, Pi, Cursor, and any MCP client.
Developers can either build workflows visually or let agents generate them programmatically. The system exposes seven MCP tools—including statewright_create_workflow and statewright_transition—that allow agents to read the public schema and construct their own state machines from context. The free tier includes 200 transitions monthly. Paid tiers exist (Pro, Team, Enterprise), though Cochran noted that pricing structures are "likely to be in flux."
But the launch wasn't entirely smooth. The licensing triggered immediate pushback. Cochran had initially released portions of the codebase under FSL-1.1-ALv2, which converts to Apache 2.0 in May 2029. After community discussion on May 12-13, he pivoted, updating the repository to adopt the canonical FSL template and publishing a Patent Pledge covering US Provisional Patent Application 64/054,240.
The patent pledge explicitly permits non-commercial use, open-source projects, academic research, and single-team self-hosted internal deployments. Commercial competitors are excluded, with a defensive termination clause if you sue. The core engine runs Apache 2.0 and is written in Rust.
Perhaps more than Cochran expected, the licensing details became as much a topic of debate as the technical approach itself.
A Crowded Space, A Specific Bet

Statewright arrives in a landscape suddenly thick with reliability infrastructure for AI agents. LangGraph offers stateful orchestration with broader adoption. HaltState bills itself as a "control plane for autonomous AI agents." Statehouse provides state and memory engines. Each tackles a slice of the reliability puzzle.
Cochran's specific bet: enforcement belongs at the protocol layer, not in the prompt. It's an architecturally clean argument. Whether it proves more reliable than alternatives in production environments—across different models, tasks, and failure modes—will require real-world adoption and independent validation of those early benchmarks.
For now, developers building AI coding agents have a free tier available. They can test the approach themselves. And in a field where most reliability claims are still more aspiration than evidence, that might be the most honest pitch of all.
