The pixel is dying. Maybe not literally—you're still reading this on a screen—but as the primary interface between product and user, it's looking increasingly mortal. A handful of companies have noticed something odd: their most active users don't browse anymore. They don't click. They reason.
San Francisco's Armature is among the first to build explicitly for this shift. The startup, which emerged from Y Combinator's Spring 2026 cohort and launched self-serve access in May, sells continuous testing for a world where your customer might be Claude, not someone using Chrome to access Claude.
The pitch is straightforward, if conceptually disorienting. If your command-line interface or Model Context Protocol server gets called thousands of times daily by agents running inside ChatGPT, Claude Code, Cursor, or Gemini, you need infrastructure that tells you when something breaks. Your Selenium tests won't catch it. Traditional monitoring tools flag server errors, not the moment an AI misinterprets your documentation and faceplants mid-workflow.
Armature says it will.
A Problem That Emerged Fast
The issue Armature targets didn't exist three years ago. It barely registered 18 months back. But as agents proliferated—particularly coding assistants hooked into developer workflows—something changed for companies shipping tools designed to be used by these systems.
When you publish an MCP server (now the dominant standard for exposing functionality to AI agents), you lose control of the user journey. There's no onboarding funnel, no carefully designed happy path. Instead, you're handing the wheel to reasoning systems that improvise, misinterpret vague prompts, and occasionally fail in ways no QA team anticipated. An API changes, documentation drifts, and suddenly agents can't complete tasks they handled yesterday.
Armature's platform runs live LLM agents through customers' MCPs and CLIs, testing across multiple harnesses and models on a regular cadence. That roster includes Claude Code, Codex, OpenClaw, ChatGPT, Gemini CLI, OpenCode, and Cursor. The automated runs aim to catch broken workflows before they reach production. The system also surfaces model "thinking" via reasoning traces and tool-call logs, offering analytics and competitive benchmarking.
Perhaps most ambitious—and still unshipped—is a planned auto-remediation MCP endpoint. The idea: let customers' internal coding agents autonomously fix agent experience issues based on run data. That feature remains on the roadmap.
The Founders
Theodore Otzenberger and Louis Scremin built Armature with the kind of résumés that suggest they've been living this particular headache. Otzenberger spent time at Palantir, known for infrastructure that operates at scale under tight security constraints. Scremin led AI automation at Joko, where he deployed production agents to 6 million users—early exposure to the operational chaos of agentic systems in the wild. The team numbered three as of May 2026, though typical for early-stage startups, details on the broader org remain scarce.
Their core thesis, stated plainly on the site: "The next user is an agent."
It's a framing that's gained traction recently, echoing posts from Netlify's architecture team and a growing community effort around something called "Agent Experience" as a design discipline. The term—AX, if you're into acronyms—caught fire in 2025 as companies realized optimizing for agents requires different principles than optimizing for humans. Different failure modes. Different success metrics.
How It Works

The setup flow reflects Armature's technical audience. Users sign up, connect an MCP URL or run a CLI in a sandbox, and the platform's agent discovers available tools, assembles a test plan, and defines evaluation criteria. Tests run in the background. Results appear in dashboards; alerts push to Slack or email. PagerDuty, incident.io, and SMS integrations are listed as "coming soon."
The platform supports any MCP transport—HTTP, SSE, or stdio—and "any CLI we can spawn in a sandbox," per the documentation. End-to-end missions might look like: "Deploy a simple app using my MCP which provides cloud services." Pass/fail tracking. Alerts when steps break. Between full runs, lightweight heartbeat checks monitor availability.
Pricing starts at $99 monthly (or $79 annually) for a Starter plan covering one MCP/CLI source, ten tool monitors checked up to every five minutes, and one end-to-end workflow tested daily. The Pro tier, at $199 per month, scales to ten sources, thirty monitors, and hourly workflow tests across four model-harness pairs. Enterprise pricing is custom. All plans come with a seven-day free trial, naturally.
Armature positions itself as distinct from agent observability tools like LangSmith or Braintrust, which help developers debug their own agents. Instead, it tests how external users' agents experience your product—a testing layer focused on how agents interact with your tools in the wild. Closer in spirit, perhaps, to Datadog Synthetics or Playwright—but designed for agentic workflows rather than click paths. The company's FAQ notes that deterministic testing misses the "off-script failures" agents introduce, which is a polite way of saying: agents are weird, and your unit tests won't save you.
Timing and Crowded Territory

The launch arrives amid a broader infrastructure buildout around agentic workflows. Honeycomb launched agent observability features on May 12, 2026. Collibra unveiled an AI Command Center on May 6. Kore.ai shipped an Agent Management Platform in March; New Relic rolled out an AI agent platform in February. CoreWeave announced a unified agentic AI platform on May 28. All position themselves around visibility, governance, or control for production AI systems.
The academic side provides some validation. An April 2026 arXiv study evaluated 177,436 agent tools tracked between November 2024 and February 2026, finding software development dominated agent tooling and MCP had emerged as the predominant standard. A separate March 2026 paper described SpecOps, a fully automated agent testing framework across GUIs, CLIs, and web interfaces—demonstrating both feasibility and cost profiles for this kind of infrastructure.
Open-source alternatives exist but remain fragmented. test-mcp appeared on Hacker News in January as a CLI for MCP server testing. mcp-agent-test hit PyPI last month with mocking and assertion features. ProofAgent offers a SaaS layer for testing and governance, though details on adoption remain thin. None have quite framed the problem as Armature does: continuous testing across multiple commercial harnesses, positioned as infrastructure for companies whose products are becoming de facto platforms for agents.
The Unanswered Questions

No customer logos are publicly listed on Armature's site as of late May 2026. No case studies detail how early adopters use the platform, which makes sense for a company weeks into public availability, fresh out of Y Combinator. The company hasn't disclosed funding beyond YC participation; third-party databases list figures flagged as AI-generated estimates, unconfirmed by Armature.
What's clearer is the strategic bet. Armature is wagering that a new category—call it AX infrastructure, agent experience tooling, whatever label sticks—will emerge as companies realize they're no longer shipping features for human users alone. The fragmentation across harnesses and models, the unpredictability of reasoning systems, the frequency of silent regressions when providers update their agents: these are operational problems, not thought experiments.
Whether Armature captures that category or gets absorbed into broader observability platforms depends on execution. And on whether "agent experience" proves durable as a concept or fades into the graveyard of overhyped enterprise buzzwords.
For now, the company is offering a specific answer to a specific question: how do you know your MCP server works the way Claude's agent expects it to?
The seven-day free trial suggests they think the problem resonates enough for teams to find out themselves. Whether it does—and whether those teams convert—will determine if Armature is early to a genuine shift or just early.
