The press release landed with an odd sort of candor. Amazon Web Services, announcing in early May that its AI agents could now execute stablecoin payments through Bedrock, included a warning that most tech companies wouldn't dare put in writing: "The stakes are high—a misconfigured payment flow doesn't just produce a bad answer. It moves real money."
There it was. The quiet part, said out loud.
That admission captures the peculiar moment enterprises find themselves navigating in 2026. They're racing—some would say stumbling—to deploy thousands of AI agents while grappling with a reality most vendors won't acknowledge: these systems fail. Constantly. And those failures cost actual money.
Enter ReasonBlocks, a two-person startup emerging from Y Combinator's Spring 2026 batch with a pitch that sounds almost quaint in its directness: What if you could catch agent failures while they're happening, before the damage compounds? The San Francisco outfit claims to cut token costs by 52% and boost accuracy by 42% on coding benchmarks—not by conjuring better models, but by stopping agents from repeating mistakes they've already solved.
The company's own testing, it should be noted. Independent validation hasn't surfaced yet, which is typical at this stage but worth flagging.
Still, the timing feels right. Or desperate, depending on how you look at it.
Sprawl, Meet Crisis
Gartner dropped a number in late April that clarified just how fast this train is moving: Fortune 500 companies will leap from running fewer than 15 agents in 2025 to more than 150,000 by 2028. A thousand-fold increase in three years. The analysts didn't mince words, warning of "agent sprawl" and sketching out six steps to manage what amounts to an infrastructure crisis on the horizon.
The industry, it seems, is already deep in it. Seventeen percent of organizations have deployed agents, according to Gartner's recent survey data, while north of 60% plan to within two years. By year's end, 40% of enterprise applications are expected to feature task-specific AI agents—up from less than 5% in 2025.
That velocity has every major platform vendor scrambling to consolidate agent infrastructure. OpenAI launched its Frontier enterprise platform in February. Google rebranded Vertex AI as the "Gemini Enterprise Agent Platform" at Cloud Next in late April. Anthropic released finance-sector domain agents with specific skills on May 6. AWS announced Bedrock Managed Agents and OpenAI Codex integration on April 28. Microsoft's Copilot Studio and "Agent 365" bundles launched May 1, reportedly priced at $99 per user for the E7 tier.
These aren't incremental product tweaks. They represent a calculated bet that enterprises will centralize the entire lifecycle—building, running, governing, and paying for agents—into integrated platforms. Because managing agents piecemeal is proving, well, untenable.
The Failure Catalog
The research paints a consistent, somewhat grim picture. Agents loop. They make redundant tool calls. They lose context mid-task. They produce cascading errors from malformed outputs. A February 2026 paper analyzing agent failures identified 12 error categories spanning tool initialization, parameter handling, execution, and result interpretation. Communication-related failures dominated in multi-agent systems, according to separate research on cloud root-cause analysis agents.
The cost problem compounds as scale increases. Claude Sonnet 4.6, one of the strongest coding models available as of this spring, charges $3 per million input tokens and $15 per million output tokens. That price-performance frontier looks attractive—until an agent hits a loop, burns through token budgets without producing anything useful, and slams into runtime limits. ReasonBlocks claims its customers see 70% fewer budget cap-hits using its monitoring stack, though again, these are figures from the company's own reporting.
Perhaps more revealing: agents are migrating beyond text generation into domains where errors have immediate, tangible consequences. AWS's AgentCore Payments integration with Coinbase and Stripe, highlighted on May 8, enables end-to-end agent payment flows using stablecoins. Misconfigured agents can now execute financial transactions. That reality is driving demand for runtime controls that go beyond after-the-fact observability.
And then there's regulation. EU AI Act obligations for general-purpose AI systems began phasing in during 2025, with high-risk category requirements intensifying through 2026 and 2027. Penalties reach up to €35 million or 7% of revenue for prohibited practices—the kind of figures that get CTO attention. NIST released a concept note on April 7 for developing a Profile for Trustworthy AI in Critical Infrastructure. Agent runtimes need audit trails, policy enforcement, and tamper-evident logging. Not as nice-to-haves. As compliance requirements.
The Two-Person Fix

ReasonBlocks positions itself as infrastructure that sits between agents and the models they call. Founders Sajeev Magesh, formerly of Stanford's computer science program, and Rohan Vij, who came out of Carnegie Mellon with research stints at ENGIE and UC Davis, have been building together for 11 years—an eternity in startup terms, and perhaps a hint at staying power.
Their pitch centers on three mechanisms: mid-run monitors that catch failures before budgets blow, compression that strips irrelevant context, and what they call a private "reasoning library" that compounds learning across production runs.
The reported results come from testing on SWE-bench Pro, a tougher benchmark developed by Scale AI last October to address contamination concerns with the original SWE-bench. Using Claude Sonnet 4.6 across 75 problems, ReasonBlocks claims its monitor and compression stack delivered 42% higher accuracy with 52% fewer tokens. The company's documentation, updated within the past week, shows the Python SDK working with LangChain, OpenAI's Agents SDK, and Anthropic's Messages and Agent SDK. Server-side monitor endpoints return steering interventions. The system supports what ReasonBlocks calls "E-trace injection" with tiered patterns, FSM-based model routing, codebase memory, and token compression with early exit (currently LangChain-only).
Solid engineering, if the numbers hold up.
They're not alone in this space. Sentinel.AI, which surfaced in a March 13 report, offers observability and reliability features including circuit breakers, DAG tracing, rollback and replay, error-budget SLOs, and dead-letter queues. AgentGuard47, an open-source package updated in April, provides runtime budget kill switches and loop detection. LangGraph, part of the LangChain ecosystem, added durable execution, checkpoints, and human-in-the-loop primitives, with a MongoDB partnership announced March 31. OpenAI's updated Agents SDK, detailed in a mid-April blog post, includes sandboxed execution, run cancellation, and approval workflows.
The distinction ReasonBlocks emphasizes is timing. Post-hoc observability tools like LangSmith, Langfuse, and Arize Phoenix analyze what went wrong after the fact. Evaluation infrastructure from Braintrust, which reportedly raised $80 million at an $800 million valuation in February, focuses on systematic testing. ReasonBlocks argues the decisive intervention happens mid-run, before failures compound.
MongoDB's April 23 engineering blog put it more bluntly than most vendors would dare: "The LLM is the smallest part of an agentic platform. Infrastructure—memory, state, orchestration, observability, evals, security, data—determines success."
The Payment Protocol Wars

The AWS-Coinbase-Stripe collaboration represents more than a technical integration. It's the opening move in what looks like a battle over agent-to-agent payment protocols. Coinbase's x402 protocol (reviving HTTP 402 "Payment Required") competes with Stripe's Machine Payments Protocol. Early security research on x402, published May 12, identified potential attack surfaces—because of course it did.
Both protocols depend on runtime infrastructure that can enforce pre-action authorization, execute rollbacks, and maintain audit trails. Exactly the capabilities agent monitoring layers are racing to provide.
Anthropic's May 6 launch of finance-sector domain agents, each with skills, connectors, and subagents, shows where vendors expect this to land. Agents handling financial workflows, compliance reviews, and transaction approvals need deterministic controls and full auditability. The company updated its subscription policy May 13 to restore third-party agent SDK usage with dedicated "Agent SDK credits," a sign that pricing models are still very much in flux.
Standards are evolving through the Linux Foundation's Agentic AI Foundation, launched in December 2025 to house the Model Context Protocol, OpenAI's goose project, and Anthropic's AGENTS.md specification. The AAIF announced a global 2026 events program in March. But standardization hit turbulence in late April when security researchers disclosed critical RCE-class vulnerabilities in multiple MCP implementations. Patches varied across vendors.
A reminder, if one was needed, that neutrally governed standards don't automatically mean production-ready security.
Platform Consolidation vs. Best-of-Breed

ServiceNow's messaging at its Knowledge 2026 conference in early May captured the industry's current arc. The company claims its internal AI specialist resolves IT service desk cases "99% faster"—a vendor metric that invites skepticism but suggests the magnitude of change. Bell Canada reported a 25% improvement in customer response time using ServiceNow telecom agents at Mobile World Congress. ServiceNow expanded its AI Control Tower to discover, observe, govern, secure, and measure AI "deployed across any system in the enterprise," integrating with NVIDIA's "AI Factory" validated design.
That governance layer is where enterprise platform bets are converging. Accenture's April 27 announcement that it's rolling out Microsoft Copilot to its 743,000-person workforce—Microsoft called it "the largest Copilot deployment to date"—signals confidence in the platform model. PwC's January case study described testing and feedback loops built into each agent feature across 230,000 users in more than 100 countries.
But the infrastructure question remains unsettled, perhaps more than the vendors would like to admit.
Platform vendors want to own the full lifecycle, from building to running to governing agents. Startups like ReasonBlocks are betting that reliability concerns will create demand for stack-agnostic layers that plug into multiple frameworks. The parallel is obvious: enterprises didn't stop at vendor-provided observability when Datadog and New Relic emerged to span cloud providers. Agent infrastructure may follow the same playbook.
IDC's FutureScape 2026 report projects that by 2027, half of enterprises will use AI agents, with agentic systems approaching half of AI spend by 2029. Sequoia's January essay, titled "2026: This is AGI," noted that scaffolding around model limits—memory handoffs, compaction, runtime controls—determines whether agents work at scale.
ReasonBlocks is placing a bet that enterprises paying millions in token costs and dealing with agent sprawl will pay for mid-run interventions that compress context, route to cheaper models when appropriate, and build a compounding knowledge base from production traces. Whether a two-person team can execute that vision against platform giants and a dozen other infrastructure startups is another question entirely.
The SWE-bench Pro numbers are promising. The broader test is whether CTO organizations drowning in agent deployments will add another vendor to their stack—or consolidate into platforms that promise to handle reliability themselves, even if those platforms haven't proven they can deliver what MongoDB called "the infrastructure that makes agents work."
The agent infrastructure market is fragmenting and consolidating at the same time. Which means 2026 will be a year of sorting: which control layers become table stakes, which vendors survive, and whether the industry learns to manage thousands of agents before it needs to manage billions.
Spoiler: nobody really knows yet. But the stakes, as AWS noted with unusual candor, are high.
