Somewhere between the chatbot craze and the next wave of autonomous AI, the infrastructure broke.
Not in a spectacular, headline-grabbing way. More like a slow-motion realization among developers that the tools built for quick demos weren't holding up when agents needed to run for hours—or days—without human intervention. State management became a headache. Partial failures turned into silent corruption. And the trust boundaries everyone assumed would hold? They didn't, not when agents started browsing the open web or executing code in sandboxes that weren't quite as isolated as advertised.
The numbers, at least, tell a story that sounds like progress. Agentic usage of OpenAI's Codex surged more than fivefold in the first half of 2026, according to company data. By August 2025, Gartner projected that 40 percent of enterprise applications would incorporate task-specific AI agents by 2026—up from less than 5 percent in 2025. Yet dig beneath the adoption frenzy and the picture gets murkier: only 17 percent of organizations have actually deployed agents in production, the firm noted, and roughly half of AI-driven digital initiatives are missing their ROI targets this year. The culprits? Collaboration gaps and data issues, the kind of unglamorous problems that don't make for good keynote slides.
Into this charged, uncertain environment steps OpenProse, a Y Combinator-backed startup positioning itself as something the industry didn't quite know it needed: "an open-source operating system for reliable long-running agents." The Spring 2026 batch company raised $1.25 million from 8-Bit Capital's Jonathan Abrams, Otis Chandler, and other angels. Its pitch is deceptively straightforward. Stop scripting agents, the founders say. Declare them instead.
The company's Markdown-first language lets developers describe desired outcomes as "contracts," while its Reactor runtime reconciles those declarations against reality, memoizing work so that—in the startup's framing—"cost scales with surprise, not wall-clock time." OpenProse claims north of 7,000 installs and reports an initial $10,000 pilot that expanded to $10,000 per month. The GitHub repository shows 1,600 stars and 122 forks as of late June, with a recent 0.15.0 release of the core "open-prose skill" on June 4. It runs as a plugin in compatible harnesses like Claude Code and OpenClaw, and the team emphasizes harness-agnosticism—a bet that developers will want portable infrastructure as the agent landscape fragments further.
Whether that bet pays off depends on whether the industry can solve a problem that's proven harder than anyone expected.
When Autonomy Gets Complicated
The shift from chatbots to autonomous agents introduces a new class of engineering challenge, one that's less about prompt design and more about the unsexy stuff of distributed systems. Unlike pass-at-one prompts, long-running agents must maintain state across hours or even days, recover gracefully from partial failures, and operate with minimal human oversight. A March 2026 academic framework titled "Beyond pass@1" proposed new reliability metrics—RDC, VAF, GDS, MOP—to capture these realities. Another paper in May introduced AgingBench, a longitudinal reliability test for agents that degrade after deployment, much like production software quietly accruing technical debt.
The problem, it turns out, is not theoretical.
In June, Microsoft disclosed "AutoJack," a remote code execution chain in AutoGen Studio triggered by malicious websites—a stark reminder that localhost trust boundaries break down when agents browse the open web. Researchers also flagged critical vulnerabilities in implementations of Anthropic's Model Context Protocol (MCP), the emerging standard for tool discovery and invocation. LangChain and LangGraph patched high-severity data exfiltration issues in March. Each incident underscored a shared weakness: agent infrastructure built for demos doesn't hold up under adversarial conditions or prolonged operation. Maybe it never could.
"Infrastructure becomes the orchestrator determining value capture," McKinsey wrote in an April analysis of agentic AI—a somewhat bloodless way of saying that the companies figuring out the runtime layer are the ones making money. The firm's data suggests that organizations building domain-based, end-to-end agentic workflows with robust runtime and operations layers are capturing ROI. Half-measures—what a Forrester report in May labeled "agentish" chatbots—remain stuck in what one executive described to me as "pilot purgatory," that liminal space where budgets get renewed but production deployments never quite materialize.
The Big Platforms Make Their Move
Every major cloud and AI vendor now offers some version of an agent runtime, though the quality and maturity vary wildly. OpenAI shipped sandboxed execution and snapshot-and-rehydration features in its Agents SDK in April, alongside WebSocket persistence improvements to its Responses API. Anthropic integrated the Claude Agent SDK into Apple's Xcode 26.3 in February, enabling longer autonomous coding tasks with MCP access. Google's Gemini Enterprise Agent Platform and Agent Development Kit received documentation updates as recently as June 24, emphasizing enterprise-scale reliability—though whether that reliability exists in practice or just in marketing copy remains an open question.
AWS added bring-your-own file systems (S3, EFS) to its Bedrock AgentCore Runtime in May for long-running workflows, and published design patterns for persistent MCP servers in February. Microsoft used its Build 2026 conference to position enterprise "Autopilots" as long-running agents, announcing general availability of its Agent Framework and tighter convergence with Semantic Kernel and the Copilot ecosystem. Even Databricks, whose data lakehouse roots might seem tangential to agents, updated its Mosaic AI Agent Framework documentation in late May with templates for stateful execution and supervisor patterns for background tasks.
The Linux Foundation formalized governance around MCP and related standards in December 2025 by launching the Agentic AI Foundation. As of mid-2026, AAIF has hosted developer summits and continues to expand its scope beyond the original protocols—a recognition, perhaps more explicit than anyone intended, that interoperability will determine who wins the agent infrastructure layer. Standards battles, after all, are rarely about technical merit alone.
Open Source and the "Agent OS" Gold Rush

Outside the cloud hyperscalers, a different breed of tooling has emerged—scrappier, more opinionated, and in some cases more interesting. LangGraph, part of the LangChain ecosystem, went generally available in May 2025 and markets itself as low-level, stateful orchestration for long-running agents, with explicit support for retries, checkpointing, and human-in-the-loop workflows. CrewAI, a multi-agent orchestration framework, released version 1.14.5 in May 2026 and has cultivated a following among teams building collaborative agent systems.
Then there are the "agent OS" projects—some credible, others speculative, a few that smell like vaporware. OpenClaw, an open-source autonomous assistant, has built a sizable ecosystem and positions itself as an OS-style platform. Newer entrants like KruxOS advertise kernel-level sandboxing with cgroups and seccomp, while OpenSpry and CoWork OS target local-first, multi-LLM environments. The maturity and security posture of these systems varies widely, to put it charitably. Splendor calls itself an "agent kernel runtime." Steadybase makes enterprise agent OS claims but lacks independent verification of deployments.
Academic research has also begun to sketch out what an agent-native OS might look like—though whether any of it will leave the lab is another question. A June paper on "Agent libOS" proposed a library-OS-inspired runtime with async scheduling, object memory, and permission grants. Another group released AOHP, an Android-based harness built on AOSP with personalized service composition and secure information flow, intended as a testbed for OS-level agent concepts. A February paper on "AI Runtime Infrastructure" outlined execution-time optimization layers—adaptive memory, failure recovery, policy enforcement—explicitly designed for long-horizon agents.
OpenProse fits somewhere in this crowded, occasionally bewildering spectrum. It doesn't claim kernel-level sandboxing or device integration. What it offers instead is authoring simplicity. Developers write contracts in portable Markdown; the Reactor runtime handles execution with content-addressed receipts that provide auditability and introspection. Founder Dan Barrett, in a May interview with Kent Beck, positioned Prose as "a real programming language for AI sessions," emphasizing practicality over ambitious architecture. The MIT license and plugin ecosystem signal a bet on composability rather than vertical integration—a choice that could prove wise or limiting, depending on whether the market consolidates around a few dominant platforms or fragments further.
When Enterprise Money Shows Up
ServiceNow's partnership with Anthropic, announced in January, made Claude the default model for the company's Build Agent, an agentic application and workflow builder. The goal: shorten time-to-value while maintaining governance across critical industries like healthcare and finance, where mistakes carry regulatory consequences. Salesforce reported more than 12,000 Agentforce customers by early 2026, with particular traction in contact center and IT service management use cases. By February, 180 organizations had replaced legacy support tools with Agentforce IT Service, according to a company investor relations release—though how many of those deployments are truly agentic versus glorified automation remains unclear.
These deployments, at least the successful ones, share a pattern. They succeed when governance, observability, and failure recovery are baked in from the start. They falter—sometimes spectacularly—when agents are bolted onto existing systems without rethinking state management, audit trails, or approval workflows. The EU AI Act's phased enforcement, with general-purpose AI model obligations taking effect around August 2026 and high-risk system deadlines extending into 2027, has added regulatory urgency. Enterprises must now map AI Act applicability to agent workflows—long-running, tool-using, code-executing—and implement required controls before the next wave of deadlines hits.
NIST released a concept note in April for an AI Risk Management Framework profile on trustworthy AI in critical infrastructure, complementing its 2024 generative AI profile. A February draft addressed AI-specific cybersecurity concerns. The message from regulators and standards bodies is consistent: if you're deploying autonomous systems, you need auditability, containment, and explainability. And you need them before something goes wrong, not after.
Gartner's June Hype Cycle for Agentic AI noted that more than 60 percent of organizations expect to deploy agents within two years, but operational readiness remains a gating factor—a polite way of saying most companies have no idea what they're getting into. Forrester's State of Agentic AI report, published in May, found that many so-called agent deployments are little more than chatbots with extra steps. TechRadar framed 2026 as "the year agents become trusted co-workers," but only where governance and observability are strong. Without those foundations, pilots get canceled and budgets evaporate, usually quietly.
Patterns Emerging From the Chaos

The agent infrastructure market, if you squint, is consolidating around a few key patterns—though whether these represent genuine convergence or just the industry's current best guess is hard to say. Declarative state graphs (LangGraph, OpenProse contracts) reduce fragility by making execution paths explicit. Checkpointing and resumability (OpenAI's snapshot APIs, Databricks' long-task continuation) let agents survive interruptions without losing hours of work. Content-addressed receipts and artifact-first logging (OpenProse's ledger, AWS's bring-your-own file systems) provide the auditability enterprises demand, or claim to demand. MCP is emerging as the de facto standard for tool integration, despite security growing pains that suggest the protocol wasn't designed with adversarial use cases in mind.
Harness-agnosticism matters more than it might seem at first glance. Enterprises will run agents across cloud vendors, on-prem data centers, and developer laptops. Proprietary SDKs create lock-in and integration tax. Open-source runtimes with portable contracts reduce switching costs and enable hybrid deployments—at least in theory. OpenProse's Reactor, with its emphasis on surprise-driven cost and memoization, represents one approach to this problem. Whether it can compete with cloud-native offerings that bundle compute, security, and compliance remains an open question, one the founders will need to answer as the company scales.
The broader narrative is harder to ignore, even for skeptics. McKinsey's April analysis suggested that winners in the agentic era will build domain-based, end-to-end workflows with robust runtime layers—a framing that conveniently positions McKinsey to sell consulting services, but isn't necessarily wrong for that. IDC projects global AI IT spend of roughly $409 billion in 2026, climbing to $700 billion by 2029, yet warns that collaboration and data issues will sink half of this year's initiatives. OpenAI's internal data shows that agentic usage is growing fastest outside traditional software developers—business analysts, marketers, operations teams, the kind of users who have no patience for brittle infrastructure. Fortune 500 companies may operate more than 150,000 agents in production by 2028, according to a TechRadar analysis citing Gartner projections.
If those projections hold—and that's a big if—the current scramble to build agent infrastructure will look, in hindsight, like the early days of containerization or serverless. Standards will harden. Security will improve, or the industry will suffer a reckoning that makes today's AutoJack disclosures look quaint. Vendors promising "agent OS" capabilities will either ship production-grade systems or fade into the landscape of overpromised demos, forgotten except by the VCs who funded them.
For now, the gap between aspiration and execution remains uncomfortably wide. Agents are here, no question. But the operating systems for agents? Those are still being written—one declarative contract, one memoized receipt, one security patch at a time.
