Every few months brings another proclamation: AI agents have arrived. GitHub rolls out agent mode. Startups tout superhuman coding benchmarks. Enterprise software companies quietly rebrand their existing automation tools with the magic words "agentic AI."
Scratch beneath the marketing veneer, though, and a more complicated picture emerges. Most of these agents still stumble when confronted with the kind of work that unfolds over hours or days—tasks requiring coordination across multiple systems, real-time judgment calls, and verifiable results in environments that refuse to stay clean and predictable.
The industry term for this challenge is "long-horizon reliability," and it may be the defining technical problem of the next phase of AI development. Building an agent that can navigate an integrated development environment? That's been solved, more or less. Building one that can autonomously catch a production error flagged by Sentry, trace it through logs, coordinate fixes across three different services, push changes through continuous integration pipelines, and notify the appropriate team members on Slack—all without human intervention?
That remains elusive. And according to Gartner, the gap between promise and performance is about to claim casualties: the research firm predicts over 40% of agentic AI projects will be canceled by the end of 2027.
When Explosive Growth Meets Stubborn Reality
The numbers initially suggest an industry on fire. IDC forecasts worldwide AI spending will hit $632 billion by 2028, representing a 29% compound annual growth rate. Gartner reports that 40% of enterprise applications will feature task-specific AI agents by 2026, up from less than 5% this year. Salesforce's Agentic Enterprise Index documents agent creation surging 119% in the first half of 2025 alone, with agent-led customer service conversations growing 22-fold.
Yet when Gartner surveyed CIOs and IT leaders this past May, only 24% had actually deployed even a handful of agents. LangChain's State of AI report offers another revealing data point: 43% of organizations now send LangGraph traces—evidence of stateful, multi-turn agent interactions—which means 57% are still building simpler, single-call LLM applications. Call-and-response, in other words. Not autonomy.
The performance benchmarks tell an even more sobering story. Back in 2023, GPT-4 achieved 14.4% success on WebArena's multi-domain website tasks, compared to 78.2% for humans. Two years later, the best model on WebGames—a benchmark of self-contained browser challenges—hits 43.1% versus human performance of 95.7%. Improvement, certainly. But hardly the stuff of autonomous workflows.
Software engineering tasks reveal similar constraints. SWE-EVO, a benchmark specifically designed to test long-horizon, multi-file software evolution, shows even frontier models with sophisticated scaffolding achieving roughly 21% success. That compares to 65% on single-issue tasks in SWE-bench Verified. The takeaway? Agents can handle discrete problems reasonably well. String multiple problems together over time, and success rates collapse.
Why Long Tasks Break Agents
The difficulty isn't mysterious, exactly. It's a function of compounding error, context management, and the sheer volume of decision points that accumulate when a task stretches across systems and time zones.
Consider what should be a straightforward software engineering workflow: an error surfaces in production monitoring. An agent needs to parse the alert, trace it through logs spanning multiple services, identify the root cause, write and test a fix, navigate CI/CD systems, and notify the relevant teammates. Each step introduces opportunities for failure, and more critically, each step depends on the accuracy of the previous one.
OSWorld-Human, an efficiency analysis of desktop agents, found that agents take 1.4 to 2.7 times more steps than necessary to complete tasks, with planning and reflection phases dominating latency. Worse, step-time increases along the trajectory. Agents don't just take more steps—they get slower and less accurate the longer they work, like a marathon runner hitting the wall.
The academic community has been methodically building benchmarks to expose these limitations. OSWorld presents 369 real desktop and web tasks across operating systems. BrowserGym standardizes web-agent evaluation across multiple benchmarks. OS-Marathon, introduced this past January, specifically targets repetitive long-horizon workflows—the kind of tasks where agents should theoretically shine but often stumble in practice.
Then there's the security dimension, which compounds complexity. Agent-SafetyBench, encompassing 349 environments and 2,000 test cases across eight risk categories, found that none of 16 tested agents scored above 60%. Research continues to document persistent vulnerabilities: prompt injection, unsafe tool invocation, the thorny problem of implementing least-privilege access when agents operate across organizational boundaries. As Anthropic's Model Context Protocol gains adoption—enabling agents to connect to diverse data sources and tools—the attack surface expands proportionally.
The Industry Fights Back

The response is taking shape across multiple fronts, though perhaps not always in the ways you'd expect.
At the model level, reinforcement learning has emerged as the leading technique for teaching agents to handle extended tasks. DeepSeek-R1, detailed in a Nature paper last year, uses a pure-RL pipeline to incentivize emergent behaviors like self-reflection and verification. OpenAI's o3 models are "trained with large-scale reinforcement learning to reason using chain of thought," though Nvidia CEO Jensen Huang noted these reasoning models can require up to 100 times more compute than previous generations. Which creates its own set of economic questions.
RL is also being applied directly to agent tasks themselves. WebAgent-R1 demonstrated end-to-end multi-turn RL for web agents, boosting a Qwen-2.5-3B model from 6.1% to 33.9% success on WebArena-Lite, and a Llama-3.1-8B from 8.5% to 44.8%. Impressive gains, certainly. But scaling RL training requires high-fidelity environments that accurately simulate real-world complexity—and building those environments is expensive, tedious work.
Which is where a new cohort of infrastructure startups is carving out territory, focused on the unglamorous but essential layer beneath the models themselves.
Polymath Labs, emerging from Y Combinator's Winter 2026 batch, is building production-grade RL environments specifically for long-horizon software engineering tasks. Founded by Dylan Ma (formerly of Hume AI and AWS) and Naren Yenuganti (who worked at Plaid and Amazon), Polymath focuses on creating realistic training environments that go beyond IDE interactions to encompass multi-service changes, CI/CD integration, ticket management in systems like Sentry and Linear, cloud metrics and logs, and coordination via tools like Slack.
"We obsess over realism," Ma told me. "Train an agent on a simplified environment and you get an agent that works in simplified environments. The moment you hit production complexity, it falls apart."
Vibrant Labs, from Y Combinator's Winter 2024 cohort, takes a similar approach for browser and computer-use agents, building RL environments to benchmark and improve long-horizon capabilities.
Meanwhile, established players are moving quickly to shore up their agent offerings. GitHub introduced "agent mode" with autonomous coding capabilities and multi-model support, including o3-mini and Gemini 2.0 Flash. Refact.ai, an open-source project, recently reported 70.4% success on SWE-bench Verified with methodology emphasizing guardrails and strategic planning tools—a reminder that scaffolding and careful engineering around base models remains as important as the models themselves.
Anthropic's Claude 4.5 series shows strong performance on OSWorld benchmarks, while the company's Model Context Protocol is rapidly becoming a de facto standard for connecting agents to tools. The protocol's broad adoption—including by OpenAI as of last year—creates a common interface layer that could accelerate development. Enterprise platforms like Workato's Enterprise MCP are emerging to handle secure connections to business systems, addressing both technical and governance requirements that matter more than many startups initially appreciated.
What Comes Next (and Who Survives It)
The next 12 to 24 months will likely determine whether long-horizon agents can cross the threshold from promising demos to production reliability. Several trends are converging that could either accelerate progress or expose deeper limitations the industry hasn't fully reckoned with.
Benchmarks are evolving beyond single-task snapshots toward continuous, multi-session evaluation. OS-Marathon's focus on repetitive workflows and SWE-EVO's emphasis on multi-file evolution tasks represent attempts to measure what actually matters for enterprise adoption, not just what makes for impressive blog posts. NIST's AI Safety Institute, which signed agreements with OpenAI and Anthropic in August 2024 for pre- and post-release model evaluations, is building standardized testing that could become de facto market requirements for high-capability agents.
Regulatory pressure, meanwhile, is arriving faster than many companies anticipated. The EU AI Act's general-purpose AI transparency and risk management rules took effect this past August 2, with broader obligations following this August 2—meaning companies deploying agents in Europe need compliance strategies now, not in some theoretical future quarter. NIST's AI Risk Management Framework, while technically voluntary, is shaping how U.S. enterprises think about agent governance and auditability in ways that will likely become functionally mandatory.
Compute economics remain challenging, perhaps prohibitively so. The RL training that shows promise for long-horizon tasks is expensive. Industry observers note that inference now dominates demand as reasoning models multiply token usage—creating pressure on specialized training infrastructure. Companies that can amortize RL training costs across multiple customers or use cases may have a durable advantage. Those that can't may find themselves squeezed.
The fundamental question isn't whether agents will eventually handle long-horizon tasks reliably. The research trajectory and commercial incentives are aligned; they almost certainly will, given enough time and capital. The question is whether the current wave of startups and enterprise deployments will survive the gap between today's capabilities and tomorrow's requirements.
Gartner's prediction of 40% project cancellations by 2027 doesn't imply universal failure. It implies winners and losers, companies that correctly anticipated where the technical bottlenecks would emerge and companies that didn't.
What seems increasingly clear—talking to founders, researchers, and enterprise buyers—is that the layer between base models and production deployments represents durable value. The infrastructure for realistic training environments, robust evaluation, and secure orchestration. The boring middleware that nobody writes Medium posts about.
As one founder in the space put it to me recently: "Agents that work in a demo are easy. Agents that work on Tuesday morning when production is down and three teams need coordination—those are hard."
The companies building toward that second reality are probably the ones to watch. Assuming, of course, they survive the next two years.
