Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 5, 2026

AI Agents Are Running Biology Labs: The Self-Driving Lab Revolution

AI Agents Are Running Biology Labs: The Self-Driving Lab Revolution
Artificial IntelligenceBiotech+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 5, 2026

The Race to Build Infrastructure for Human Digital Twins

The Race to Build Infrastructure for Human Digital Twins
Digital TwinsBiotech+3

Founders Mentioned

Dylan Ma

Polymath Labs

saas icon
SaaS

Naren Yenuganti

Polymath Labs

saas icon
SaaS

Dylan Ma

Polymath Labs

saas icon
SaaS

Naren Yenuganti

Polymath Labs

saas icon
SaaS
SaaS iconSaaS
February 5, 2026
YcArtificial IntelligenceB2b SaasDeveloper Tools

The Agent Reliability Gap: Why Long-Horizon AI Remains Unsolved

As enterprises rush to deploy AI agents, a YC-backed lab is tackling the hardest problem: making autonomous systems reliable for complex, multi-day tasks spanning real production tools.

The Agent Reliability Gap: Why Long-Horizon AI Remains Unsolved

When McKinsey announced it would deploy 25,000 AI agents—more than half the size of its human consulting workforce—the January 2026 declaration landed like a gauntlet. Here was one of the world's most influential management firms betting its reputation on autonomous software. The future, McKinsey seemed to say, had arrived.

Dan Priest didn't buy it.

The Chief AI Officer at PwC, McKinsey's arch-rival, responded with barely concealed disdain. Agent counts were vanity metrics, he argued publicly. What mattered was whether the technology actually worked: task mastery, adoption, oversight. The subtext was unmistakable—anyone can claim to deploy thousands of agents. Making them useful is another matter entirely.

That tension, it turns out, captures something fundamental about the moment enterprises find themselves in. Nearly every large company is racing toward AI agents—software that can work autonomously across systems, completing complex tasks without constant human intervention. Cloudera's survey of roughly 1,500 IT leaders across 14 countries found that 96% plan to expand agent use within a year. Salesforce reported 119% growth in agents created during the first half of 2025, with agent-led service conversations climbing about 70% month-over-month from January through June.

The momentum is undeniable. Whether it's sustainable is another question.

Scratch beneath the deployment numbers and a more complicated picture emerges—one that should give even the most enthusiastic executives pause. TechRepublic discovered in January that despite all the activity, only 8.6% of enterprises actually have agents running in production. That figure had doubled in late 2025, which sounds impressive until you consider what Gartner warned six months earlier: more than 40% of agentic AI projects could be canceled by 2027 due to complexity and poor fit with business needs.

The problem isn't that AI agents can't do anything useful. In narrow, well-defined domains, they perform admirably. Klarna's AI assistant, launched in February 2024, handled the equivalent workload of 700 to 853 customer service agents in its first month, managing two-thirds of customer chats. Those numbers are real.

But customer service chatbots are the easy part.

When Simple Tasks Become Complex Journeys

The real test comes when work extends beyond a single session or tool—when tasks unfold over hours or days, coordinating across multiple services and adapting to shifting conditions. This is what researchers call the "long-horizon" problem, and it remains largely unsolved.

Consider a software engineer's actual workday. They don't just write code in an IDE. They push changes across multiple microservices, configure CI/CD pipelines, triage alerts in Sentry, coordinate through Slack, analyze logs in Google Cloud Platform. Each step depends on the previous one. Context matters. Timing matters. A mistake in step three can cascade through everything that follows.

Current AI agents struggle mightily with this kind of complexity. The benchmarks tell a stark story, perhaps starker than most companies deploying these systems realize. SWE-bench Pro, released in September 2025 to measure performance on long-horizon enterprise software engineering tasks, shows frontier models achieving less than 25% success rates. SWE-EVO, introduced in December to test multi-file evolution tasks, found GPT-5 combined with OpenHands reaching approximately 21% accuracy.

These aren't incremental gaps. They represent a fundamental barrier to autonomous operation.

OSWorld, a benchmark for computer-use agents, showed improvement from roughly 5% success rates in 2024 to vendor-reported ranges of 40% to 70% in 2025. Progress, certainly. But even these gains come with significant caveats—tool utilization remains low, and grounding in complex multi-step workflows proves brittle. Some researchers, in papers like "The SWE-Bench Illusion," have questioned whether reported progress reflects genuine capability or clever optimization for narrow test cases.

LangChain's State of Agent Engineering report from January 2026 surfaced an uncomfortable disconnect: while 57% of the 1,300-plus respondents claimed to have agents in production, only 52% ran offline evaluations. The enthusiasm for deployment appears to be outrunning the discipline of rigorous testing. Which perhaps explains why Gartner simultaneously urges enterprises to establish agentic strategies within three to six months while warning of massive project cancellations ahead.

The industry has reached what Forrester calls the "hard-hat work" phase of AI adoption—the unglamorous slog of making things actually work. After the hype of 2024 and early 2025, enterprises are discovering that deploying reliable autonomous agents is fundamentally harder than anticipated. Forrester predicts that 25% of planned AI spending will slip into 2027 as organizations grapple with return-on-investment scrutiny and the operational realities of agent governance.

Building Better Training Grounds

The technical root cause is becoming clearer, even if the solution isn't. Current training and evaluation approaches don't adequately simulate realistic production conditions. Agents trained on synthetic data or simplified scenarios fail when confronted with the temporal dependencies, tool orchestration, and outcome verification requirements of actual work.

The industry is responding through two parallel efforts that may or may not converge: platformization by the tech giants and specialized training infrastructure from smaller, more focused players.

On the platform side, the major cloud providers launched comprehensive agent management systems throughout 2025 and early 2026. OpenAI unveiled its Frontier enterprise platform in February, bringing in early customers including Intuit, State Farm, Thermo Fisher, and Uber. Google's Vertex AI Agent Builder reached general availability in March 2025. AWS introduced AgentCore for Bedrock with policy controls in December.

These platforms converge around similar building blocks: memory management, evaluation frameworks, observability, governance. Microsoft's Ignite 2025 announcements included multi-agent workflows and Model Context Protocol integration for Copilot. In December, Anthropic donated the Model Context Protocol to the newly formed Agentic AI Foundation under the Linux Foundation, with backing from across the hyperscalers. Over 10,000 public MCP servers have been adopted across ChatGPT, Gemini, Copilot, and VS Code.

Meanwhile, a different bet is emerging from applied research labs. San Francisco-based Polymath Labs, part of Y Combinator's Winter 2026 batch, is building what it describes as production-grade reinforcement learning environments for long-horizon tasks. The company's founders—Dylan Ma, previously at Hume AI and AWS, and Naren Yenuganti, previously at Plaid and Amazon, both UC Berkeley alumni—are starting with software engineering across the full development lifecycle.

Their thesis centers on a specific gap: existing benchmarks and training environments don't capture the temporal unfolding and dependency enforcement of real production systems. Polymath is building RL environments that behave like actual production tooling—spanning multi-service code changes, CI/CD systems, error tracking platforms like Sentry, project management in Linear, cloud metrics in GCP, team coordination through Slack. The company has partnerships with what it calls "frontier labs" to customize and scale these environments.

This represents a departure from the prevailing approach of wrapping better prompts and orchestration around existing foundation models. Instead, it treats agent reliability as fundamentally a training data problem—one requiring environments sophisticated enough to teach complex, multi-day workflows.

The timing aligns with renewed interest in reinforcement learning for AI agents. A February 2025 paper titled "LOOP: Reinforcement Learning for Long-Horizon Interactive LLM Agents" demonstrated that training agents directly in target digital environments outperformed larger baseline models. OpenAI's launch of reinforcement fine-tuning for reasoning models provides another data point; the company published a case study in October 2025 of Doppel using RFT with GPT-5 to reduce analyst workload by 80%. Google and AWS have introduced their own reinforcement learning capabilities.

The academic work reinforces what practitioners are discovering: long-horizon tasks require agents that can plan over extended time horizons, maintain context across tool boundaries, and learn from verifiable outcomes rather than next-token prediction alone.

The Compliance Reckoning

Digital illustration for article section "The Compliance Reckoning" in "The Agent Reliability Gap: Why Long-Horizon AI Remains Unsolved" - A professional and conceptual digital illustration representing the impending compliance reckoning f...

Running parallel to these technical developments, regulatory and compliance requirements are reshaping the economics of agent deployment in ways that most companies may not have fully absorbed.

The EU AI Act's major provisions take effect August 2, 2026—less than six months away—with obligations around high-risk AI systems, governance, and substantial penalties. Organizations placing general-purpose AI systems before that date must comply by August 2027. Standards bodies have moved quickly: ISO/IEC 42001:2023 provides an AI management system framework, while ISO/IEC 42005:2025 addresses AI system impact assessment.

In the United States, NIST released its AI Risk Management Framework for Generative AI in July 2024. OMB Memorandum M-24-10 from March 2024 mandated Chief AI Officers, inventories, and minimum risk practices for federal AI use.

Security remains a persistent thorn. Prompt injection attacks against browser agents continue to demonstrate vulnerability. ITPro reported in January 2026 on prompt injection risks in OpenAI's Atlas system, underscoring that these aren't theoretical problems. The Cloud Security Alliance published its Agentic AI Red Teaming Guide with updates through October 2025, while OWASP's Top-10 for LLM Applications includes "Excessive Agency" as a distinct risk category.

The regulatory and security environment creates both challenges and opportunities. Enterprises need auditable, observable agent systems with clear outcome verification—which happens to align with the kind of training and evaluation infrastructure required to achieve long-horizon reliability in the first place.

What Comes Next

Digital illustration for article section "What Comes Next" in "The Agent Reliability Gap: Why Long-Horizon AI Remains Unsolved" - A conceptual digital visualization representing the rapid expansion of enterprise AI ecosystems, fea...

Gartner predicts that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. By 2028, the firm expects one-third of agentic implementations to involve collaborating agents, with cross-application agent ecosystems becoming the norm.

The industry is moving from simple assistants toward what analysts call "Outcome-as-Agentic-Solution"—vendors delivering measurable business outcomes through autonomous agents rather than selling software licenses. This shift places enormous pressure on reliability. An agent that occasionally hallucinates in a chat interface is frustrating. An agent that corrupts production systems or leaks sensitive data during a multi-day workflow is existential.

Several factors will determine which approaches prevail, though the smart money would probably hedge across multiple bets. First, whether training-focused solutions like specialized RL environments prove more effective than inference-time orchestration and tool governance. The research suggests both matter, but the balance remains open. Second, how quickly standardization occurs around protocols like MCP and evaluation frameworks. The Linux Foundation's involvement signals movement toward industry-wide coordination, though history suggests these processes take longer than anyone hopes. Third, whether enterprises develop sufficient expertise to properly scope, deploy, and oversee agent systems—the people problem that no amount of technology can solve alone.

What's becoming evident is that the next phase of AI deployment won't be measured in agent counts or API calls. It will be measured in tasks completed reliably across realistic time horizons, with verifiable outcomes and appropriate governance.

For engineering leaders evaluating agent strategies, the implications are straightforward if uncomfortable. Pilots will continue proliferating—that's unavoidable given competitive pressure and executive enthusiasm. But production success requires infrastructure that doesn't yet exist at scale. The organizations that solve long-horizon reliability first—whether through applied research labs building better training environments, platforms providing better governance tooling, or some combination—will define what autonomous AI means in practice rather than in press releases.

The consulting firms deploying thousands of agents are betting that incremental improvement in foundation models will eventually close the reliability gap. The research labs building production-grade training environments are betting that gap requires different infrastructure entirely.

Both could be right. Or both could be oversimplifying. The enterprises caught in between need to decide how much risk they're willing to accept while the industry figures it out. McKinsey's 25,000 agents make for an impressive headline. Whether they make for impressive work remains to be seen.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • AI Agents Are Running Biology Labs: The Self-Driving Lab Revolution
  • The Race to Build Infrastructure for Human Digital Twins
  • The $4.5B Seed Round: How AI Labs Rewrote Venture Capital Rules
  • Drones That Charge on Power Lines Promise Infinite Flight
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.