The industry's go-to coding benchmark, SWE-bench Verified, has faced growing scrutiny over contamination issues and flawed test cases. Frontier models had learned to game the system rather than solve real problems.
But here's what matters: the limitations of existing benchmarks have exposed a more fundamental problem. As AI labs push toward agents capable of handling workflows that stretch across days, not minutes, they're discovering that existing evaluation frameworks can't capture what actually happens in production environments. The messy, interconnected reality of enterprise work—where small mistakes cascade and nothing happens in isolation—remains stubbornly difficult to measure.
A two-person startup out of Y Combinator thinks it has an answer.
The 25% Problem
Polymath Labs, founded by Dylan Ma and Naren Yenuganti, released Horizon-SWE on February 2, 2026. Unlike traditional coding benchmarks, it evaluates agents across an entire production stack: Slack threads, Linear tickets, Notion documentation, Sentry error alerts, Prometheus metrics, CI/CD pipelines, and a live monorepo-backed application. The kind of environment an on-call engineer navigates every day.
The results? Sobering, to put it mildly. Leading models achieve roughly 25% pass rates on end-to-end tasks. Claude Opus 4.6 tops the leaderboard at 25.5%—with a margin of error of ±3.6%—as of early February. Even on a partial-credit composite score that rewards incremental progress, the best systems barely clear 60%.
These aren't just numbers on a leaderboard. They represent the gap between what AI agents can demonstrate on isolated coding tasks and what enterprises actually need: the ability to diagnose an incident, coordinate across multiple tools, update documentation, and push a fix through CI without breaking production. It's one thing to patch a bug in a vacuum. It's another to do it in a system where everything connects to everything else.
The Adoption Paradox
The shift from copilots to autonomous agents is happening faster than the technology can keep up. A January survey from the Mayfield CXO Network found that 42% of enterprises now run AI agents in production, with another 30% in pilot phases. Gartner predicts that by 2028, 60% of brands will deploy agentic AI for one-to-one customer interactions. BCC Research pegs the AI agents market at $8 billion in 2025, projecting growth to $48.3 billion by 2030—a 43.3% compound annual growth rate.
Yet adoption remains uneven. Lyzr's Q1 2026 report reveals that 62% of companies exploring agents lack a clear starting point. Another 32% stall after initial pilots.
The problem isn't interest or investment. It's reliability.
Agents that perform well on benchmarks often falter when handed real workflows spanning hours or days, where small errors compound and recovery paths aren't obvious. Klarna's widely cited customer service agent—deployed in early 2024 and handling two-thirds of chats in its first month, equivalent to 700 full-time employees—offered an early proof of concept. The company reported a $39-$40 million profit improvement in 2024, later revised upward to $60 million in annual savings by 2025.
But even Klarna's success story came with caveats. The company has since adopted a hybrid approach, blending agent autonomy with human oversight for edge cases. Simple, repetitive support tasks, it turns out, are a long way from the open-ended, multi-tool workflows enterprises want to automate next.
When Time Becomes the Problem

The challenge isn't just technical complexity. It's time.
METR, an AI safety research organization, reported in March 2025 that the duration of tasks agents can complete autonomously has been doubling roughly every seven months. Extrapolate that trend, and week-long tasks become plausible in the near term. But longer horizons introduce compounding failure modes that don't show up in short benchmarks.
Consider a typical software engineering workflow: An on-call engineer receives a Sentry alert. Investigates metrics in Grafana. Correlates the issue with recent commits. Discusses root cause in Slack. Opens a Linear ticket. Drafts a fix. Runs tests locally. Pushes code. Monitors CI/CD. Updates Notion documentation.
Each step depends on the previous ones. A single misstep—misinterpreting a metric, linking the wrong commit, breaking a test—can cascade into wasted effort. Or worse, introduce new bugs.
Recent benchmarks underscore the problem. LongCLI-Bench, published February 15, evaluates agents on command-line programming tasks requiring dozens of steps; state-of-the-art systems achieve pass rates below 20%, with many failing in the first few moves. DeepPlanning, released January 26, tests multi-day travel planning and multi-product shopping under verifiable constraints. Current agents struggle to satisfy even basic requirements.
AgentLAB, a February 18 benchmark, demonstrates that long-horizon agents are vulnerable to adversarial attacks—intent hijacking, task injection, memory poisoning—across 28 environments and 644 test cases.
The common thread? Planning and state-tracking. Agents that excel at single-turn reasoning—answering a question, generating code—often lack the memory, error recovery, and hierarchical planning needed to navigate workflows that unfold over hours. Andrej Karpathy noted in November 2025 that fully autonomous agents remain a decade-scale challenge, citing compounding error rates and the absence of continual learning mechanisms.
It's a sobering assessment. Maybe also a realistic one.
The Hidden Infrastructure Play
Behind the scenes, AI labs are betting heavily on a solution: reinforcement learning environments. Unlike supervised datasets, which teach agents to mimic human outputs, RL environments let agents learn through trial and error in realistic settings.
The catch? Building high-fidelity environments—complete with tool integrations, verifiable outcomes, and non-trivial state dynamics—is labor-intensive and expensive.
Wing VC, an early-stage venture firm, noted in January that labs are now treating RL environments as a distinct budget line item, with projected spending in the tens of millions annually. Growth expectations through 2026: 3-5×. TechCrunch reported in September 2025 that demand for training environments has spawned a mini-ecosystem. Companies like Mechanize (focused on software engineering tasks and rumored to be working with Anthropic), Prime Intellect (positioning itself as "Hugging Face for RL environments"), and data providers such as Surge, Mercor, and Scale are all pivoting into environment creation.
The competition is crowding the space. But most players still rely on manual curation.
Polymath Labs sees an opening in automation. According to the company's Y Combinator profile, its mission is to build "world generation models and systems to automate and align environment creation," with an end goal of generating "realistic, long-horizon environments from a text description alone." Starting with software engineering—where the need is acute and the feedback loops are measurable—the founders aim to reduce the human labor required to spin up production-grade RL training and evaluation setups.
Two Founders, One Benchmark
Dylan Ma, previously at Hume AI and AWS, and Naren Yenuganti, who spent time at Plaid and Amazon, both studied at UC Berkeley before founding Polymath Labs. Their January 10 blog post describes a philosophy centered on "realism, verifiable outcomes, and long-horizon workflows across tools."
Horizon-SWE embodies that philosophy. Unlike earlier benchmarks that isolate coding tasks or rely on static datasets, Horizon-SWE simulates an entire software engineering lifecycle. Agents interact with live services, respond to multi-tool incidents, and produce outcomes that can be verified end-to-end.
The benchmark's design is deliberately production-grade. Tasks involve real integrations—Slack for communication, Linear for project tracking, Notion for documentation, Sentry for error monitoring, Prometheus and Grafana for observability, and a full CI/CD pipeline tied to a monorepo. Agents must navigate this ecosystem autonomously, making decisions that mirror what an on-call engineer would face. The scoring system includes both pass/fail criteria and partial-credit composites, rewarding incremental progress even when tasks aren't fully resolved.
As of the February 6 leaderboard update, Claude Opus 4.6 leads on the binary pass rate at 25.5%. On the composite partial-credit metric, which captures progress across sub-tasks, the best score is 60.4.
The results suggest that even frontier models have significant headroom before they're ready for unsupervised, long-horizon production work. Polymath's public communications emphasize that these numbers aren't just academic. They reflect the gap between today's agents and the autonomous systems enterprises are eager to deploy.
The company's positioning is explicitly tied to the evolving benchmark landscape. As concerns grow about contamination and the limitations of existing coding benchmarks, labs and enterprise tool vendors now need third-party benchmarks to validate agent capabilities, and Horizon-SWE offers a standardized, realistic alternative.
The broader question is whether Polymath can scale its environment generation approach beyond software engineering. Customer service, data analysis, compliance workflows—other domains where long-horizon autonomy would unlock value.
The Regulation Factor

The next year will test whether automated RL environment generation can become a sustainable business. Wing VC's analysis suggests the market is there: as labs scale training runs and expand into new domains, they'll need environments faster than human curators can build them.
DeepMind's Genie series, which generates playable 3D worlds from text prompts, and WebWorld, a February 2026 system trained on over a million web interactions, hint at the potential for synthetic environment generation at scale. Polymath's bet is that the same techniques can be adapted for enterprise workflows, where verifiability and realism matter more than visual fidelity.
Regulation may accelerate demand. The EU AI Act's high-risk provisions and enforcement mechanisms take effect August 2, 2026, and general-purpose AI obligations were already in force as of August 2, 2025. In the U.S., California's SB-53 transparency law and Colorado's AI Act (effective June 30, 2026) add compliance pressure. NIST's draft Cyber AI Profile, released in December 2025 with a comment period that closed January 30, ties AI adoption to cybersecurity planning.
All of this points toward a future where enterprises need auditable, verifiable agent behaviors. Precisely what realistic RL environments and benchmarks like Horizon-SWE aim to provide.
The open question is whether the current crop of models will hit diminishing returns on long-horizon tasks. Or whether architectural improvements—hierarchical planning, better memory systems, process-level reward shaping—will push pass rates substantially higher.
HiPER, a hierarchical reinforcement learning system published February 18, shows promise on sparse-reward tasks by explicitly assigning credit across levels of a plan hierarchy. WebWorld's 30+ step horizons and improved WebArena performance suggest that world models trained on diverse interaction data can generalize better than static benchmarks.
If these trends hold, the ~25% pass rates on Horizon-SWE could climb quickly. Though it's unclear whether incremental gains will suffice for true production autonomy or whether a more fundamental rethink is required.
What Founders Need to Know
For founders building agent products, the message is straightforward: long-horizon reliability is now the bottleneck.
Enterprises are moving beyond proof-of-concept demos. Buyers—increasingly business leaders rather than engineering teams—want measurable outcomes tied to KPIs. Integration, governance, and audit trails matter as much as raw capability. The Lyzr report's finding that 62% of enterprises lack a clear starting point suggests a services opportunity around deployment and monitoring, not just tooling.
For AI labs, the shift from copilots to agents is forcing a reallocation of capital. If Wing VC's estimates hold, spending on RL environments could grow 3-5× through 2026, with tens of millions in annual budget. That's real money, redirecting labor and compute away from dataset curation and toward environment engineering.
Polymath's automation thesis—if it works—could compress that timeline and lower costs. But the company is competing against well-funded incumbents like Mechanize and Prime Intellect, plus the in-house efforts of labs that may prefer to build their own infrastructure.
---
The broader arc is clear, even if the specifics remain uncertain. AI agents are moving from narrow, supervised tasks to open-ended, multi-step workflows. The evaluation frameworks that guided the last generation of models are breaking down—contaminated by overuse and misaligned with the behaviors enterprises actually need.
New benchmarks, new environments, new training methodologies are emerging to fill the gap. Polymath Labs is placing an early bet on automated world generation as the scalable solution.
Whether that bet pays off depends on execution, timing, and how quickly the rest of the ecosystem catches up. But one thing seems certain: the 25% problem isn't going away on its own.
