Twenty-five percent. That's the success rate of Claude Opus 4.6—currently the strongest large language model in the world—when asked to complete real software engineering tasks from start to finish. Not coding puzzles. Not algorithm challenges. The messy, multi-step work that consumes most of an engineer's day: coordinating tools, managing deployments, debugging across systems, actually shipping something that runs.
The figure comes from Polymath Labs, a Y Combinator startup that published results from its Horizon-SWE benchmark in early February. And it crystallizes what the AI industry has been reluctant to say plainly: long-horizon autonomy, the ability to execute complex workflows over hours or days, remains stubbornly out of reach. Agents can dazzle with discrete tricks—writing functions, answering questions, generating boilerplate. String those actions into a coherent, multi-hour process that navigates a production environment? That's where things fall apart.
It's a problem the entire industry is feeling, even if the marketing materials suggest otherwise.
Why the Frontier Hits a Wall
Dylan Ma and Naren Yenuganti started Polymath Labs with this gap in mind. Ma had worked as a research engineer at Hume AI and at AWS; Yenuganti put in time at Plaid and Amazon. Both came out of UC Berkeley research programs. Both watched agents perform impressively in controlled demos, then stumble when asked to handle the sprawl and friction of real systems.
"We build simulated worlds for training and evaluating AI agents on long-horizon, multi-tool tasks in any domain," the company explains. Their first major public effort, Horizon-SWE, went live on February 2 (with an update on February 6) and targets software engineering specifically—a domain where the chasm between laboratory success and production reliability is particularly wide.
The benchmark constructs what amounts to a containerized software company, complete with 73 tools spanning the full engineering lifecycle. Agents must navigate codebases, coordinate deployments, manage infrastructure, verify correctness. Tasks mirror the real work that happens outside the IDE, where engineers spend much of their time wrangling interconnected systems rather than writing pristine code from scratch.
Polymath ran its evaluations using OpenHands, an open-source agent harness designed for standardized testing. The results? Claude Opus 4.6 topped the leaderboard at 25.5% on strict pass/fail criteria. The rest of the frontier models clustered lower: Claude Opus 4.5 at 22.1%, Sonnet 4.5 at 20.8%, Gemini 3 Pro at 19.1%. Even GPT-5.2 Codex, trained explicitly for coding tasks, managed just 18.0%.
The benchmark includes a partial-credit track that scores deployment competence, engineering quality, and correctness separately. On that composite metric, Opus 4.6 reaches 60.4. Opus 4.5 scores 54.9. GPT-5.2 Codex hits 50.1. The models can, in other words, handle individual sub-tasks with some proficiency. But threading them together into a successful end-to-end workflow—that's another matter entirely.
The Pattern Repeats Elsewhere
This isn't an isolated finding. METR, the AI safety nonprofit, updated its Time Horizon evaluation on January 29. Their analysis suggests Claude Opus 4.6 can complete tasks with a 50% success rate when those tasks take roughly 14.5 hours—a meaningful improvement over earlier models, certainly, but far short of the days-long workflows that human engineers execute routinely without drama.
An ICLR 2026 survey on web agents noted that planning failures remain the most common error source across contemporary systems, whether tested on WebArena or OSWorld. The models are getting better at isolated maneuvers. Connecting those maneuvers into coherent, resilient strategies? That skill still lags.
The disconnect has spurred a wave of academic work. HiPER, published in February, achieved 97.4% success on ALFWorld and 83.3% on WebShop using hierarchical reinforcement learning with explicit credit assignment—an approach that tries to teach agents which actions, buried deep in a long sequence, actually mattered. HiMAC, released March 1, introduced critic-free hierarchical policy optimization. MiRA, which appeared March 20, reported nearly a 37-point absolute gain on WebArena-Lite when applying milestone-based reward shaping to Gemma3-12B.
NVIDIA's ToolOrchestra, detailed in a November 2025 paper and discussed again in a January blog post, demonstrated that a small RL-trained "orchestrator" could coordinate expert models and tools more effectively than relying on monolithic prompting alone. The emphasis was on sample efficiency—ToolOrchestra showed gains without requiring enormous rollout budgets.
The common thread across these efforts: scaling up context windows or parameter counts won't solve the long-horizon problem on its own. What's needed, apparently, is better credit assignment, hierarchical planning, and reward structures that can bridge delayed outcomes with intermediate progress signals.
The Commercial Pressure Is Real

This isn't just an academic curiosity. Gartner predicted last October that 40% of enterprise applications would embed task-specific AI agents by this year, up from under 5% in 2025. McKinsey's State of AI 2025 report found 62% of survey respondents already experimenting with agents. A March article citing a Microsoft Security executive claimed 80% of Fortune 500 companies are using agents in daily operations—though that particular stat comes from a vendor source and warrants a degree of skepticism.
The hyperscalers are moving aggressively regardless. Microsoft announced Copilot agents and domain agents at Ignite in November 2024, with steady updates throughout 2025. AWS made multi-agent collaboration generally available on Bedrock in March of last year. Salesforce has been pushing Agentforce hard, launching an AgentExchange marketplace that went live by March 10 of this year. The IRS, reportedly, is now running Agentforce in some capacity—a deployment that surfaced in reports dating to November 2025.
Cognition's Devin, the autonomous software engineering agent that grabbed headlines in 2024, has been evolving quietly through enterprise pilots. Goldman Sachs was testing it as of mid-2025, according to reports at the time. Cognizant announced a partnership with Cognition on January 28.
But enthusiasm is colliding with practical limits. OpenAI published a post on February 23 explaining why it no longer treats SWE-bench Verified as a sole evaluation metric, citing contamination concerns and performance ceiling issues. Days later, on February 28 (updated March 2), the company outlined red lines for its collaboration with the Department of Defense—no autonomous weapons, no mass domestic surveillance, no high-stakes automated decisions. Even at the frontier, full autonomy remains constrained by safety and reliability questions.
Perhaps more than the labs expected.
An Emerging Ecosystem
Polymath positions itself within an infrastructure layer that's still taking shape—startups building environments to train, evaluate, and deploy long-horizon agents. A September TechCrunch piece profiled companies like Surge, Mercor, Prime Intellect, and Mechanize, all focused on environment generation for agent training. The article captured both the optimism, quoting a16z partner Jennifer Li on rising demand, and the skepticism: OpenAI engineer Sherwin Wu, for instance, said he was "short on RL environment startups" due to the persistent risk of reward hacking.
Scale AI has been active here too, contributing to SWE-bench Pro orchestration and releasing the MCP-Atlas benchmark in January. The Model Context Protocol itself—which hit its one-year anniversary on November 25—has emerged as something resembling a standard interface for tool discovery and invocation. Polymath's Horizon-SWE relies on MCP for multi-tool coordination.
Regulatory and standards efforts are accelerating alongside the technical work. NIST launched an AI Agent Standards Initiative on February 17, signaling federal interest in interoperable, secure agent ecosystems. The EU AI Act, which entered into force in August 2024, has phased obligations coming due throughout this year. The Cloud Security Alliance published an Agentic AI Red Teaming Guide last May, a document increasingly cited in security guidance as 2026 has progressed.
Even something as basic as web access—table stakes for many agent applications—faces tightening legal constraints. Google's lawsuit against SerpApi, filed December 19, alleges DMCA anti-circumvention related to SearchGuard protections. A hearing is reportedly set for May 19. The case could establish precedents that shape how agents interact with search and scraping infrastructure going forward.
What 25% Actually Means

Polymath's focus on "production-grade systems," "verifiable outcomes," and "realism" reflects a broader maturation in how the industry evaluates agents. The company noted in a January 10 post that it works with frontier labs to customize and scale simulated environments—suggesting these worlds may become training grounds for the next wave of models.
But the 25% completion rate on Horizon-SWE isn't just a technical benchmark. It's a statement about where autonomy actually stands, as opposed to where the hype cycle suggests it should be. Gartner and Forrester both frame this year as an inflection point for multiagent systems, with governance and security cited as key adoption bottlenecks. Yet the more fundamental barrier remains technical: getting agents to reliably execute multi-hour, multi-tool workflows without human intervention.
Whether Polymath's simulated worlds accelerate progress on that front depends largely on adoption by the labs doing the heavy computational lifting. The company emerged from Y Combinator's Winter 2026 batch with a two-person team, according to its YC profile. It hasn't publicly announced funding beyond the accelerator. But the problem it's tackling—and the benchmark it's built to track progress—addresses a gap the entire industry is confronting, whether it admits it or not.
The race to long-horizon autonomy isn't won by models that ace coding interviews or generate clever one-liners in Slack. It's won by systems that can coordinate 73 tools across a containerized software company, maintain state through dozens of decision points, and ship working code to production. Right now, even the most capable model in the world manages that feat one time in four.
Which is to say: not often enough.
