Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechMarch 28, 2026

Space Solar Arrays Hit Inflection Point as YC-Backed Startup Enters Race

Space Solar Arrays Hit Inflection Point as YC-Backed Startup Enters Race
YcSpace Tech+3
SaaS iconSaaSMarch 28, 2026

Why On-Device AI Is Moving From Hype to Reality in 2026

Why On-Device AI Is Moving From Hype to Reality in 2026
YcEdge Computing+3
SaaS iconSaaS
March 28, 2026
YcAi AgentsAi BenchmarkingAutonomous SystemsB2b Saas

How Polymath Labs Is Solving AI's Long-Horizon Problem

YC-backed startup tackles critical barrier for autonomous AI agents with new benchmark showing frontier models complete only 25% of complex software tasks.

How Polymath Labs Is Solving AI's Long-Horizon Problem

Twenty-five percent. That's the success rate of Claude Opus 4.6—currently the strongest large language model in the world—when asked to complete real software engineering tasks from start to finish. Not coding puzzles. Not algorithm challenges. The messy, multi-step work that consumes most of an engineer's day: coordinating tools, managing deployments, debugging across systems, actually shipping something that runs.

The figure comes from Polymath Labs, a Y Combinator startup that published results from its Horizon-SWE benchmark in early February. And it crystallizes what the AI industry has been reluctant to say plainly: long-horizon autonomy, the ability to execute complex workflows over hours or days, remains stubbornly out of reach. Agents can dazzle with discrete tricks—writing functions, answering questions, generating boilerplate. String those actions into a coherent, multi-hour process that navigates a production environment? That's where things fall apart.

It's a problem the entire industry is feeling, even if the marketing materials suggest otherwise.

Why the Frontier Hits a Wall

Dylan Ma and Naren Yenuganti started Polymath Labs with this gap in mind. Ma had worked as a research engineer at Hume AI and at AWS; Yenuganti put in time at Plaid and Amazon. Both came out of UC Berkeley research programs. Both watched agents perform impressively in controlled demos, then stumble when asked to handle the sprawl and friction of real systems.

"We build simulated worlds for training and evaluating AI agents on long-horizon, multi-tool tasks in any domain," the company explains. Their first major public effort, Horizon-SWE, went live on February 2 (with an update on February 6) and targets software engineering specifically—a domain where the chasm between laboratory success and production reliability is particularly wide.

The benchmark constructs what amounts to a containerized software company, complete with 73 tools spanning the full engineering lifecycle. Agents must navigate codebases, coordinate deployments, manage infrastructure, verify correctness. Tasks mirror the real work that happens outside the IDE, where engineers spend much of their time wrangling interconnected systems rather than writing pristine code from scratch.

Polymath ran its evaluations using OpenHands, an open-source agent harness designed for standardized testing. The results? Claude Opus 4.6 topped the leaderboard at 25.5% on strict pass/fail criteria. The rest of the frontier models clustered lower: Claude Opus 4.5 at 22.1%, Sonnet 4.5 at 20.8%, Gemini 3 Pro at 19.1%. Even GPT-5.2 Codex, trained explicitly for coding tasks, managed just 18.0%.

The benchmark includes a partial-credit track that scores deployment competence, engineering quality, and correctness separately. On that composite metric, Opus 4.6 reaches 60.4. Opus 4.5 scores 54.9. GPT-5.2 Codex hits 50.1. The models can, in other words, handle individual sub-tasks with some proficiency. But threading them together into a successful end-to-end workflow—that's another matter entirely.

The Pattern Repeats Elsewhere

This isn't an isolated finding. METR, the AI safety nonprofit, updated its Time Horizon evaluation on January 29. Their analysis suggests Claude Opus 4.6 can complete tasks with a 50% success rate when those tasks take roughly 14.5 hours—a meaningful improvement over earlier models, certainly, but far short of the days-long workflows that human engineers execute routinely without drama.

An ICLR 2026 survey on web agents noted that planning failures remain the most common error source across contemporary systems, whether tested on WebArena or OSWorld. The models are getting better at isolated maneuvers. Connecting those maneuvers into coherent, resilient strategies? That skill still lags.

The disconnect has spurred a wave of academic work. HiPER, published in February, achieved 97.4% success on ALFWorld and 83.3% on WebShop using hierarchical reinforcement learning with explicit credit assignment—an approach that tries to teach agents which actions, buried deep in a long sequence, actually mattered. HiMAC, released March 1, introduced critic-free hierarchical policy optimization. MiRA, which appeared March 20, reported nearly a 37-point absolute gain on WebArena-Lite when applying milestone-based reward shaping to Gemma3-12B.

NVIDIA's ToolOrchestra, detailed in a November 2025 paper and discussed again in a January blog post, demonstrated that a small RL-trained "orchestrator" could coordinate expert models and tools more effectively than relying on monolithic prompting alone. The emphasis was on sample efficiency—ToolOrchestra showed gains without requiring enormous rollout budgets.

The common thread across these efforts: scaling up context windows or parameter counts won't solve the long-horizon problem on its own. What's needed, apparently, is better credit assignment, hierarchical planning, and reward structures that can bridge delayed outcomes with intermediate progress signals.

The Commercial Pressure Is Real

Digital illustration for article section "The Commercial Pressure Is Real" in "How Polymath Labs Is Solving AI's Long-Horizon Problem" - A minimalist and conceptual representation of rapid commercial pressure and enterprise adoption, fea...

This isn't just an academic curiosity. Gartner predicted last October that 40% of enterprise applications would embed task-specific AI agents by this year, up from under 5% in 2025. McKinsey's State of AI 2025 report found 62% of survey respondents already experimenting with agents. A March article citing a Microsoft Security executive claimed 80% of Fortune 500 companies are using agents in daily operations—though that particular stat comes from a vendor source and warrants a degree of skepticism.

The hyperscalers are moving aggressively regardless. Microsoft announced Copilot agents and domain agents at Ignite in November 2024, with steady updates throughout 2025. AWS made multi-agent collaboration generally available on Bedrock in March of last year. Salesforce has been pushing Agentforce hard, launching an AgentExchange marketplace that went live by March 10 of this year. The IRS, reportedly, is now running Agentforce in some capacity—a deployment that surfaced in reports dating to November 2025.

Cognition's Devin, the autonomous software engineering agent that grabbed headlines in 2024, has been evolving quietly through enterprise pilots. Goldman Sachs was testing it as of mid-2025, according to reports at the time. Cognizant announced a partnership with Cognition on January 28.

But enthusiasm is colliding with practical limits. OpenAI published a post on February 23 explaining why it no longer treats SWE-bench Verified as a sole evaluation metric, citing contamination concerns and performance ceiling issues. Days later, on February 28 (updated March 2), the company outlined red lines for its collaboration with the Department of Defense—no autonomous weapons, no mass domestic surveillance, no high-stakes automated decisions. Even at the frontier, full autonomy remains constrained by safety and reliability questions.

Perhaps more than the labs expected.

An Emerging Ecosystem

Polymath positions itself within an infrastructure layer that's still taking shape—startups building environments to train, evaluate, and deploy long-horizon agents. A September TechCrunch piece profiled companies like Surge, Mercor, Prime Intellect, and Mechanize, all focused on environment generation for agent training. The article captured both the optimism, quoting a16z partner Jennifer Li on rising demand, and the skepticism: OpenAI engineer Sherwin Wu, for instance, said he was "short on RL environment startups" due to the persistent risk of reward hacking.

Scale AI has been active here too, contributing to SWE-bench Pro orchestration and releasing the MCP-Atlas benchmark in January. The Model Context Protocol itself—which hit its one-year anniversary on November 25—has emerged as something resembling a standard interface for tool discovery and invocation. Polymath's Horizon-SWE relies on MCP for multi-tool coordination.

Regulatory and standards efforts are accelerating alongside the technical work. NIST launched an AI Agent Standards Initiative on February 17, signaling federal interest in interoperable, secure agent ecosystems. The EU AI Act, which entered into force in August 2024, has phased obligations coming due throughout this year. The Cloud Security Alliance published an Agentic AI Red Teaming Guide last May, a document increasingly cited in security guidance as 2026 has progressed.

Even something as basic as web access—table stakes for many agent applications—faces tightening legal constraints. Google's lawsuit against SerpApi, filed December 19, alleges DMCA anti-circumvention related to SearchGuard protections. A hearing is reportedly set for May 19. The case could establish precedents that shape how agents interact with search and scraping infrastructure going forward.

What 25% Actually Means

Digital illustration for article section "What 25% Actually Means" in "How Polymath Labs Is Solving AI's Long-Horizon Problem" - A minimalist and modern conceptual image featuring a pristine, highly detailed architectural diorama...

Polymath's focus on "production-grade systems," "verifiable outcomes," and "realism" reflects a broader maturation in how the industry evaluates agents. The company noted in a January 10 post that it works with frontier labs to customize and scale simulated environments—suggesting these worlds may become training grounds for the next wave of models.

But the 25% completion rate on Horizon-SWE isn't just a technical benchmark. It's a statement about where autonomy actually stands, as opposed to where the hype cycle suggests it should be. Gartner and Forrester both frame this year as an inflection point for multiagent systems, with governance and security cited as key adoption bottlenecks. Yet the more fundamental barrier remains technical: getting agents to reliably execute multi-hour, multi-tool workflows without human intervention.

Whether Polymath's simulated worlds accelerate progress on that front depends largely on adoption by the labs doing the heavy computational lifting. The company emerged from Y Combinator's Winter 2026 batch with a two-person team, according to its YC profile. It hasn't publicly announced funding beyond the accelerator. But the problem it's tackling—and the benchmark it's built to track progress—addresses a gap the entire industry is confronting, whether it admits it or not.

The race to long-horizon autonomy isn't won by models that ace coding interviews or generate clever one-liners in Slack. It's won by systems that can coordinate 73 tools across a containerized software company, maintain state through dozens of decision points, and ship working code to production. Right now, even the most capable model in the world manages that feat one time in four.

Which is to say: not often enough.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Space Solar Arrays Hit Inflection Point as YC-Backed Startup Enters Race
  • Why On-Device AI Is Moving From Hype to Reality in 2026
  • YC-Backed Ritivel Automates Drug Regulatory Filings with AI
  • The Robot Cowboys: How AI Drones Are Transforming Cattle Ranching
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.