Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
SaaS iconSaaSMarch 2, 2026

AI Agents Dominate Y Combinator's First Spring 2025 Demo Day

AI Agents Dominate Y Combinator's First Spring 2025 Demo Day
YcAi Agents+3
Healthtech & Biotech iconHealthtech & BiotechMarch 2, 2026

Lab-Grown Neurons Playing DOOM Signal Bio-Computing's Commercial Dawn

Lab-Grown Neurons Playing DOOM Signal Bio-Computing's Commercial Dawn
BiotechAi Hardware+3

Founders Mentioned

Dylan Ma

Polymath Labs

saas icon
SaaS

Naren Yenuganti

Polymath Labs

saas icon
SaaS

Dylan Ma

Polymath Labs

saas icon
SaaS

Naren Yenuganti

Polymath Labs

saas icon
SaaS
SaaS iconSaaS
March 2, 2026
YcAi AgentsReinforcement LearningAi BenchmarkingEnterprise Ai

YC's Polymath Labs Tackles AI Agents' Long-Horizon Challenge

As AI agents struggle with multi-step tasks, YC W26 company Polymath Labs automates RL environment generation. New benchmark shows leading models achieve only 25% pass rates.

YC's Polymath Labs Tackles AI Agents' Long-Horizon Challenge

The industry's go-to coding benchmark, SWE-bench Verified, has faced growing scrutiny over contamination issues and flawed test cases. Frontier models had learned to game the system rather than solve real problems.

But here's what matters: the limitations of existing benchmarks have exposed a more fundamental problem. As AI labs push toward agents capable of handling workflows that stretch across days, not minutes, they're discovering that existing evaluation frameworks can't capture what actually happens in production environments. The messy, interconnected reality of enterprise work—where small mistakes cascade and nothing happens in isolation—remains stubbornly difficult to measure.

A two-person startup out of Y Combinator thinks it has an answer.

The 25% Problem

Polymath Labs, founded by Dylan Ma and Naren Yenuganti, released Horizon-SWE on February 2, 2026. Unlike traditional coding benchmarks, it evaluates agents across an entire production stack: Slack threads, Linear tickets, Notion documentation, Sentry error alerts, Prometheus metrics, CI/CD pipelines, and a live monorepo-backed application. The kind of environment an on-call engineer navigates every day.

The results? Sobering, to put it mildly. Leading models achieve roughly 25% pass rates on end-to-end tasks. Claude Opus 4.6 tops the leaderboard at 25.5%—with a margin of error of ±3.6%—as of early February. Even on a partial-credit composite score that rewards incremental progress, the best systems barely clear 60%.

These aren't just numbers on a leaderboard. They represent the gap between what AI agents can demonstrate on isolated coding tasks and what enterprises actually need: the ability to diagnose an incident, coordinate across multiple tools, update documentation, and push a fix through CI without breaking production. It's one thing to patch a bug in a vacuum. It's another to do it in a system where everything connects to everything else.

The Adoption Paradox

The shift from copilots to autonomous agents is happening faster than the technology can keep up. A January survey from the Mayfield CXO Network found that 42% of enterprises now run AI agents in production, with another 30% in pilot phases. Gartner predicts that by 2028, 60% of brands will deploy agentic AI for one-to-one customer interactions. BCC Research pegs the AI agents market at $8 billion in 2025, projecting growth to $48.3 billion by 2030—a 43.3% compound annual growth rate.

Yet adoption remains uneven. Lyzr's Q1 2026 report reveals that 62% of companies exploring agents lack a clear starting point. Another 32% stall after initial pilots.

The problem isn't interest or investment. It's reliability.

Agents that perform well on benchmarks often falter when handed real workflows spanning hours or days, where small errors compound and recovery paths aren't obvious. Klarna's widely cited customer service agent—deployed in early 2024 and handling two-thirds of chats in its first month, equivalent to 700 full-time employees—offered an early proof of concept. The company reported a $39-$40 million profit improvement in 2024, later revised upward to $60 million in annual savings by 2025.

But even Klarna's success story came with caveats. The company has since adopted a hybrid approach, blending agent autonomy with human oversight for edge cases. Simple, repetitive support tasks, it turns out, are a long way from the open-ended, multi-tool workflows enterprises want to automate next.

When Time Becomes the Problem

Digital illustration for article section "When Time Becomes the Problem" in "YC's Polymath Labs Tackles AI Agents' Long-Horizon Challenge" - A surreal conceptual photograph featuring a minimalist, elongated hourglass centered on a smooth, re...

The challenge isn't just technical complexity. It's time.

METR, an AI safety research organization, reported in March 2025 that the duration of tasks agents can complete autonomously has been doubling roughly every seven months. Extrapolate that trend, and week-long tasks become plausible in the near term. But longer horizons introduce compounding failure modes that don't show up in short benchmarks.

Consider a typical software engineering workflow: An on-call engineer receives a Sentry alert. Investigates metrics in Grafana. Correlates the issue with recent commits. Discusses root cause in Slack. Opens a Linear ticket. Drafts a fix. Runs tests locally. Pushes code. Monitors CI/CD. Updates Notion documentation.

Each step depends on the previous ones. A single misstep—misinterpreting a metric, linking the wrong commit, breaking a test—can cascade into wasted effort. Or worse, introduce new bugs.

Recent benchmarks underscore the problem. LongCLI-Bench, published February 15, evaluates agents on command-line programming tasks requiring dozens of steps; state-of-the-art systems achieve pass rates below 20%, with many failing in the first few moves. DeepPlanning, released January 26, tests multi-day travel planning and multi-product shopping under verifiable constraints. Current agents struggle to satisfy even basic requirements.

AgentLAB, a February 18 benchmark, demonstrates that long-horizon agents are vulnerable to adversarial attacks—intent hijacking, task injection, memory poisoning—across 28 environments and 644 test cases.

The common thread? Planning and state-tracking. Agents that excel at single-turn reasoning—answering a question, generating code—often lack the memory, error recovery, and hierarchical planning needed to navigate workflows that unfold over hours. Andrej Karpathy noted in November 2025 that fully autonomous agents remain a decade-scale challenge, citing compounding error rates and the absence of continual learning mechanisms.

It's a sobering assessment. Maybe also a realistic one.

The Hidden Infrastructure Play

Behind the scenes, AI labs are betting heavily on a solution: reinforcement learning environments. Unlike supervised datasets, which teach agents to mimic human outputs, RL environments let agents learn through trial and error in realistic settings.

The catch? Building high-fidelity environments—complete with tool integrations, verifiable outcomes, and non-trivial state dynamics—is labor-intensive and expensive.

Wing VC, an early-stage venture firm, noted in January that labs are now treating RL environments as a distinct budget line item, with projected spending in the tens of millions annually. Growth expectations through 2026: 3-5×. TechCrunch reported in September 2025 that demand for training environments has spawned a mini-ecosystem. Companies like Mechanize (focused on software engineering tasks and rumored to be working with Anthropic), Prime Intellect (positioning itself as "Hugging Face for RL environments"), and data providers such as Surge, Mercor, and Scale are all pivoting into environment creation.

The competition is crowding the space. But most players still rely on manual curation.

Polymath Labs sees an opening in automation. According to the company's Y Combinator profile, its mission is to build "world generation models and systems to automate and align environment creation," with an end goal of generating "realistic, long-horizon environments from a text description alone." Starting with software engineering—where the need is acute and the feedback loops are measurable—the founders aim to reduce the human labor required to spin up production-grade RL training and evaluation setups.

Two Founders, One Benchmark

Dylan Ma, previously at Hume AI and AWS, and Naren Yenuganti, who spent time at Plaid and Amazon, both studied at UC Berkeley before founding Polymath Labs. Their January 10 blog post describes a philosophy centered on "realism, verifiable outcomes, and long-horizon workflows across tools."

Horizon-SWE embodies that philosophy. Unlike earlier benchmarks that isolate coding tasks or rely on static datasets, Horizon-SWE simulates an entire software engineering lifecycle. Agents interact with live services, respond to multi-tool incidents, and produce outcomes that can be verified end-to-end.

The benchmark's design is deliberately production-grade. Tasks involve real integrations—Slack for communication, Linear for project tracking, Notion for documentation, Sentry for error monitoring, Prometheus and Grafana for observability, and a full CI/CD pipeline tied to a monorepo. Agents must navigate this ecosystem autonomously, making decisions that mirror what an on-call engineer would face. The scoring system includes both pass/fail criteria and partial-credit composites, rewarding incremental progress even when tasks aren't fully resolved.

As of the February 6 leaderboard update, Claude Opus 4.6 leads on the binary pass rate at 25.5%. On the composite partial-credit metric, which captures progress across sub-tasks, the best score is 60.4.

The results suggest that even frontier models have significant headroom before they're ready for unsupervised, long-horizon production work. Polymath's public communications emphasize that these numbers aren't just academic. They reflect the gap between today's agents and the autonomous systems enterprises are eager to deploy.

The company's positioning is explicitly tied to the evolving benchmark landscape. As concerns grow about contamination and the limitations of existing coding benchmarks, labs and enterprise tool vendors now need third-party benchmarks to validate agent capabilities, and Horizon-SWE offers a standardized, realistic alternative.

The broader question is whether Polymath can scale its environment generation approach beyond software engineering. Customer service, data analysis, compliance workflows—other domains where long-horizon autonomy would unlock value.

The Regulation Factor

Digital illustration for article section "The Regulation Factor" in "YC's Polymath Labs Tackles AI Agents' Long-Horizon Challenge" - A conceptual, surrealistic composition illustrating the rapid expansion of automated reinforcement l...

The next year will test whether automated RL environment generation can become a sustainable business. Wing VC's analysis suggests the market is there: as labs scale training runs and expand into new domains, they'll need environments faster than human curators can build them.

DeepMind's Genie series, which generates playable 3D worlds from text prompts, and WebWorld, a February 2026 system trained on over a million web interactions, hint at the potential for synthetic environment generation at scale. Polymath's bet is that the same techniques can be adapted for enterprise workflows, where verifiability and realism matter more than visual fidelity.

Regulation may accelerate demand. The EU AI Act's high-risk provisions and enforcement mechanisms take effect August 2, 2026, and general-purpose AI obligations were already in force as of August 2, 2025. In the U.S., California's SB-53 transparency law and Colorado's AI Act (effective June 30, 2026) add compliance pressure. NIST's draft Cyber AI Profile, released in December 2025 with a comment period that closed January 30, ties AI adoption to cybersecurity planning.

All of this points toward a future where enterprises need auditable, verifiable agent behaviors. Precisely what realistic RL environments and benchmarks like Horizon-SWE aim to provide.

The open question is whether the current crop of models will hit diminishing returns on long-horizon tasks. Or whether architectural improvements—hierarchical planning, better memory systems, process-level reward shaping—will push pass rates substantially higher.

HiPER, a hierarchical reinforcement learning system published February 18, shows promise on sparse-reward tasks by explicitly assigning credit across levels of a plan hierarchy. WebWorld's 30+ step horizons and improved WebArena performance suggest that world models trained on diverse interaction data can generalize better than static benchmarks.

If these trends hold, the ~25% pass rates on Horizon-SWE could climb quickly. Though it's unclear whether incremental gains will suffice for true production autonomy or whether a more fundamental rethink is required.

What Founders Need to Know

For founders building agent products, the message is straightforward: long-horizon reliability is now the bottleneck.

Enterprises are moving beyond proof-of-concept demos. Buyers—increasingly business leaders rather than engineering teams—want measurable outcomes tied to KPIs. Integration, governance, and audit trails matter as much as raw capability. The Lyzr report's finding that 62% of enterprises lack a clear starting point suggests a services opportunity around deployment and monitoring, not just tooling.

For AI labs, the shift from copilots to agents is forcing a reallocation of capital. If Wing VC's estimates hold, spending on RL environments could grow 3-5× through 2026, with tens of millions in annual budget. That's real money, redirecting labor and compute away from dataset curation and toward environment engineering.

Polymath's automation thesis—if it works—could compress that timeline and lower costs. But the company is competing against well-funded incumbents like Mechanize and Prime Intellect, plus the in-house efforts of labs that may prefer to build their own infrastructure.

---

The broader arc is clear, even if the specifics remain uncertain. AI agents are moving from narrow, supervised tasks to open-ended, multi-step workflows. The evaluation frameworks that guided the last generation of models are breaking down—contaminated by overuse and misaligned with the behaviors enterprises actually need.

New benchmarks, new environments, new training methodologies are emerging to fill the gap. Polymath Labs is placing an early bet on automated world generation as the scalable solution.

Whether that bet pays off depends on execution, timing, and how quickly the rest of the ecosystem catches up. But one thing seems certain: the 25% problem isn't going away on its own.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • AI Agents Dominate Y Combinator's First Spring 2025 Demo Day
  • Lab-Grown Neurons Playing DOOM Signal Bio-Computing's Commercial Dawn
  • AI Foundation Models Solve Satellite Blind Spots in Climate Monitoring
  • GrazeMate's AI Drones Replace Cowboys, Helicopters for Cattle Herding
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.