Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 3, 2026

DesignVerse raises $5.5M to automate enterprise software

DesignVerse raises $5.5M to automate enterprise software
Ai AutomationEnterprise Software+3
SaaS iconSaaSOctober 3, 2026

OSCP raises $6M for GPS-free navigation sensors

OSCP raises $6M for GPS-free navigation sensors
PhotonicsSensor Tech+3
SaaS iconSaaSMarch 5, 2026

ArmorCode Raises $16M to Scale AI Security Governance Platform

ArmorCode Raises $16M to Scale AI Security Governance Platform
Ai GovernanceCyber Security+3
Climate / Social Tech iconClimate / Social TechMarch 5, 2026

Physics-Informed AI Emerges as New Weapon for Building Energy Crisis

Physics-Informed AI Emerges as New Weapon for Building Energy Crisis
Building Energy ManagementData Center Efficiency+3
SaaS iconSaaS
March 5, 2026
YcAi BenchmarkingAgi ResearchArtificial Intelligence

How This YC Startup's AI Benchmark Stumps OpenAI and Anthropic

ARC Prize Foundation, a YC W26 company, built the industry's toughest AGI test—used by every major AI lab. Now they're launching interactive challenges to combat model overfitting.

How This YC Startup's AI Benchmark Stumps OpenAI and Anthropic

OpenAI had its pick of achievements when it came time to debut the o3 model in December 2024. The company could have highlighted performance on MMLU, that venerable benchmark most frontier systems now breeze through. Or any number of other standardized tests that have become routine victories.

Instead, OpenAI chose to spotlight a single metric: ARC-AGI.

The decision said more than any press release could. ARC-AGI, a fiendishly difficult series of visual puzzles designed to measure genuine reasoning rather than memorized patterns, has become something of an obsession among the major labs. It's the test that refuses to be gamed. The one that even the most sophisticated models still mostly fail.

And now, the nonprofit behind that benchmark is betting that static tests—no matter how clever—have reached their expiration date.

The Test No One Can Ace

ARC Prize Foundation operates in a peculiar space. Part standards body, part research provocateur, the organization exists to build what co-founder Mike Knoop calls "AI benchmarks that measure general intelligence and inspire new ideas." Knoop, who previously co-founded Zapier, runs ARC alongside François Chollet, the Google engineer who created Keras and has emerged as one of the field's most vocal skeptics of pure scaling approaches.

Their flagship creation, the ARC-AGI benchmark series, has become the reasoning test that labs can't afford to skip. When ARC-AGI-2 launched in March 2025, frontier models stumbled badly. OpenAI's o1 and DeepSeek's R1—systems designed explicitly for complex reasoning—scored in the low single digits. The benchmark wasn't merely difficult; it was exposing something fundamental about how current models actually think.

By December, after a competition that attracted 1,455 teams submitting more than 15,000 entries via Kaggle, the best private evaluation score had crawled to just 24%. That winning entry, from a team calling itself NVARC, managed this feat at a cost of 20 cents per task. Not terrible, but hardly the stuff of artificial general intelligence.

Commercial systems fared only marginally better. Anthropic's Claude Opus 4.5, equipped with 64,000 thinking tokens and running in a verified configuration, hit 37.6% accuracy. The cost per task: $2.20. A refinement harness built by a team called Poetiq, running on Google's Gemini 3 Pro, eventually reached 54%—at roughly $30 per puzzle. Greg Kamradt, ARC's president, observed in March that ARC-AGI-2 "significantly raises the bar for AI." Perhaps an understatement.

The benchmark's credibility is evident in how labs have adopted it, sometimes grudgingly. OpenAI, Anthropic, Google DeepMind, and xAI all reported ARC-AGI results on their 2025 model cards, according to ARC's technical documentation. It has become an unofficial litmus test—the assessment you cite when claiming your system can actually reason.

What Makes Intelligence So Hard to Pin Down

Chollet introduced the ARC-AGI concept back in 2019 with a specific thesis: truly intelligent systems should exhibit what he calls "fluid intelligence." Not the regurgitation of training data. Not pattern matching at scale. But the ability to acquire new skills efficiently when confronted with genuinely novel problems.

The puzzles themselves are deceptively simple—grid-based visual challenges requiring minimal prior knowledge. Each one presents a kind of mini-mystery, solvable by a human with a few minutes of thought but stubbornly resistant to models trained on internet-scale data. The core insight is almost philosophical: intelligence should generalize. A system shouldn't need to have encountered every conceivable pattern during training to figure out new ones in the wild.

That design principle, however elegant, now faces an uncomfortable reality. ARC's December analysis of competition results flagged what they termed "knowledge-dependent overfitting." Frontier models, after ingesting trillions of tokens during pretraining, appear to have absorbed some of the statistical regularities in ARC-AGI's puzzle formats. Private test sets help, but they can't entirely prevent models from leveraging patterns encoded deep in their weights.

Chollet, who left Google in November 2024 to work on ARC full-time, has long maintained that scaling LLMs alone won't produce AGI. The 2025 results vindicate his skepticism, at least to a point. Winning solutions didn't come from bigger models. They came from clever test-time refinement loops—techniques that let models iteratively improve answers by trading compute at inference time for better reasoning. Architectural ingenuity over raw scale.

When Static Tests Stop Working

Digital illustration for article section "When Static Tests Stop Working" in "How This YC Startup's AI Benchmark Stumps OpenAI and Anthropic"

Which brings us to ARC-AGI-3, slated for release sometime early this year. The new benchmark abandons static puzzles entirely. Instead, it offers roughly 100 interactive game-like environments where agents must explore, plan, manage working memory, acquire goals, and demonstrate something approximating alignment. Performance gets measured using Relative Human Action Efficiency—essentially, how many actions an AI system takes compared to a human solving the same task.

The shift is both practical and ideological. Practically speaking, interactive environments are far harder to overfit. Models can't memorize solutions when each environment demands real-time adaptation. Ideologically, the move aligns with a growing consensus that static benchmarks have saturated. The 2025 Stanford AI Index noted that older assessments like MMLU and GSM8K have become less informative as models approach or exceed human performance through a murky combination of genuine capability and dataset contamination.

Labs and researchers have responded by turning to harder, more specialized tests. GPQA for scientific reasoning. SWE-bench Verified for software agents. Humanity's Last Exam for knowledge-heavy safety scenarios. ARC-AGI-3 fits into this wave but with a distinct focus on general intelligence rather than domain expertise.

The 2025 International AI Safety Report, cited by governments and safety institutes worldwide, explicitly called out reproducibility and contamination concerns in current evaluations. Interactive benchmarks, by design, address both.

The Evaluation Arms Race

ARC Prize isn't operating in a vacuum. The evaluation ecosystem has exploded in the past two years, fragmenting into dozens of specialized efforts. MLCommons launched AILuminate, a safety-focused suite. NIST is running generative AI pilot evaluations. Stanford's HELM project maintains sprawling leaderboards. Scale AI introduced a commercial evaluation platform in April 2025, betting that enterprise buyers now demand rigorous assessment tools before writing checks.

Most of these efforts target specific modalities or safety risks. ARC Prize occupies a narrower niche—arguably more fundamental. It's not asking "Can your model diagnose diseases?" or "Can it write secure code?" It's asking the more elemental question: Can it think?

That framing has resonated with researchers hungry for alternatives to the LLM scaling paradigm. The 2025 ARC Prize winners demonstrated that small, recursive models with strong test-time adaptation can outperform far larger systems. One paper-award recipient, the Tiny Recursive Model team, achieved 45% on ARC-AGI-1 and roughly 8% on ARC-AGI-2 using a model with just seven million parameters. These results suggest architectural choices and inference-time search matter as much as—possibly more than—pretraining scale.

Knoop has emphasized this point in interviews, including a 2024 Sequoia podcast. ARC Prize, he argues, exists to "drive new ideas, not just bigger LLMs." The 2025 competition seems to have succeeded on that front. Refinement harnesses, iterative reasoning loops, and test-time adaptation emerged as dominant techniques. Multiple top teams converged independently on similar strategies, suggesting they'd stumbled onto something real.

The Next Frontier, Maybe

Whether ARC-AGI-3's interactive environments will prove more or less stubborn than static puzzles remains an open question. The benchmark's preview materials describe scenarios requiring agents to maintain working memory, execute multi-step plans, and handle tasks where goals aren't explicitly stated upfront. Early documentation suggests a focus on alignment—agents must infer the correct objective from context, not just maximize a reward signal.

Digital illustration for article section "The Next Frontier, Maybe" in "How This YC Startup's AI Benchmark Stumps OpenAI and Anthropic"

The timing aligns with broader institutional momentum toward standardized evaluation. The 2025 International AI Safety Report, NIST initiatives, and pre-deployment frameworks from AI Safety Institutes in the U.S. and U.K. all point toward governments demanding more rigorous testing regimes before frontier systems reach deployment. ARC's March submission to the White House Office of Science and Technology Policy advocated for benchmark-driven research, a theme that's only gained traction since.

For AI labs, the writing on the wall is clear enough. Leaderboard gaming and dataset memorization have diminishing returns. Interactive, adversarial, and continuously evolving benchmarks will define the next generation of capability assessments. ARC Prize Foundation—a Y Combinator-backed nonprofit led by two of the most credible figures in applied AI—finds itself well-positioned to set the terms of that evolution.

The real question for 2026 is whether any model will crack 75% on ARC-AGI-3, or whether interactive reasoning proves just as stubborn a barrier as the static puzzles before it. History suggests betting against progress is unwise. But then again, history also suggests that genuine intelligence—the kind that generalizes across truly novel problems—remains frustratingly out of reach.

For now, at least, the puzzles are winning.

More stories

  • DesignVerse raises $5.5M to automate enterprise software
  • OSCP raises $6M for GPS-free navigation sensors
  • ArmorCode Raises $16M to Scale AI Security Governance Platform
  • Physics-Informed AI Emerges as New Weapon for Building Energy Crisis
  • How iLabService Built 7,700-Lab Network After $15M Series A
  • Command Code Raises $5M for AI Coding Agent That Learns Your Style
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.