A small nonprofit is quietly reshaping how the world's most powerful AI companies measure progress toward artificial general intelligence. The question is whether their benchmark can survive its own success.
When Google DeepMind updated its Gemini 3.1 Pro model card this past February, eagle-eyed observers noticed something new tucked among the usual performance metrics. Between MMLU scores and HumanEval results sat an entry that would have meant little to most readers: ARC-AGI-2, adorned with a small "ARC Prize Verified" badge.
DeepMind wasn't breaking new ground—it was catching up. By year's end, all four major AI labs (Anthropic, Google DeepMind, OpenAI, and xAI) had begun featuring ARC-AGI scores on their public-facing model cards, according to the nonprofit that administers the test. What started as an academic exercise in 2019 has become perhaps the closest thing this fractious industry has to a shared measure of genuine machine intelligence.
In a field where benchmark manipulation and breathless self-promotion are something of a contact sport, that's worth examining.
An Unlikely Standards Body
The organization behind this quiet victory is the ARC Prize Foundation, a nonprofit led by Mike Knoop—cofounder of the automation company Zapier—and François Chollet, the Google researcher who built Keras and later published a paper titled "On the Measure of Intelligence" that still gets cited in AI circles today. Greg Kamradt, who spent time at Salesforce, serves as president.
The foundation's structure is unconventional, maybe deliberately so. It went through Y Combinator's Winter 2026 cohort with just four employees, an unusual résumé line for what is essentially a standards-setting body. Then again, Knoop and Chollet both come from the startup world, where credibility often matters less than velocity.
Their donor list reflects the blurred lines of modern AI funding. xAI and Google appear as lab sponsors. Individual backers include economist Tyler Cowen, DoorDash's Andy Fang, HubSpot's Dharmesh Shah, and Box CEO Aaron Levie—a roster that suggests both technical credibility and deep-pocketed belief in the mission. The academic advisory panel features Melanie Mitchell, Guy Van den Broeck, Todd Gureckis, and Vishal Misra, lending scholarly heft to what might otherwise look like another Silicon Valley enthusiasm project.
Knoop tends to frame the foundation's work in the language of tech optimism: "increase the idea rate" around AGI. But the execution leans heavily academic. Prize winners must open-source their solutions, for one thing. And the benchmark itself is explicitly designed to resist the brute-force scaling strategies that have defined AI progress since the transformer revolution.
Why This Test Is Different
Most benchmarks measure what a model knows. ARC-AGI measures how efficiently it learns something it's never seen before.
That distinction is more than semantic. Chollet's original paper introduced what he called "skill-acquisition efficiency"—essentially, the ability to extract abstract rules from minimal examples and apply them elsewhere. The test presents visual puzzles that most humans solve in seconds but that have confounded even cutting-edge language models. Each task requires spotting a pattern from a handful of demonstrations, then generalizing it to a novel case. No amount of pre-training helps here. Either your system can reason its way through, or it can't.
"Easy for humans, hard for AI," Chollet has said more than once, and the results bear him out. Human test-takers score near 85% on the original ARC-AGI-1 evaluation set. When the foundation tested major reasoning models last June on the newer ARC-AGI-2 benchmark, most scored in the single digits. Not 10% or 15%. Single digits.
The design constraints are intentional, almost severe. Tests run offline—no internet, no massive databases to query mid-answer. Compute is artificially limited to force genuine reasoning rather than exhaustive trial-and-error. The benchmark evolves, too. After ARC-AGI-1 gained traction in 2024, the foundation released ARC-AGI-2 in March 2025, explicitly engineered to be harder for machines while remaining straightforward for humans. The semi-private and private test sets—120 tasks apiece—include only puzzles that at least two humans can solve within two attempts, a safeguard against unfairness.
What Last Year's Competition Revealed

The 2025 Kaggle competition, which ran from March through November, laid bare just how far the field has to go. Some 1,455 teams submitted over 15,000 entries. The top private-set score—posted by a team called NVARC—reached 24.03% on ARC-AGI-2, at a per-task cost of roughly 20 cents.
Pause on that number. The winning approach, open-sourced per contest rules, solved less than a quarter of the puzzles that most humans handle without effort. This came after nine months of intense work, 90 research paper submissions, and a collective prize pool of $1 million. (The $700,000 grand prize, which requires hitting 85% accuracy within compute limits, remains unclaimed.)
The best verified commercial model as of December? Claude Opus 4.5 (Thinking, 64k context), managing 37.6% on ARC-AGI-2. Poetiq's refinement harness, built atop Gemini 3 Pro, hit 54%—though that falls into the "Model Refinements" category, meaning it's really an orchestrated system rather than a single model working alone.
What emerged from the competition wasn't just incremental progress. Winners converged on what the foundation now calls "refinement loops"—iterative program synthesis and verification systems that treat language models as reasoning engines rather than answer machines. It represents a shift from scale-driven pre-training to test-time adaptation. Which suggests, perhaps, that the next leap in AI won't come from bigger datasets but from smarter inference-time architectures.
Interactive Reasoning Arrives This Spring
Maybe sensing that even their toughest static puzzles will eventually fall to clever enough systems, the foundation is launching something different on March 25: ARC-AGI-3, a fully interactive benchmark spanning over 1,000 levels across more than 150 novel game environments.
The developer toolkit shipped in late January. Early preview results? Humbling. Where ARC-AGI-2 tested pattern recognition, ARC-AGI-3 demands exploration, planning, memory, goal acquisition, and something approaching alignment—the ability to figure out what the benchmark wants without being told. Graph-based exploration baselines solve a median of 30 out of 52 preview levels. Language models, which dominate static Q&A tasks, struggle badly with the open-ended nature of interactive challenges.
The foundation is also introducing action-efficiency scoring—not merely whether you solve the task, but how many moves it takes. That matters in practical terms. An AI system that burns through 10,000 steps to accomplish what a human does in 50 isn't particularly useful, even if it technically succeeds.
A Crowded Field

ARC-AGI doesn't operate in isolation. The evaluation landscape has grown crowded lately, each framework emphasizing different facets of intelligence.
CAIS and Scale AI published Humanity's Last Exam (HLE) in Nature in late January—2,500 expert-crafted questions designed to be unsearchable and ferociously difficult. Stanford's HELM project offers standardized, multi-scenario comparisons. MLCommons' MLPerf has added reasoning-oriented tests. The UK AI Safety Institute open-sourced its Inspect platform, consolidating agent benchmarks like GAIA and SWE-Bench. NIST's GenAI program continues running active evaluation rounds.
Yet none have achieved quite the model-card traction of ARC-AGI—that crucial real estate where labs showcase capabilities to investors, customers, and increasingly, regulators. Perhaps the nonprofit structure lends credibility in a way corporate-backed benchmarks can't. Perhaps it's the "Verified" tag that DeepMind and others now display beside their scores, signaling third-party validation in an industry rife with self-reported metrics. Or perhaps Chollet's original insight—measure learning efficiency, not accumulated knowledge—simply cuts closest to what people mean when they say "general intelligence."
The Policy Play

Knoop has been vocal about the foundation's policy ambitions. He submitted recommendations to the White House's AI Action Plan last July, arguing for a federal hub for AI benchmarking, diverse independent evaluators, and U.S. leadership in global evaluation standards. NIST, the recently renamed CAISI (formerly the U.S. AI Safety Institute), the Department of Energy, and the National Science Foundation are already tasked with leading evaluation science under current administration frameworks.
Across the Atlantic, the EU AI Act's standardization pathway is developing harmonized standards. The first AI quality management system standard (prEN 18286) entered public inquiry last October. Enforcement begins in August. The UK's AISI continues expanding evaluation infrastructure. Policy pressure is building, in other words, for exactly the kind of transparency the ARC Prize Foundation has championed—standardized reporting of both accuracy and efficiency, model-card disclosures, third-party verification.
Whether ARC-AGI-3's interactive challenges reset the field's understanding of current capabilities—or whether they too eventually fall to increasingly sophisticated refinement systems—remains an open question. But for now, when AI labs want to make claims about general intelligence, they're reaching for a benchmark run by a nonprofit with four employees and backed by researchers who believe rigor matters more than hype.
The March 25 launch will be watched closely. Not just for the scores themselves, but for what they reveal about how far we really are—or aren't—from machines that learn the way humans do. If nothing else, it should make for more honest conversations than the industry typically allows itself.
