The metric appeared quietly at first—tucked into Google DeepMind's February announcement of Gemini 3 Deep Think like any other performance number: 84.6% on something called ARC-AGI-2, verified by the ARC Prize Foundation. A week later, Anthropic devoted an entire table in Claude 4.6's system card to ARC-AGI-2 scores. By then, OpenAI had already featured the benchmark prominently in its GPT-5.2 rollout the previous December.
Somewhere along the way, without much fanfare, a nonprofit that emerged from Y Combinator's Winter 2026 cohort had accomplished what dozens of academic benchmarks couldn't: it got every major AI lab to care about the same test.
That's unusual. The AI industry churns out benchmarks the way Hollywood produces sequels—constantly, competitively, and often with diminishing returns. Most fade into obscurity, cited in a few papers before researchers move on. But the ARC Prize Foundation, co-founded by François Chollet (the creator of Keras and the original ARC benchmark) and Mike Knoop (Zapier's co-founder), built something stickier. Their benchmark measures what models haven't memorized. Which turns out to be exactly what the industry needed to measure—even if it didn't know it yet.
Colored Grids and Fluid Intelligence
The premise is deceptively simple. ARC-AGI presents colored-grid transformation puzzles that require abstract reasoning to solve. No massive training corpus exists for these tasks. No leaked test sets circulate on GitHub. According to the foundation's technical reports, the benchmark aims to measure "fluid intelligence"—the ability to adapt to genuinely novel situations using minimal examples, much the way humans fumble through unfamiliar problems.
This matters more than it might seem. As AI models began posting superhuman scores on traditional benchmarks—reading comprehension tests, math problems, coding challenges—an uncomfortable question surfaced: what do those numbers actually mean? Are models reasoning, or just retrieving patterns they've seen thousands of times during training?
ARC-AGI doesn't let models hide behind memorization. When ARC-AGI-2 launched in March 2025, designed with stricter anti-contamination measures and carefully calibrated human baselines, most leading reasoning models scored in the single digits. The gap was... clarifying.
The Competition Strategy
The foundation's path to relevance began with money. In 2024, they launched a $1 million challenge that attracted 1,430 teams and nearly 18,000 entries. The top solution on the private evaluation set reached 55.5%, up from 33% at the year's start.
Nobody claimed the grand prize.
That failure—the unclaimed million—might have been the best thing that could have happened. It signaled that ARC-AGI wasn't just another benchmark to be gamed and forgotten. The problems were genuinely hard.
The 2025 competition, run on the tougher ARC-AGI-2 benchmark, drew similar numbers: 1,455 teams, over 15,000 entries. The top Kaggle submission managed 24% on the private split at a cost of $0.20 per task. The best verified commercial model as of last December was Claude Opus 4.5 (the Thinking variant, 64k context window) at 37.6%, though at roughly $2.20 per task. A refinement pipeline built by a team called Poetiq using Gemini 3 Pro eventually hit 54% verified—at approximately $30 per task.
These competitions did more than crown winners and pay out prizes. They established methodology. They created a research community that spoke a common language. And crucially, they demonstrated to frontier labs that ARC-AGI exposed meaningful capability gaps in their models—gaps that traditional benchmarks had stopped revealing months or years earlier.
According to the ARC Prize 2025 technical report published in January, all four major frontier labs—Anthropic, Google DeepMind, OpenAI, and xAI—reported ARC-AGI scores in their public model cards during 2025. That cross-lab acknowledgment marked a quiet turning point. When competitors agree on a metric, it's either collusion or validation.
In this case, probably validation.
Evolution Under Pressure

The benchmark itself hasn't stood still. ARC-AGI-1, which originated from Chollet's 2019 research, became heavily studied by 2024. Researchers developed increasingly sophisticated approaches to attack it. Concerns about saturation on public splits mounted—the curse of any successful benchmark.
ARC-AGI-2 addressed this with new task distributions, separate public, semi-private, and private splits, and stricter contamination controls. An arxiv paper published in May 2025 outlined the calibration process used to ensure tasks genuinely tested reasoning rather than pattern matching. Whether it succeeded is still being debated, though early results suggested the new version was indeed harder.
Now the foundation is preparing ARC-AGI-3, scheduled for a March 25 launch. The developer toolkit released in late January signals a significant shift: from static puzzles to interactive, game-like environments. According to foundation documentation, version 3 will test exploration, planning, memory, and long-horizon goal achievement. The benchmark will score models on efficiency relative to human performance, not just accuracy.
That's an ambitious leap. Whether it works remains to be seen.
The foundation's approach to distribution and accessibility as it scales suggests serious thinking about infrastructure and reach. Or perhaps just hedging against the possibility that hosting infrastructure gets expensive.
Why Frontier Labs Actually Listened
Plenty of benchmarks get published. Most get ignored. ARC-AGI succeeded because it filled a gap that had started to feel embarrassing: the distance between what people meant when they said "reasoning" and what existing tests actually measured.
Chollet has long argued—sometimes to the point of repetition—that LLM scaling alone won't yield AGI. He emphasizes fluid intelligence and program synthesis. That philosophical stance positions ARC-AGI as a counterweight to benchmarks that essentially reward scale and memorization. When a model scores 90%+ on most academic tests but can't break 50% on ARC-AGI-2, that tells developers something specific about their architecture's limitations.
The foundation's analysis of OpenAI's o3 model, published in April 2025, illustrated this neatly. The o3-preview's high-compute run on ARC-AGI-1 back in December 2024 showed strong performance. But its results on ARC-AGI-2 were markedly lower—evidence, the foundation argued, that the newer version was measuring something genuinely distinct.
Knoop's analysis of the 2025 competition results identified what he called "the year of the refinement loop," highlighting deep learning-guided program synthesis and test-time adaptation as emerging meta-strategies. These aren't just competition techniques. They're research directions that labs are actively pursuing, with funding and headcount attached.
The Regulatory Tailwind
Timing helped. The EU AI Act's regulatory framework has been developing, with obligations for general-purpose AI providers rolling out through this coming August. The US AI Safety Institute and the International Network of AI Safety Institutes, both launched in November 2024, are working to establish shared principles for risk assessment.
Regulators need standardized, reproducible evaluations. They need numbers they can audit. In March 2025, the ARC Prize Foundation submitted recommendations to the US Office of Science and Technology Policy's AI Action Plan, advocating for benchmark-based capability tracking and procurement criteria.
The foundation's verified leaderboard system—which distinguishes between self-reported scores and third-party verified results—aligns perfectly with what regulators want: auditable claims about AI capabilities. Anthropic's February system card demonstrates how labs are already adapting, including multiple verified benchmark scores alongside their performance claims.
Whether regulators will actually mandate specific benchmarks remains unclear. But if they do, ARC-AGI has positioned itself well.
What Comes Next

The ARC-AGI-3 launch later this month will test whether the benchmark can maintain its resistance to saturation as it shifts to interactive environments. The foundation has committed to running annual competitions until ARC is "defeated," according to statements from Knoop in past results announcements. That's either admirably stubborn or a clever marketing strategy. Maybe both.
Perhaps the real test is whether ARC-AGI remains relevant as AI capabilities advance in unexpected directions. The benchmark's evolution from static puzzles to agentic tasks suggests the foundation understands that measuring intelligence requires adapting the yardstick as models improve. Or at least staying one step ahead.
For now, when the world's leading AI labs want to claim progress toward AGI, they're increasingly pointing to a four-person nonprofit's colored-grid puzzles. That's either a remarkable validation of Chollet's decade-long argument about fluid intelligence, or a sign that the industry is still groping for better ways to measure what actually matters.
Probably both. These things usually are.
