Something peculiar appeared in OpenAI's GPT-5.2 model card last December. Amid the familiar metrics—coding benchmarks, mathematical reasoning scores—sat a line most engineers wouldn't have recognized two years earlier: "ARC-AGI-1 (Verified): >90%."
Google DeepMind followed. Then Anthropic. By early this year, according to the ARC Prize 2025 technical report, all four frontier labs had woven this once-obscure benchmark into their public documentation. ARC Prize Verified labels, methodology links, dedicated sections explaining "adaptive thinking" protocols. The pattern felt deliberate.
What started as an academic curiosity in 2019 has quietly become something more consequential: a de facto standard for measuring general reasoning in artificial intelligence. And in an industry increasingly wary of benchmarks that can be gamed and regulators who demand verifiable metrics, the timing of that ascent is worth examining.
A Test You Can't Study For
François Chollet designed the Abstraction and Reasoning Corpus with a specific irritant in mind: AI systems were getting brilliant at mimicking intelligence without actually reasoning. His solution? Colorful grid puzzles that require spotting abstract patterns with minimal examples. No massive training sets to memorize. No shortcuts. Just raw problem-solving.
For years, the benchmark languished at the margins of research. State-of-the-art models barely cracked 33% accuracy on ARC-AGI-1's private test set—a humbling result that suggested something about the difference between pattern matching at scale and genuine adaptability.
Then Mike Knoop entered the picture. The Zapier cofounder teamed up with Chollet to launch the ARC Prize Foundation, a nonprofit dedicated to building what they called "enduring" benchmarks. They put real stakes behind it: an open Kaggle competition with prizes unlocking at 85% accuracy. Money talks, especially in AI research.
The 2024 competition pushed frontier performance from 33% to 55.5%. Researchers threw everything at the problem—program synthesis, test-time training, refinement loops. The grand prize went unclaimed, but something shifted. By mid-2025, paper submissions had nearly doubled to 90 entries. The academic community was paying attention.
When Verification Became Currency

May 17 of last year marked a turning point, though you'd be forgiven for missing it. The foundation released ARC-AGI-2, a harder variant designed to resist brute-force approaches. More sophisticated evaluation protocols. Human testing baselines. Tasks that widened the scoring range and exposed the gap between clever hacks and genuine reasoning.
But the real inflection came last November with ARC Prize Verified. The program tackles what's become AI evaluation's most stubborn problem: benchmark integrity. Companies submit models for testing on hidden datasets. An independent academic panel—Todd Gureckis, Guy Van den Broeck, Melanie Mitchell, Vishal Misra among them—certifies the results. Sponsors including xAI, Google, Nous Research, and Prime Intellect fund development of harder variants.
The timing wasn't accidental. Public datasets get memorized. Solutions leak on GitHub. Private test sets become semi-public through careful API probing. Contamination concerns had poisoned trust in traditional benchmarks just as labs needed to demonstrate reasoning capabilities to investors, regulators, and an increasingly skeptical public. ARC Prize Verified offered something suddenly precious: measurement you could actually trust.
The Numbers Tell Stories

The verified scores reveal more than headline percentages suggest. OpenAI's GPT-5.2, released December 11, claimed over 90% on ARC-AGI-1 but dropped to roughly 53% on the harder variant. Google DeepMind's Gemini 3.1 Pro Thinking High hit 77.1% on ARC-AGI-2. Anthropic's Claude Sonnet 4.6 reached 86.5% on ARC-AGI-1 and 60.4% on the tougher version.
Those gaps matter—perhaps more than the labs initially intended to reveal. They show that reasoning capabilities scale unevenly, that test-time compute budgets create meaningful trade-offs, that harder tasks expose limitations even frontier models can't brute-force their way past. Anthropic's system card specified their scores used 120,000 thinking tokens on "High effort" mode. This kind of transparency around compute resources has become standard practice, possibly because ARC-AGI makes the costs painfully explicit.
The open competition tells a different story entirely. Last year's ARC Prize Grand Prize—unlocking at 85% on ARC-AGI-2's private set—went unclaimed. The top score hovered around 24% for solutions costing $0.20 per task. NVIDIA Kaggle Grandmasters led the public leaderboard at 27.64% using cost-efficient ensembles of 4-billion parameter models. One team, CompressARC, achieved roughly 20% on ARC-AGI-1 with just 76,000 parameters, though at the cost of lengthy inference on consumer GPUs.
The chasm between frontier models and competition entries suggests test-time compute, synthetic data generation, and architectural innovations all contribute meaningfully to performance. It also suggests the benchmark still has considerable headroom. Which is the point.
Regulation Changes the Equation
Two dates loom over the industry's evaluation infrastructure. The EU AI Act's general provisions take effect August 2 of this year. The UK AI Safety Institute's January 2025 report, though later marked "withdrawn" on official government pages, highlighted urgent needs for robust capability testing in the wake of new reasoning models.
The Act doesn't mandate specific benchmarks—it's more subtle than that. It requires conformity assessments and presumptions of compliance through harmonized standards. Labs building frontier models need auditable evidence of what their systems can and cannot do. This creates market pull for third-party verification programs exactly like ARC Prize Verified, without regulators having to specify them directly.
The shift is already visible in how companies present performance data. Google DeepMind's Gemini 3 pages explicitly label results as "ARC Prize Verified" and link to detailed methodology. OpenAI marks its ARC-AGI numbers as "Verified" in product launches. Anthropic's system card attributes scores directly to "the ARC Prize Foundation." This kind of third-party attestation wasn't common in model cards three years ago. Now it feels mandatory.
And it's not just ARC-AGI. A constellation of verification-focused benchmarks emerged recently: Humanity's Last Exam with its own verification ecosystem, SWE-Bench Verified for software repair, OSWorld-Verified for computer use, GDPval for real-world work tasks. The UK AI Safety Institute released Inspect, an open-source evaluation platform, in 2024. The pattern points toward an industry reluctantly accepting that self-reported metrics won't survive regulatory scrutiny.
What This Actually Measures

The benchmark isn't perfect, and its creators don't pretend otherwise. Cost variance remains an issue. Some researchers argue ARC-AGI focuses too narrowly on visual-spatial reasoning. Others point out that 85% still feels arbitrary as a threshold for anything resembling "general intelligence." Early analysis of OpenAI's o3 preview noted significant differences between preview and public versions, plus expenses that scale uncomfortably with desired accuracy.
But those critiques somewhat miss the point. ARC-AGI succeeded not because it perfectly captures intelligence—whatever that means—but because it created a credible, verifiable measuring stick precisely when the industry needed one. François Chollet, in a February profile in Le Monde, stressed that ARC-AGI exists to test adaptability to genuine novelty. The kind of skill acquisition that separates intelligence from statistical correlation. Mike Knoop, speaking on a Sequoia podcast, described the benchmark's motivation as forcing systems to reason with few examples and zero training-set overlap.
The foundation isn't resting. ARC-AGI-3, focused on interactive reasoning, is in development with fresh sponsor backing. The Verified program adds structure as measurement demands intensify. And perhaps most tellingly, research teams are building neurosymbolic approaches, tiny recursive models, and program-synthesis hybrids specifically to crack ARC-AGI's challenges. When a benchmark changes what researchers build, it's moved beyond simple measurement.
As models grow more capable and regulation tightens, the question shifts from "can we measure reasoning?" to "can we trust those measurements?" The fact that every major lab now reports ARC-AGI scores—on hidden test sets, certified by independent academics, with methodology disclosed—suggests the answer is starting to converge. Whether that convergence represents genuine progress or just a new form of consensus theater remains to be seen. But for now, those colorful grid puzzles have become something nobody expected in 2019: required reading for anyone building at the frontier.
