OpenAI's launch of GPT-5.2 came with an odd flourish. Alongside the usual cascade of capability claims sat a table of benchmark scores bearing a single word: "Verified." Not by OpenAI's internal testers, mind you. By a nonprofit most industry watchers had never heard of.
Three months later, Google DeepMind followed suit. Gemini 3 Deep Think scored 84.6% on something called ARC-AGI-2, the company announced—"verified by the ARC Prize Foundation."
In less than two years, a Y Combinator-linked nonprofit co-founded by François Chollet and Mike Knoop has become, improbably, AI's unofficial referee. Today the ARC Prize Foundation verifies intelligence claims from every major frontier lab: OpenAI, Anthropic, Google DeepMind, xAI. According to the foundation's January 2026 technical report, its benchmark—ARC-AGI—has evolved from an academic curiosity into what they're now calling "an industry standard," appearing in model cards and launch announcements across the sector.
This wasn't foreordained. It happened because the AI industry needed something it conspicuously lacked: a test that couldn't be gamed.
The Test That Wouldn't Break
Chollet, creator of the widely adopted Keras framework, first published the Abstraction and Reasoning Corpus in 2019 with a specific ambition. He wanted to measure "general fluid intelligence" through colored grid puzzles requiring few-shot learning and abstraction. The tasks looked deceptively simple. They proved wickedly hard to solve. For years, state-of-the-art models barely moved the needle.
Then the scaling era arrived. Large language models started approaching human-level performance on traditional benchmarks through brute compute and pattern matching. Chollet watched this unfold and concluded that most tests were measuring memorization, not intelligence.
His solution? Make the test harder.
In May 2025, ARC Prize launched ARC-AGI-2, featuring increased task complexity. When labs threw massive compute at ARC-AGI-1—OpenAI's unreleased o3 was the first to achieve a qualifying score under extreme compute conditions—Chollet and Knoop had already formalized the nonprofit in early 2024. By the time they announced ARC-AGI-2 publicly in mid-2025, they'd settled on a strategy: keep the benchmark ahead of the models. Always.
The 2025 competition drew 1,455 teams and 15,154 submissions. All 90 research papers were released open source. The top score on the private leaderboard hit 24% at roughly $0.20 per task. The Grand Prize—which requires 85% accuracy—remained unclaimed.
Meanwhile, frontier labs began reporting their own scores. xAI's Grok 4 managed 15.9% in July 2025. Anthropic's Claude Opus 4.5 hit 37.6% by November. OpenAI's GPT-5.2 reached 52.9% in December.
By February 2026, when Google announced Gemini 3 Deep Think had achieved 84.6%, something had quietly shifted. Labs weren't just running the benchmark anymore. They were coordinating with ARC Prize for verified results on launch day.
Verification as Currency
That "Verified" label carries more weight than it might seem. In an industry where benchmark contamination has become a recurring embarrassment, third-party verification offers plausible deniability and borrowed credibility. OpenAI raised concerns about SWE-bench Verified in February 2026, citing contamination issues. Stanford's HELM benchmarks provide holistic capability dashboards, but they're retrospective assessments. MLCommons' AILuminate focuses on safety rather than raw capability.
ARC Prize found an opening: same-day verification of reasoning capabilities at model launch. According to a 2025 case study by infrastructure provider Vunda, ARC Prize runs concurrent evaluations of 400 tasks across more than seven model providers, integrating results into launch announcements in real time.
It's operational benchmarking as a service, really. Wrapped in academic legitimacy.
The foundation's funding model reflects this hybrid identity. Named donors include Vercel ($30,000), Nous Research ($25,000), Box CEO Aaron Levie ($25,000), and Andreessen Horowitz (amount undisclosed). Enough to run competitions and maintain infrastructure. Not enough, presumably, to compromise independence. The nonprofit structure helps. So does the pedigree: Knoop co-founded Zapier; Chollet created one of the most widely used deep learning frameworks in existence.
In March 2025, ARC Prize submitted a response to the White House Office of Science and Technology Policy's AI Action Plan RFI. The document positioned benchmarking as a U.S. strategic capability and recommended establishing a benchmarking hub within NIST's AI Safety Institute. Whether or not that materializes, the timing was strategic. The EU AI Act enters general application in August 2026, with phased obligations running through 2027. International networks of AI safety institutes—launched in November 2024—are conducting joint testing exercises. Governments want standardized, contamination-resistant evaluations.
ARC Prize is building exactly that.
The Moving Target

The benchmark keeps evolving because it has to. On March 5, 2026, ARC Prize published ARC-TGI, introducing 461 human-validated task generators designed to prevent overfitting and enable scalable sampling. Instead of hand-authoring a fixed set of puzzles, the generators produce new variations that maintain core reasoning requirements while preventing memorization.
Three weeks later—March 25, 2026—ARC-AGI-3 launches. It represents a fundamental departure from static puzzles. The new benchmark places agents into interactive environments with more than 1,000 "games" or levels where they must explore, learn, plan, and adapt. There's a developer harness. APIs. Scorecards. It's closer to how humans actually demonstrate intelligence: through open-ended problem-solving in dynamic environments.
This mirrors a broader shift across AI evaluation, where benchmarks are moving from static knowledge tests to agentic scenarios. Scale's MCP-Atlas (launched December 19, 2025) tests models on real tool servers. Mercor's APEX-Agents (published January 20, 2026) evaluates cross-application professional tasks. OpenAI's BrowseComp (April 2025) measures web browsing agents. OSWorld-Verified tests computer use through GUI interactions.
ARC Prize isn't alone in this transition. But it may be furthest ahead in one specific dimension: making the transition illegible to gaming. Interactive environments resist contamination better than static datasets. Generated tasks resist memorization better than hand-authored ones.
The question is whether ARC can maintain that edge as models grow more capable.
The Marketing Problem

Current ARC-AGI-2 scores tell a particular story. Google's Gemini 3 Deep Think leads at 84.6%, tantalizingly close to the 85% Grand Prize threshold. OpenAI's GPT-5.2 sits at 52.9%. Anthropic's Claude Opus 4.6, based on secondary reports from February 2026, appears to have reached approximately 68.8%. These numbers are verified, timestamped, comparable.
They're also incomplete.
ARC-AGI-2 measures abstract reasoning on grid puzzles. It doesn't capture everything—coding ability, long-context understanding, multimodal reasoning, real-world task completion. Labs now report a constellation of scores across multiple benchmarks: GPQA Diamond for expert-level science questions, AIME 2025 for competition math, MMMU-Pro for multimodal understanding, Humanity's Last Exam for closed-ended knowledge questions.
The risk? That verified ARC scores become a marketing proxy for general intelligence rather than what they actually measure: sample-efficient learning and abstract pattern recognition. Chollet's original critique of scaling-only approaches—that they don't produce true AGI—remains central to ARC's philosophy. The benchmark is designed to fail models that rely purely on scale and memorization.
Whether that philosophy survives contact with an industry increasingly focused on agentic capabilities and real-world performance is, well, unclear. ARC-AGI-3's shift to interactive environments suggests the foundation understands this tension. But the transition from puzzles to agents also makes the benchmark harder to explain, harder to verify, potentially more subjective.
Staying Ahead

For now, ARC Prize occupies valuable territory. Independent enough to be credible. Operational enough to be useful. Difficult enough to remain relevant as models improve. The "Verified" label appearing in launch announcements from OpenAI and Google signals that frontier labs consider third-party validation worth the coordination cost.
That could change, of course. Labs could build internal evaluation teams that match ARC's credibility. Governments could mandate standardized testing through institutions like NIST's AI Safety Institute or the UK's AISI. Competitors could emerge with better benchmarks or faster verification pipelines.
Or the benchmark treadmill could simply continue. ARC-AGI-3 gives way to ARC-AGI-4, which gives way to ARC-AGI-5, each iteration racing to stay ahead of model capabilities. The foundation's technical report claims they've achieved industry standard status. That status depends on maintaining a test that's hard enough to matter but solvable enough to show progress.
The 2026 competition launches with ARC-AGI-3 and the new generator framework. If frontier models crack 85% on ARC-AGI-2 before year's end—Google is already close—Chollet and Knoop will need to move fast.
The test only works if it can't be beaten easily. Once a nonprofit becomes the official scorekeeper for AGI, staying ahead of the competition stops being a choice. It becomes an existential requirement.
