Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechMarch 25, 2026

Immutrin Raises $87M Series A for Antibody Therapy to Clear Amyloid

Immutrin Raises $87M Series A for Antibody Therapy to Clear Amyloid
BiotechDrug Development+2
Healthtech & Biotech iconHealthtech & BiotechMarch 25, 2026

DrHouse Raises $3.5M Seed to Scale 24/7 Telemedicine Platform

DrHouse Raises $3.5M Seed to Scale 24/7 Telemedicine Platform
TelehealthHealthtech+2
SaaS iconSaaS
March 25, 2026
Ai BenchmarkingAgi ResearchAi GovernanceAi

How ARC-AGI Became AI's Gold Standard for Intelligence Testing

All frontier AI labs now report scores on ARC-AGI, the nonprofit benchmark measuring general reasoning. As regulation looms, its new verification program tackles AI's biggest evaluation problem.

How ARC-AGI Became AI's Gold Standard for Intelligence Testing

Something peculiar appeared in OpenAI's GPT-5.2 model card last December. Amid the familiar metrics—coding benchmarks, mathematical reasoning scores—sat a line most engineers wouldn't have recognized two years earlier: "ARC-AGI-1 (Verified): >90%."

Google DeepMind followed. Then Anthropic. By early this year, according to the ARC Prize 2025 technical report, all four frontier labs had woven this once-obscure benchmark into their public documentation. ARC Prize Verified labels, methodology links, dedicated sections explaining "adaptive thinking" protocols. The pattern felt deliberate.

What started as an academic curiosity in 2019 has quietly become something more consequential: a de facto standard for measuring general reasoning in artificial intelligence. And in an industry increasingly wary of benchmarks that can be gamed and regulators who demand verifiable metrics, the timing of that ascent is worth examining.

A Test You Can't Study For

François Chollet designed the Abstraction and Reasoning Corpus with a specific irritant in mind: AI systems were getting brilliant at mimicking intelligence without actually reasoning. His solution? Colorful grid puzzles that require spotting abstract patterns with minimal examples. No massive training sets to memorize. No shortcuts. Just raw problem-solving.

For years, the benchmark languished at the margins of research. State-of-the-art models barely cracked 33% accuracy on ARC-AGI-1's private test set—a humbling result that suggested something about the difference between pattern matching at scale and genuine adaptability.

Then Mike Knoop entered the picture. The Zapier cofounder teamed up with Chollet to launch the ARC Prize Foundation, a nonprofit dedicated to building what they called "enduring" benchmarks. They put real stakes behind it: an open Kaggle competition with prizes unlocking at 85% accuracy. Money talks, especially in AI research.

The 2024 competition pushed frontier performance from 33% to 55.5%. Researchers threw everything at the problem—program synthesis, test-time training, refinement loops. The grand prize went unclaimed, but something shifted. By mid-2025, paper submissions had nearly doubled to 90 entries. The academic community was paying attention.

When Verification Became Currency

Digital illustration for article section "When Verification Became Currency" in "How ARC-AGI Became AI's Gold Standard for Intelligence Testing" - A conceptual, minimalist close-up of a heavy, elegantly minted metallic token representing verificat...

May 17 of last year marked a turning point, though you'd be forgiven for missing it. The foundation released ARC-AGI-2, a harder variant designed to resist brute-force approaches. More sophisticated evaluation protocols. Human testing baselines. Tasks that widened the scoring range and exposed the gap between clever hacks and genuine reasoning.

But the real inflection came last November with ARC Prize Verified. The program tackles what's become AI evaluation's most stubborn problem: benchmark integrity. Companies submit models for testing on hidden datasets. An independent academic panel—Todd Gureckis, Guy Van den Broeck, Melanie Mitchell, Vishal Misra among them—certifies the results. Sponsors including xAI, Google, Nous Research, and Prime Intellect fund development of harder variants.

The timing wasn't accidental. Public datasets get memorized. Solutions leak on GitHub. Private test sets become semi-public through careful API probing. Contamination concerns had poisoned trust in traditional benchmarks just as labs needed to demonstrate reasoning capabilities to investors, regulators, and an increasingly skeptical public. ARC Prize Verified offered something suddenly precious: measurement you could actually trust.

The Numbers Tell Stories

Digital illustration for article section "The Numbers Tell Stories" in "How ARC-AGI Became AI's Gold Standard for Intelligence Testing" - A minimalist, conceptual close-up of an elegant, vintage analog split-flap display resting on a smoo...

The verified scores reveal more than headline percentages suggest. OpenAI's GPT-5.2, released December 11, claimed over 90% on ARC-AGI-1 but dropped to roughly 53% on the harder variant. Google DeepMind's Gemini 3.1 Pro Thinking High hit 77.1% on ARC-AGI-2. Anthropic's Claude Sonnet 4.6 reached 86.5% on ARC-AGI-1 and 60.4% on the tougher version.

Those gaps matter—perhaps more than the labs initially intended to reveal. They show that reasoning capabilities scale unevenly, that test-time compute budgets create meaningful trade-offs, that harder tasks expose limitations even frontier models can't brute-force their way past. Anthropic's system card specified their scores used 120,000 thinking tokens on "High effort" mode. This kind of transparency around compute resources has become standard practice, possibly because ARC-AGI makes the costs painfully explicit.

The open competition tells a different story entirely. Last year's ARC Prize Grand Prize—unlocking at 85% on ARC-AGI-2's private set—went unclaimed. The top score hovered around 24% for solutions costing $0.20 per task. NVIDIA Kaggle Grandmasters led the public leaderboard at 27.64% using cost-efficient ensembles of 4-billion parameter models. One team, CompressARC, achieved roughly 20% on ARC-AGI-1 with just 76,000 parameters, though at the cost of lengthy inference on consumer GPUs.

The chasm between frontier models and competition entries suggests test-time compute, synthetic data generation, and architectural innovations all contribute meaningfully to performance. It also suggests the benchmark still has considerable headroom. Which is the point.

Regulation Changes the Equation

Two dates loom over the industry's evaluation infrastructure. The EU AI Act's general provisions take effect August 2 of this year. The UK AI Safety Institute's January 2025 report, though later marked "withdrawn" on official government pages, highlighted urgent needs for robust capability testing in the wake of new reasoning models.

The Act doesn't mandate specific benchmarks—it's more subtle than that. It requires conformity assessments and presumptions of compliance through harmonized standards. Labs building frontier models need auditable evidence of what their systems can and cannot do. This creates market pull for third-party verification programs exactly like ARC Prize Verified, without regulators having to specify them directly.

The shift is already visible in how companies present performance data. Google DeepMind's Gemini 3 pages explicitly label results as "ARC Prize Verified" and link to detailed methodology. OpenAI marks its ARC-AGI numbers as "Verified" in product launches. Anthropic's system card attributes scores directly to "the ARC Prize Foundation." This kind of third-party attestation wasn't common in model cards three years ago. Now it feels mandatory.

And it's not just ARC-AGI. A constellation of verification-focused benchmarks emerged recently: Humanity's Last Exam with its own verification ecosystem, SWE-Bench Verified for software repair, OSWorld-Verified for computer use, GDPval for real-world work tasks. The UK AI Safety Institute released Inspect, an open-source evaluation platform, in 2024. The pattern points toward an industry reluctantly accepting that self-reported metrics won't survive regulatory scrutiny.

What This Actually Measures

Digital illustration for article section "What This Actually Measures" in "How ARC-AGI Became AI's Gold Standard for Intelligence Testing" - A minimalist and conceptual still life representing the complex measurement of visual-spatial reason...

The benchmark isn't perfect, and its creators don't pretend otherwise. Cost variance remains an issue. Some researchers argue ARC-AGI focuses too narrowly on visual-spatial reasoning. Others point out that 85% still feels arbitrary as a threshold for anything resembling "general intelligence." Early analysis of OpenAI's o3 preview noted significant differences between preview and public versions, plus expenses that scale uncomfortably with desired accuracy.

But those critiques somewhat miss the point. ARC-AGI succeeded not because it perfectly captures intelligence—whatever that means—but because it created a credible, verifiable measuring stick precisely when the industry needed one. François Chollet, in a February profile in Le Monde, stressed that ARC-AGI exists to test adaptability to genuine novelty. The kind of skill acquisition that separates intelligence from statistical correlation. Mike Knoop, speaking on a Sequoia podcast, described the benchmark's motivation as forcing systems to reason with few examples and zero training-set overlap.

The foundation isn't resting. ARC-AGI-3, focused on interactive reasoning, is in development with fresh sponsor backing. The Verified program adds structure as measurement demands intensify. And perhaps most tellingly, research teams are building neurosymbolic approaches, tiny recursive models, and program-synthesis hybrids specifically to crack ARC-AGI's challenges. When a benchmark changes what researchers build, it's moved beyond simple measurement.

As models grow more capable and regulation tightens, the question shifts from "can we measure reasoning?" to "can we trust those measurements?" The fact that every major lab now reports ARC-AGI scores—on hidden test sets, certified by independent academics, with methodology disclosed—suggests the answer is starting to converge. Whether that convergence represents genuine progress or just a new form of consensus theater remains to be seen. But for now, those colorful grid puzzles have become something nobody expected in 2019: required reading for anyone building at the frontier.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • Immutrin Raises $87M Series A for Antibody Therapy to Clear Amyloid
  • DrHouse Raises $3.5M Seed to Scale 24/7 Telemedicine Platform
  • AI Models Predict Missing Biology Data to Speed Drug Discovery
  • The Race to Archive Human Skill for Robot Training
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.