Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Climate / Social Tech iconClimate / Social TechMarch 16, 2026

PolyCycl Raises Series A to Scale Plastic-to-Oil Tech Across India

PolyCycl Raises Series A to Scale Plastic-to-Oil Tech Across India
Chemical RecyclingCircular Economy+3
SaaS iconSaaSMarch 16, 2026

Isembard Lands $50M Series A to Build 25 AI-Powered Factories

Isembard Lands $50M Series A to Build 25 AI-Powered Factories
Series AManufacturing+3
SaaS iconSaaS
March 16, 2026
Ai BenchmarkingAgi ResearchArtificial IntelligenceAi Governance

How ARC Prize Became the Gold Standard for Measuring AI Intelligence

Four frontier AI labs now report ARC-AGI scores in model cards. With ARC-AGI-3 launching March 25, the nonprofit's benchmarks are shaping the race toward AGI—and policy.

How ARC Prize Became the Gold Standard for Measuring AI Intelligence

When OpenAI unveiled o3 in December 2024, something unusual happened. The company didn't simply release benchmark scores and walk away. Instead, it partnered with a small nonprofit to verify the results in real time, broadcasting the process with a degree of transparency that caught even veteran observers off guard.

The benchmark at the center of this production—ARC-AGI—had been kicking around since 2019, largely ignored by an industry more interested in chasing state-of-the-art scores on established tests. But something shifted. Within months, three more frontier labs followed OpenAI's lead, embedding ARC-AGI scores in their official model cards. By early 2026, a test originally designed to measure fluid intelligence through abstract pattern recognition had quietly become the closest thing artificial intelligence has to a universal yardstick for general intelligence.

The trajectory is, to put it mildly, striking.

ARC Prize Foundation—founded in 2024 by Zapier co-founder Mike Knoop and Google engineer François Chollet—entered Y Combinator's Winter 2026 batch with a team of four people. Just four. Yet its benchmarks now appear in documentation from OpenAI, Anthropic, Google DeepMind, and xAI. When DeepMind published performance data for Gemini 3.1 Deep Think in February 2026, the company explicitly labeled its ARC-AGI-2 score of 84.6% as "ARC Prize verified." Anthropic's Opus 4.6 release the same month included similar verification language in its system card. OpenAI's GPT-5.2 model pages list "ARC-AGI-1 (Verified)" and "ARC-AGI-2 (Verified)" scores with the kind of matter-of-fact precision usually reserved for industry standards.

This isn't your typical benchmark proliferation story, where every lab spins up its own preferred metric. It's a standardization play happening in real time, and it's worth paying attention to.

The Intelligence Test AI Still Struggles to Game

Chollet introduced ARC in a 2019 paper with the blunt title "On the Measure of Intelligence." His argument? Most benchmarks test task-specific skill rather than skill-acquisition efficiency—the hallmark of what psychologists call fluid intelligence. The distinction matters more than it might seem at first glance.

A model trained on millions of math problems can ace graduate-level tests without demonstrating the ability to learn a genuinely novel task. It's pattern matching at scale, impressive but narrow. ARC-AGI tasks, by contrast, are deliberately designed to be human-easy but AI-hard: abstract visual puzzles requiring pattern recognition with minimal prior knowledge. The kind of thing a reasonably bright ten-year-old might solve in seconds while a billion-parameter model spins its wheels.

The early results bore this out in humbling fashion. A 2020 Kaggle competition yielded a winning score around 21%. By late 2024, state-of-the-art on ARC-AGI-1's private evaluation had climbed to 55.5%, according to ARC Prize's technical report published December 5, 2024. Progress was real but incremental, the kind of slow grind that makes for boring headlines.

Then o3 hit 87.5% in high-compute mode.

The number itself was impressive. The cost? That raised eyebrows. ARC Prize's blog post on o3 included detailed compute budgets and cost-per-task breakdowns, making the tradeoff explicit: reasoning capability scaled, yes, but so did infrastructure expense. This wasn't magic—it was brute force applied with uncommon intelligence.

The 2025 competition introduced ARC-AGI-2, designed specifically to counter contamination and overfitting pressures. Over 1,455 teams submitted 15,154 entries. The top score on the private evaluation reached 24%, per ARC Prize's January 15, 2026 technical report.

A reset, in other words. Deliberate and necessary. As models saturate older benchmarks—memorizing patterns, exploiting quirks—ARC evolves to stay one step ahead.

Four Labs, One Unlikely Standard

The adoption pattern accelerated through 2025 in ways that surprised even insiders. According to ARC Prize's 2025 technical report, "four frontier AI labs (Anthropic, Google DeepMind, OpenAI, xAI) reported ARC-AGI performance in public model cards in 2025." The language is careful—reported, not merely tested internally. Model cards, for those unfamiliar, are the industry's way of signaling what metrics actually matter. What gets measured gets managed, as the saying goes.

DeepMind's Gemini 3.1 Deep Think page doesn't just list an ARC-AGI-2 score; it links to a methodology PDF explaining the verification process in granular detail. Anthropic's Opus 4.6 news page states the team "submitted for official verification" with a 120,000-token thinking budget—a specific constraint that suggests this wasn't an afterthought. OpenAI's GPT-5.2 pages include ARC-AGI scores across multiple localized versions, published between December 2025 and January 2026.

The consistency is telling. It suggests internal benchmarking protocols now treat ARC-AGI as table stakes, the kind of metric you'd better be ready to discuss when your model drops.

xAI's public-facing documentation is less explicit, which tracks with the company's general preference for opacity. But ARC Prize's 2025 report names xAI alongside the other three labs. Grok 4 reportedly scored somewhere around 15.8% to 16% on ARC-AGI-2 in mid-2025, according to public reporting and the ARC Prize technical summary. The foundation lists xAI as a donor under "AI Lab Donors" on its homepage, alongside Google—a detail that might mean little, or might mean quite a lot depending on how cynical you're feeling.

The verification process itself appears to be tightening. ARC Prize maintains a public leaderboard with methodology notes, cost caps, and preview flags—all the scaffolding of a maturing ecosystem. Labs now disclose "reasoning budgets," specifying the number of thinking tokens or inference passes allowed. What was once a black box has become a semi-transparent arms race, which is progress of a sort.

DeepMind and Anthropic both note effort budgets in recent releases, a tacit acknowledgment that raw scores without context are incomplete signals. Perhaps even misleading ones.

The Benchmark Crisis No One Wanted to Talk About

The timing here isn't coincidental, though whether that's coordination or convergence is harder to say.

In November 2025, The Guardian reported on a UK AISI and academic review that found flaws across more than 440 AI benchmarks, citing reproducibility and contamination issues that should have been embarrassing for an industry that prides itself on rigor. The International AI Safety Report 2026, published in February, highlighted an "evaluation gap" where pre-deployment tests fail to predict real-world performance. The report noted rapid gains on math, science, and coding benchmarks but warned of "jagged" capabilities—models that excel on some tasks while failing spectacularly on superficially similar ones—and widespread test saturation.

Translation: the industry had a measurement problem, and everyone knew it.

Regulatory momentum is building, which tends to focus minds. The U.S. National Institute of Standards and Technology created the Center for AI Standards and Innovation (CAISI) in 2025 to lead evaluation standards. On February 13, 2026, NIST announced an International Network for Advanced AI Measurement, Evaluation, and Science, establishing consensus on automated evaluation best practices—the kind of boring but essential infrastructure work that doesn't make headlines but shapes entire industries.

The EU AI Act's obligations for general-purpose AI providers took effect August 2, 2025, requiring transparency and, for systemic-risk models, state-of-the-art evaluations. Vague language, to be sure, but the directional pressure is clear.

This regulatory scaffolding creates demand for credible, independent benchmarks. ARC Prize positions itself as precisely that: a nonprofit with third-party verification, donor transparency (conveniently listed on the homepage), and a mission described as providing "a North Star to open AGI" via human-calibrated, enduring benchmarks. Whether that mission statement survives contact with reality remains to be seen, but the foundation's leadership brings a mix of technical depth and startup operational experience that's hard to dismiss. Chollet created the benchmark. Knoop built a company that scaled. President Greg Kamradt rounds out the team.

Four people. Again, worth repeating.

Technical Evolution and the Quiet Neurosymbolic Revival

The methods driving ARC-AGI progress are themselves revealing, if you're the sort who finds technique interesting.

ARC Prize's 2024 report highlighted deep-learning-guided program synthesis and test-time training—approaches that blend neural networks with classical symbolic reasoning. The 2025 report emphasized "refinement loops," iterative program optimization guided by feedback signals, and something called Tiny Recursive Models (TRM). One TRM approach, using roughly 7 million parameters with on-the-fly refinement, achieved approximately 45% on ARC-AGI-1 and 8% on ARC-AGI-2, according to an arXiv paper from late 2025.

These compact, neurosymbolic architectures suggest a path distinct from pure scaling. You don't need a trillion parameters if your architecture is clever enough, the research implies. Whether that thesis holds up under pressure is an open question.

Yet the frontier labs are exploring scale-plus-reasoning hybrids that make different bets. OpenAI's o3 demonstrated that compute-intensive search, applied to a reasoning-capable base model, can yield step-function improvements on abstract tasks. The cost tradeoff—o3's high-compute mode ran at 172 times the inference budget of its high-efficiency setting—underscores a strategic tension that won't resolve neatly. Do labs optimize for raw capability or practical deployment economics?

The answer likely depends on use case: research validation versus customer-facing products. Or, put differently, what looks good in a press release versus what you can actually bill customers for.

The 2025 competition prize pool exceeded $725,000, according to ARC Prize's competition page. The 2024 grand prize of $600,000 went unclaimed, though smaller awards totaling over $125,000 were distributed. These aren't trivial sums for a nonprofit. The foundation's donor list includes AI labs and individuals—Tyler Cowen, Jeff Fang, Aman Shah, Aaron Levie—suggesting a coalition spanning academia, venture capital, and industry.

Money talks, as it usually does.

The Interactive Reasoning Shift (and Why It Matters)

Digital illustration for article section "The Interactive Reasoning Shift (and Why It Matters)" in "How ARC Prize Became the Gold Standard for Measuring AI Intelligence" - A conceptual and minimalist modern digital illustration representing interactive reasoning and goal-...

ARC-AGI-3, announced for March 25, 2026, represents the next inflection point. The foundation describes it as the "first interactive reasoning benchmark," measuring exploration, goal-directedness, and memory in novel game environments.

This moves beyond static pattern recognition into agentic territory—the kind of sequential decision-making required for multi-step planning. Solving a puzzle is one thing. Navigating an environment where your actions change the state space? That's closer to the messy, open-ended challenges of real-world deployment.

The shift mirrors broader industry trends. The International AI Safety Report 2026 notes that agent capability on multi-hour software tasks has roughly doubled every seven months. Extrapolations are uncertain, obviously, but suggest potential for multi-day autonomy by 2030. METR (Model Evaluation and Threat Research) tracks similar "time-horizon" metrics, with updates published in January and February 2026. The doubling times align with ARC's evolution cadence: from static puzzles (ARC-AGI-1) to harder static puzzles (ARC-AGI-2) to interactive environments (ARC-AGI-3).

Other benchmarks compete for mindshare, naturally. Humanity's Last Exam (HLE), launched by FutureHouse, tests dynamic research-grade Q&A. As of February 12, 2026, Gemini 3 Deep Think scored 48.4% with no tools, according to FutureHouse's site. GPQA Diamond, MMLU-Pro, MMMU-Pro, and FrontierMath cover graduate-level STEM and multimodal reasoning. Chatbot Arena's Elo rankings serve as a de facto live model comparator, widely used for reward model calibration.

Yet none have achieved the cross-lab standardization that ARC-AGI now commands. Not even close.

What This Actually Means (Beyond the Hype)

Digital illustration for article section "What This Actually Means (Beyond the Hype)" in "How ARC Prize Became the Gold Standard for Measuring AI Intelligence" - A conceptual and visually striking representation of deep reasoning and abstract problem-solving, fe...

For founders building on frontier models, ARC-AGI scores are becoming a proxy for reasoning depth. A model that scores 80%-plus on ARC-AGI-2 signals strong generalization—useful for applications requiring abstract problem-solving rather than rote pattern matching. Conversely, wildly divergent scores between ARC-AGI-1 and ARC-AGI-2 might indicate overfitting or contamination risk. Worth noting, at minimum.

Investors tracking AGI timelines should pay attention to the verification infrastructure. Third-party benchmarking is rare in an industry where self-reported metrics dominate and skepticism is warranted by default. ARC Prize's "Verified" badge on lab model cards introduces accountability. It's not foolproof—methodologies can be gamed, as they always are—but it's a step toward independently auditable progress claims.

Which matters, if you believe audits matter.

Policymakers face a thornier calculus. The EU AI Act and NIST's CAISI are converging on evaluation mandates, but the field lacks consensus on what "general intelligence" even means. ARC-AGI offers one operationalization: fluid intelligence as measured by skill-acquisition efficiency. Whether that definition survives contact with regulatory frameworks—where political considerations often trump technical precision—remains an open question. The International AI Safety Report's warning about the "evaluation gap" suggests that even strong benchmark performance may not predict safety outcomes, which is the sort of insight that should give everyone pause.

The next six months will clarify ARC Prize's durability. If ARC-AGI-3 launches on schedule March 25 and labs adopt it as quickly as they did ARC-AGI-2, the foundation's position as a de facto standard-setter solidifies. If labs fragment across competing benchmarks—or if ARC-AGI-3's interactive format proves too expensive to run at scale—the standardization window may close faster than it opened.

For now, though, the trend line is clear.

A four-person nonprofit, barely two years old, has convinced the world's most resource-rich AI labs to submit to its tests. That's not just a benchmark story. It's a governance experiment, playing out in public, with stakes higher than anyone involved probably anticipated.

Whether it holds is another matter entirely.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • PolyCycl Raises Series A to Scale Plastic-to-Oil Tech Across India
  • Isembard Lands $50M Series A to Build 25 AI-Powered Factories
  • Strand AI Predicts Missing Biology Data With Foundation Models
  • AI Co-Scientists Are Automating Research End-to-End. Can They Be Trusted?
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.