Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

SaaS iconSaaSOctober 4, 2026

Subvocal launches under-chin wearable for silent computer control

Subvocal launches under-chin wearable for silent computer control
YcBrain Computer Interface+3
SaaS iconSaaSOctober 4, 2026

DoD Solution raises $2M for AI drone navigation in war zones

DoD Solution raises $2M for AI drone navigation in war zones
Defense TechDrone Tech+3
Healthtech & Biotech iconHealthtech & BiotechMarch 12, 2026

China's BCI Boom: How Brain-Computer Startups Are Racing to Market

China's BCI Boom: How Brain-Computer Startups Are Racing to Market
Climate / Social Tech iconClimate / Social TechMarch 12, 2026

How Self-Charging Drones Could Transform Grid Inspections

How Self-Charging Drones Could Transform Grid Inspections
YcDrone Tech+3
SaaS iconSaaS
March 12, 2026
YcAi BenchmarkingAgi ResearchNonprofit TechArtificial Intelligence

How a YC Nonprofit Became AI's Official Scorekeeper

ARC Prize Foundation's AGI benchmark now verifies every major AI lab's intelligence claims—from OpenAI to DeepMind. As models get smarter, the test gets harder.

How a YC Nonprofit Became AI's Official Scorekeeper

OpenAI's launch of GPT-5.2 came with an odd flourish. Alongside the usual cascade of capability claims sat a table of benchmark scores bearing a single word: "Verified." Not by OpenAI's internal testers, mind you. By a nonprofit most industry watchers had never heard of.

Three months later, Google DeepMind followed suit. Gemini 3 Deep Think scored 84.6% on something called ARC-AGI-2, the company announced—"verified by the ARC Prize Foundation."

In less than two years, a Y Combinator-linked nonprofit co-founded by François Chollet and Mike Knoop has become, improbably, AI's unofficial referee. Today the ARC Prize Foundation verifies intelligence claims from every major frontier lab: OpenAI, Anthropic, Google DeepMind, xAI. According to the foundation's January 2026 technical report, its benchmark—ARC-AGI—has evolved from an academic curiosity into what they're now calling "an industry standard," appearing in model cards and launch announcements across the sector.

This wasn't foreordained. It happened because the AI industry needed something it conspicuously lacked: a test that couldn't be gamed.

The Test That Wouldn't Break

Chollet, creator of the widely adopted Keras framework, first published the Abstraction and Reasoning Corpus in 2019 with a specific ambition. He wanted to measure "general fluid intelligence" through colored grid puzzles requiring few-shot learning and abstraction. The tasks looked deceptively simple. They proved wickedly hard to solve. For years, state-of-the-art models barely moved the needle.

Then the scaling era arrived. Large language models started approaching human-level performance on traditional benchmarks through brute compute and pattern matching. Chollet watched this unfold and concluded that most tests were measuring memorization, not intelligence.

His solution? Make the test harder.

In May 2025, ARC Prize launched ARC-AGI-2, featuring increased task complexity. When labs threw massive compute at ARC-AGI-1—OpenAI's unreleased o3 was the first to achieve a qualifying score under extreme compute conditions—Chollet and Knoop had already formalized the nonprofit in early 2024. By the time they announced ARC-AGI-2 publicly in mid-2025, they'd settled on a strategy: keep the benchmark ahead of the models. Always.

The 2025 competition drew 1,455 teams and 15,154 submissions. All 90 research papers were released open source. The top score on the private leaderboard hit 24% at roughly $0.20 per task. The Grand Prize—which requires 85% accuracy—remained unclaimed.

Meanwhile, frontier labs began reporting their own scores. xAI's Grok 4 managed 15.9% in July 2025. Anthropic's Claude Opus 4.5 hit 37.6% by November. OpenAI's GPT-5.2 reached 52.9% in December.

By February 2026, when Google announced Gemini 3 Deep Think had achieved 84.6%, something had quietly shifted. Labs weren't just running the benchmark anymore. They were coordinating with ARC Prize for verified results on launch day.

Verification as Currency

That "Verified" label carries more weight than it might seem. In an industry where benchmark contamination has become a recurring embarrassment, third-party verification offers plausible deniability and borrowed credibility. OpenAI raised concerns about SWE-bench Verified in February 2026, citing contamination issues. Stanford's HELM benchmarks provide holistic capability dashboards, but they're retrospective assessments. MLCommons' AILuminate focuses on safety rather than raw capability.

ARC Prize found an opening: same-day verification of reasoning capabilities at model launch. According to a 2025 case study by infrastructure provider Vunda, ARC Prize runs concurrent evaluations of 400 tasks across more than seven model providers, integrating results into launch announcements in real time.

It's operational benchmarking as a service, really. Wrapped in academic legitimacy.

The foundation's funding model reflects this hybrid identity. Named donors include Vercel ($30,000), Nous Research ($25,000), Box CEO Aaron Levie ($25,000), and Andreessen Horowitz (amount undisclosed). Enough to run competitions and maintain infrastructure. Not enough, presumably, to compromise independence. The nonprofit structure helps. So does the pedigree: Knoop co-founded Zapier; Chollet created one of the most widely used deep learning frameworks in existence.

In March 2025, ARC Prize submitted a response to the White House Office of Science and Technology Policy's AI Action Plan RFI. The document positioned benchmarking as a U.S. strategic capability and recommended establishing a benchmarking hub within NIST's AI Safety Institute. Whether or not that materializes, the timing was strategic. The EU AI Act enters general application in August 2026, with phased obligations running through 2027. International networks of AI safety institutes—launched in November 2024—are conducting joint testing exercises. Governments want standardized, contamination-resistant evaluations.

ARC Prize is building exactly that.

The Moving Target

Digital illustration for article section "The Moving Target" in "How a YC Nonprofit Became AI's Official Scorekeeper" - A clean, minimal conceptual composition representing an evolving benchmark and moving target, featur...

The benchmark keeps evolving because it has to. On March 5, 2026, ARC Prize published ARC-TGI, introducing 461 human-validated task generators designed to prevent overfitting and enable scalable sampling. Instead of hand-authoring a fixed set of puzzles, the generators produce new variations that maintain core reasoning requirements while preventing memorization.

Three weeks later—March 25, 2026—ARC-AGI-3 launches. It represents a fundamental departure from static puzzles. The new benchmark places agents into interactive environments with more than 1,000 "games" or levels where they must explore, learn, plan, and adapt. There's a developer harness. APIs. Scorecards. It's closer to how humans actually demonstrate intelligence: through open-ended problem-solving in dynamic environments.

This mirrors a broader shift across AI evaluation, where benchmarks are moving from static knowledge tests to agentic scenarios. Scale's MCP-Atlas (launched December 19, 2025) tests models on real tool servers. Mercor's APEX-Agents (published January 20, 2026) evaluates cross-application professional tasks. OpenAI's BrowseComp (April 2025) measures web browsing agents. OSWorld-Verified tests computer use through GUI interactions.

ARC Prize isn't alone in this transition. But it may be furthest ahead in one specific dimension: making the transition illegible to gaming. Interactive environments resist contamination better than static datasets. Generated tasks resist memorization better than hand-authored ones.

The question is whether ARC can maintain that edge as models grow more capable.

The Marketing Problem

Digital illustration for article section "The Marketing Problem" in "How a YC Nonprofit Became AI's Official Scorekeeper" - A close-up, minimalist conceptual representation of a competitive data race, featuring three sleek, ...

Current ARC-AGI-2 scores tell a particular story. Google's Gemini 3 Deep Think leads at 84.6%, tantalizingly close to the 85% Grand Prize threshold. OpenAI's GPT-5.2 sits at 52.9%. Anthropic's Claude Opus 4.6, based on secondary reports from February 2026, appears to have reached approximately 68.8%. These numbers are verified, timestamped, comparable.

They're also incomplete.

ARC-AGI-2 measures abstract reasoning on grid puzzles. It doesn't capture everything—coding ability, long-context understanding, multimodal reasoning, real-world task completion. Labs now report a constellation of scores across multiple benchmarks: GPQA Diamond for expert-level science questions, AIME 2025 for competition math, MMMU-Pro for multimodal understanding, Humanity's Last Exam for closed-ended knowledge questions.

The risk? That verified ARC scores become a marketing proxy for general intelligence rather than what they actually measure: sample-efficient learning and abstract pattern recognition. Chollet's original critique of scaling-only approaches—that they don't produce true AGI—remains central to ARC's philosophy. The benchmark is designed to fail models that rely purely on scale and memorization.

Whether that philosophy survives contact with an industry increasingly focused on agentic capabilities and real-world performance is, well, unclear. ARC-AGI-3's shift to interactive environments suggests the foundation understands this tension. But the transition from puzzles to agents also makes the benchmark harder to explain, harder to verify, potentially more subjective.

Staying Ahead

Digital illustration for article section "Staying Ahead" in "How a YC Nonprofit Became AI's Official Scorekeeper" - A minimalist and conceptual representation of third-party validation and staying ahead, featuring a ...

For now, ARC Prize occupies valuable territory. Independent enough to be credible. Operational enough to be useful. Difficult enough to remain relevant as models improve. The "Verified" label appearing in launch announcements from OpenAI and Google signals that frontier labs consider third-party validation worth the coordination cost.

That could change, of course. Labs could build internal evaluation teams that match ARC's credibility. Governments could mandate standardized testing through institutions like NIST's AI Safety Institute or the UK's AISI. Competitors could emerge with better benchmarks or faster verification pipelines.

Or the benchmark treadmill could simply continue. ARC-AGI-3 gives way to ARC-AGI-4, which gives way to ARC-AGI-5, each iteration racing to stay ahead of model capabilities. The foundation's technical report claims they've achieved industry standard status. That status depends on maintaining a test that's hard enough to matter but solvable enough to show progress.

The 2026 competition launches with ARC-AGI-3 and the new generator framework. If frontier models crack 85% on ARC-AGI-2 before year's end—Google is already close—Chollet and Knoop will need to move fast.

The test only works if it can't be beaten easily. Once a nonprofit becomes the official scorekeeper for AGI, staying ahead of the competition stops being a choice. It becomes an existential requirement.

More stories

  • Subvocal launches under-chin wearable for silent computer control
  • DoD Solution raises $2M for AI drone navigation in war zones
  • China's BCI Boom: How Brain-Computer Startups Are Racing to Market
  • How Self-Charging Drones Could Transform Grid Inspections
  • The AI Race to Make Gene Therapies Safer
  • Mining Parasite Evolution: YC-Backed Ditto Bio's Autoimmune Gambit
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.