Josh Dillon left Google DeepMind in late 2025 after more than a decade building some of the company's most sophisticated AI infrastructure. He'd created TensorFlow Probability, led foundation model pre-training at Luma AI, watched billions flow into making models larger and hungrier. Then he walked away.
His bet? That the entire industry has been scaling the wrong variable.
Now Dillon and his co-founder Dan Kondratyuk—who spent years at Google working on multimodal systems including VideoPoet—are making a claim audacious enough to stop scrolling VCs mid-pitch deck. Their Y Combinator-backed startup, Rekursiv.ai, says it can match state-of-the-art performance on one of artificial intelligence's most punishing reasoning benchmarks at up to 10,000 times lower cost than frontier language models.
Ten thousand times.
If that number survives contact with independent verification, it would represent something closer to a phase transition than an incremental improvement. Tasks that currently drain tens or hundreds of thousands of dollars in compute could be measured in pocket change. The economics of AI research itself—who can afford to do it, at what scale, how fast—would fundamentally reorder.
The claim sits prominently on the company's Y Combinator S26 profile, where the founders describe autonomous systems that "ideate, run experiments, and discover new knowledge" through recursive improvement loops. The profile includes the up to 10,000× lower cost claim, which is not independently verified yet. But it arrives at an awkward moment for the industry. Explosive growth projections are colliding with grinding realities around cost, governance, and actual return on investment. And critically—perhaps unsurprisingly—that 10,000× multiplier hasn't been independently verified yet.
The Pitch: Machines That Invent Better Machines
Dillon and Kondratyuk aren't proposing to make models bigger or run them longer. They're chasing what they describe as a new scaling law: let AI systems discover better algorithms, which then make the next round of research faster and cheaper, which enables better algorithms still. Recursive improvement, not just parameter inflation.
Their March 2026 blog post offers the most concrete public evidence. An AI-scientist loop autonomously produced what they call a novel improvement to Sudoku-Extreme, hitting 97% accuracy on a single RTX 5090 GPU in 0.67 hours. Total project cost, Dillon wrote, compared to "staying at a hotel for several days." The post walks through multiple hypothesis-experiment-critique cycles, with the system generating and refining its own code. It's a proof of concept, not a peer-reviewed result, and demonstrates a concrete low-cost result for Sudoku-Extreme that is not related to their ARC claims. But it illustrates the broader pitch: machines inventing methods humans wouldn't have thought of.
The ARC-AGI claim—the 10,000× figure that's drawing all the attention—sits on shakier ground.
According to Rekursiv's YC profile, their AI scientists matched state-of-the-art accuracy on ARC-AGI in "a few days of self-directed research" at up to 10,000× lower cost. The Abstraction and Reasoning Corpus has become something of a gold standard for fluid intelligence and generalization. The third edition, which launched March 25, 2026, introduced interactive, open-world environments where early frontier models scored below 1% while humans sail past 95%. Cost per task isn't an afterthought on ARC leaderboards—it's a first-class metric. The public leaderboard only displays systems costing under $10,000 per run.
As of late July 2026, no independent replication of Rekursiv's exact cost multiplier has surfaced in public sources.
That absence doesn't mean the claim is wrong. It means we're in the uncomfortable zone where extraordinary assertions await extraordinary evidence.
The Reasoning Tax
The emphasis on cost reflects a broader industry reckoning. Test-time compute—the resources burned during inference, especially for reasoning models that explore multiple solution paths—can send bills into the stratosphere. OpenAI's o-series models, with their extended reasoning chains and majority-vote sampling strategies, triggered widespread discussion about what some have started calling the "reasoning tax."
The cost spread is enormous, almost comically so. A July 7, 2026 paper demonstrated cost-effective agent harnesses achieving 57.5% pass@2 on ARC-AGI-1 at roughly $0.25 per task, and 67.25% pass@2 at $0.62 per task. Yet research-grade runs can top $100,000, according to Bento Labs, which contrasted $100,000-plus research runs with $500 to $1,000 validation passes, emphasizing the cost dispersion across different evaluation contexts.
Anthropic's Automated Alignment Researchers, detailed in an April 2026 paper, cost about $22 per autonomous research-hour. Total token and training expenses for their experiment ran around $18,000. The system didn't beat human baselines in all settings—a useful reminder that autonomy doesn't automatically equal superiority.
Industry pricing is in constant flux. Multiple sources cite DeepSeek R1 as offering reasoning-tier capabilities at 18 to 36 times lower cost than comparable OpenAI o-series pricing, though such comparisons are heavily vendor-dependent and change with each provider update. Rough estimates suggest a 5 to 10 times annual improvement in price-for-performance at the frontier, driven by hardware advances and algorithmic efficiency.
It's a market where order-of-magnitude cost swings can hinge on orchestrator design, sampling depth, and when you locked in your API contract. Which makes evaluating a 10,000× claim... complicated.
Black Box, With Glimpses

Rekursiv's architecture remains largely proprietary. The public signals point to closed-loop research systems—iterative cycles of hypothesis generation, experiment execution, and self-critique, with the AI proposing modifications to its own code. Their YC profile emphasizes the recursive engine: "each new method makes the next round of research faster and cheaper." They claim to be the "first system to hit 100% on Sudoku with a neural net trained only on I/O pairs." That's technically distinct from the ARC-AGI work but demonstrates end-to-end autonomy on constrained reasoning tasks.
The broader research community has been converging on similar ideas, though from different angles. A July 2026 survey mapped 1,250 papers from 2024 to 2026 on recursive self-improvement, cataloging what improves (weights, orchestration harness, knowledge base) and the degree of loop closure. Sakana AI's "AI Scientist," published in Nature in March 2026, automates everything from ideation through experiments to paper-writing and peer review, claiming costs of $6 to $15 per paper—though independent analyses flagged quality concerns that dampened some of the enthusiasm.
Google DeepMind's Gemini Deep Think research agent, announced in February 2026, tackles scientific discovery with multi-agent collaboration. Meta's capacity-efficiency agents, rolled out April 16, compress roughly 10 hours of manual infrastructure diagnosis into 30 minutes. Andrej Karpathy's open-source AutoResearch, released March 7, lets researchers autonomously run ML experiments on a single GPU and saw rapid adoption.
But none have publicly staked out a 10,000× cost reduction on ARC-AGI. That's Rekursiv's alone, for now.
The distinction matters. Rekursiv isn't just automating literature review or experiment orchestration. They're positioning the system as an inventor of novel algorithms. That's a much higher bar than simply running experiments faster.
The Hype Cycle Meets the Trust Deficit
The agentic AI sector is riding a wave of real investment and equally real skepticism. Grand View Research pegged the AI agents market at $7.63 billion in 2025, forecasting $182.97 billion by 2033—a compound annual growth rate of 49.6%. Another report from 360iResearch estimated autonomous agents at $4.67 billion in 2025, growing to $15.77 billion by 2032. Morningstar projected "agentic automation" hitting $55 billion by 2036. Deloitte's TMT Predictions, published in November 2025, anticipated roughly $8.5 billion in 2026, climbing to around $35 billion by 2030.
The numbers are large and they vary wildly, which tells you something about how early-stage this all still is.
Enterprise adoption surveys tell a parallel story—bullish intent, cautious execution. Zapier's October 2025 poll found 84% of respondents planning to increase AI agent investments in 2026, with top use cases including document analysis and support triage. Deloitte's "State of AI in the Enterprise 2026," published in January, noted that agentic usage is scaling but governance maturity lags badly. Only about 20% of organizations reported mature agent oversight.
Forrester's June 2026 analysis was blunter: three-quarters of leaders are adopting agentic AI, but "companies are chasing, few are catching." Security decision-makers flagged agentic AI as a top concern. Forrester estimated that 25% of planned AI spend might defer to 2027 absent clearer ROI and control mechanisms.
That's the environment Rekursiv is stepping into. Huge addressable market, regulatory tightening, trust deficit. Microsoft's Work Trend Index 2026 surveyed 20,000 knowledge workers from February to April and documented significant agent adoption alongside persistent questions about human agency and organizational readiness. The Five Eyes agencies issued joint guidance on May 1 urging careful adoption of agentic AI services. NIST announced an AI Agent Standards Initiative on February 17, aiming for interoperable and secure agent ecosystems.
TechRadar Pro's July coverage highlighted what it called "the trust problem with agentic AI"—which, the outlet argued, is really a data problem. Rushing deployment could cost enterprises dearly if systems hallucinate, leak proprietary information, or simply waste cycles on dead-end explorations.
The "reasoning tax" and verification costs are now budget line items. Vendors emphasize orchestrators, verifiers, and selective re-exploration to manage dollars per successful task. OpenAI's July "scorecard" post explicitly stressed cost-per-successful-task as a strategic metric, not just raw accuracy.
What Works (So Far)

A few reference points help ground the conversation. Anthropic's Automated Alignment Researchers offer a useful benchmark. At roughly $22 per research-hour, the system is cheaper than hiring a postdoc. But it didn't consistently outperform human baselines in controlled comparisons—a result that's instructive. Cost savings can evaporate if output quality requires extensive human correction.
Bento Labs' April self-learning layer for ARC-AGI-3 agents reported a 2.6× improvement in scores for the same budget. The company contrasted $100,000-plus research runs with $500 to $1,000 validation passes, illustrating how staging and incremental refinement can bend the cost curve. Arc.computer, in a November 2025 CRM agent study, demonstrated 3.1 to 4.5 times cost reduction versus baselines through continual learning and evaluation loops.
Rekursiv's Sudoku-Extreme result fits this pattern but doesn't directly validate the ARC claim. Sudoku is a constrained problem with well-defined rules. ARC-AGI tests open-ended, novel reasoning in environments the model hasn't seen. The 0.67 GPU-hours on an RTX 5090 is impressively low, and the fact that the AI invented a method Dillon "wouldn't have" suggests genuine novelty.
But it's not the same task domain. Extrapolating from one to the other requires caution—perhaps more than the founders' early messaging conveys.
The Validation Gauntlet
Rekursiv's claim arrives at an inflection point for AI economics. If self-improving research loops can systematically compress discovery costs by three or four orders of magnitude, the implications ripple across drug discovery, materials science, chip design, and any domain where hypothesis generation and testing currently bottleneck on human bandwidth or compute budgets. The first company to reliably deliver such compression at scale wouldn't just lower costs. It would redefine the competitive landscape.
But the claim needs validation. The ARC Prize 2026 competition has milestones on June 30 and September 30, with a $2 million prize pool. Community reports of incremental ARC-AGI-3 improvements have hovered around 7.8% for some preview systems—still far below human performance. Rekursiv's exact evaluation protocol remains unspecified in public sources. Which ARC track? What pass@k setting? Which baseline models for the cost comparison?
Cost-per-task is highly sensitive to sampling depth, reasoning budget, and provider pricing windows. A 10,000× advantage in one configuration might collapse to 10× or 100× under different assumptions. That's not dishonesty; it's the nature of benchmarking in a rapidly moving field. But it means the headline number carries a lot of embedded asterisks.
The research trajectory suggests that knowledge-centric self-improvement—where systems explicitly audit and refine their own methods—will dominate near-term work on lowering cost per validated discovery. The ICLR 2026 Workshop on AI with Recursive Self-Improvement, held April 26, was the first focused venue on the topic. Multiple July papers advanced the state of the art. The sector is moving fast, and Rekursiv's founders have the pedigree to execute.
But extraordinary claims demand extraordinary evidence. The phrase has become a cliché, but in this case it's just true.
What Happens Next

For venture capitalists, the question is whether autonomous research systems can deliver ROI before governance frameworks, security concerns, and integration costs eat the savings. For enterprise AI decision-makers, it's whether these systems can operate reliably enough to trust with production research budgets. For researchers, it's whether the algorithms invented by AI scientist teams generalize beyond the benchmarks they're trained on.
And for everyone, it's whether a 10,000× cost cut is real, reproducible, and sustainable—or a fleeting artifact of early-stage optimization on a narrow task.
The answers will unfold over the next few quarters. Rekursiv has the spotlight, Y Combinator backing, and a bold thesis. What it needs now is independent replication, transparent methodology, and results that hold up under scrutiny from researchers who didn't build the system.
The economics of AI research may be on the verge of a phase change. Self-improving loops that compress discovery costs by orders of magnitude would represent a genuine breakthrough, not just an incremental win. Or we may be watching another startup learn that "up to 10,000×" can hide a remarkable number of caveats in two words.
Dillon and Kondratyuk have made their claim. The scientific community will do what it does best: try to break it. If the numbers survive, the implications stretch far beyond one benchmark or one startup. If they don't, well—there's always the next scaling law.
