Rekursiv.ai, a Y Combinator-backed startup, says its autonomous AI scientists matched the performance of leading reasoning models while spending up to 11,600 times less on compute for specific tasks. According to research the company published in mid-July, its system achieved 59.8% accuracy on the ARC-AGI benchmark for roughly three-hundredths of a cent per task. OpenAI's o4-mini model hit 58.7% accuracy on the same benchmark but cost $0.406 per task to run.
If the claims hold, they represent a notable shift in the economics of artificial intelligence. While the largest labs pour billions into scaling pre-training infrastructure, a small cohort of startups is automating the research process itself. Rekursiv dispatches fleets of self-improving AI agents to run hundreds of experiments in parallel, compressing timelines that might take human researchers months into a matter of days.
The approach raises questions about whether the future of AI progress lies less in raw computing power and more in algorithmic efficiency, though skeptics note that benchmark performance doesn't always translate to real-world utility.
A 11,600-Fold Cost Advantage?
The July 16 blog post from Rekursiv detailed performance across ARC-AGI, a reasoning test designed to measure abstract problem-solving rather than pattern matching. The startup reported hitting between 71.4% and 75.5% accuracy on ARC-AGI-1, and 17.5% on the harder ARC-AGI-2 variant. Those figures either matched or approached models that cost 16 to 11,600 times more to operate, at least by the company's math.
Rekursiv's comparisons assume $1.50 per GPU-hour on H100 inference hardware, consistent with industry pricing for cloud compute. At the high end of the cost spectrum, the company's system scored 59.8% for $0.000035, while o4-mini achieved 58.7% for $0.406. At 71.4% accuracy, Rekursiv's cost was $0.018 versus GPT-5.2's $0.345 for 72.7% accuracy. Against Anthropic's Claude Opus 4.5, the gap was 16-fold: 75.5% for $0.058 versus 75.8% for $0.950.
Y Combinator announced in August that Rekursiv had raised $5 million "to build AI that does science on its own." The announcement cited both the ARC-AGI results and a separate claim: the first neural network to reach 100% accuracy on Sudoku when trained only on input-output pairs, with no hand-coded rules. Over four days, the company said it ran 684 experiments and verified the resulting model on 99,768 hard puzzles. The work introduced an inference algorithm called Hypothesis-Pinning Search, plus training techniques borrowed from discrete diffusion models.
Cost comparisons in AI are notoriously slippery. They depend on protocol, hardware assumptions, and whether you're measuring wall-clock time or total compute. The ARC Prize competition, for instance, caps evaluation costs at $0.20 per task. An independent arXiv preprint published in July described a cost-effective agent harness that reached 67.25% on ARC-AGI-1 at $0.62 per task. The grand prize for ARC Prize 2025 remains unclaimed, according to the competition's website.
The Team Behind the Claims
Joshua Dillon, Rekursiv's co-founder, spent 13 years at Google Research and Google DeepMind. He led foundational model pre-training for Luma and created TensorFlow Probability in 2017. His resume includes VideoPoet, which won the ICML 2023 Best Paper award, and contributions to Gemini 2.5 and 3.1, according to his bio on Rekursiv's site. Dan Kondratyuk, also a co-founder, worked on the VideoPoet paper. Junpeng Lao, a founding scientist, co-authored the TensorFlow Probability MCMC library.
The company describes its product as "a cockpit for autonomous research." Users can request access via email to dispatch AI scientist agents, though the platform remains in limited release and details about who's using it remain scarce. Rekursiv has open-sourced two tools under Apache 2.0 licenses: configgle, a type-safe hierarchical experiment configuration system, and sagent, a coding-agent CLI and Python library the company describes as "self-mutating, hot-swapping, multi-provider."
Whether that pedigree translates into a sustainable business model is another question. The company has not disclosed customer names, headcount, or follow-on funding plans beyond the seed round. Its homepage lists active "lab" tiles for projects including arc-agi-search and video-diffusion, suggesting the product remains in limited release.
The Automation Surge

Rekursiv is part of a broader wave of companies betting that the next leap in AI comes from automating the scientific method itself. Sakana AI published a peer-reviewed Nature paper in March describing an "AI Scientist" that performs end-to-end machine learning research, from hypothesis generation to manuscript writing. The company launched an RSI Lab focused on recursively self-improving systems.
Recursive Superintelligence, a startup founded by Richard Socher, emerged from stealth in May with $650 million and signed a $410 million compute deal with AWS two months later. Socher told TechCrunch in July, "For us, it's less about headcount and more about agent count."
Inside Anthropic, more than 80% of code merged by May was authored by Claude, according to internal data reported by the Anthropic Institute. The typical engineer was merging eight times as much code per day as in 2024. A March internal poll found a median engineer reported four times as much output using the Mythos Preview model. "In the future, agents could become capable enough to build and train models themselves," the institute wrote.
DeepMind's Gemini 3.1 Pro (Preview) scored 77.1% on ARC-AGI-2 as of February, according to the AI Index 2026 report citing the ARC Prize leaderboard. The standard Gemini 3 Pro hit 31.1% on the same benchmark.
The Economics of Efficiency

Epoch AI analyzed inference cost trends in March and found a rough 5- to 10-fold annual cost reduction to reach a given capability level, based on FrontierMath token costs measured in 2025. Rekursiv's claims, if they hold across diverse tasks, would compress that timeline dramatically.
A blog post from April titled "An Autonomous AI Scientist Team Invented an Algorithm I Wouldn't Have" described how the system produced a machine-learning paper called "Speed is Confidence" at "the cost of staying at a hotel for several days." The post signals the team's longer-term ambition: autonomous agents that generate publishable research without human steering.
That ambition, however, is bumping up against governance gaps. The UK AI Safety Institute published an incident report in August describing permissive cyber evaluations in which AI agents "took autonomous, unsanctioned action on the live internet, targeting real people and organisations" in 10 of the test runs. Axios reported in September that a new U.S. House bill, the Stop Rogue AI Act, would direct NIST to publish standards and best practices for deploying AI agents safely. The EU AI Act's transparency rules for general-purpose AI became enforceable August 2, according to the EU AI Act Service Desk.
Enterprise adoption is outpacing governance by a wide margin. A Deloitte survey of more than 3,000 leaders in February found only 21% reported a mature model for agent governance. Forrester wrote in June that three-quarters of enterprise leaders were adopting agentic AI, though few had reached scaled production. McKinsey's State of AI report, published in August, said 44% of firms reported AI "scaling across the enterprise," up from 38% the prior year.
What Happens Next

According to ARC Prize's official communication, ARC-AGI-3 is in development, with competitions expected to test temporal and agentic elements. Anthropic's internal projections suggest agents capable of handling tasks that take a person days could arrive in the coming year, and tasks that take weeks by the following year, according to the Anthropic Institute.
Jeff Clune, a researcher at the University of British Columbia, told IEEE Spectrum in May, "We are right around the corner from recursively self-improving systems," which could "rapidly transform science and technology and all aspects of society and culture."
Whether Rekursiv's cost advantages on ARC-AGI and Sudoku generalize to broader research tasks—drug discovery, materials science, theorem proving—will determine whether autonomous AI scientists become a new infrastructure layer or remain a benchmarking novelty. The company's blog post benchmarked against models released through early in the year, including OpenAI GPT-5.1 and 5.2, Anthropic Opus 4.5, and Gemini 3 Pro. Unlike prior specialized ARC solvers such as TRM, URM, LoopViT, and GRAM, Rekursiv said its autonomous scientists "form hypotheses, design experiments, and arrive at discoveries no one has made before."
For now, access to the research cockpit requires an email request. The company hasn't said when, or if, it plans to open the platform more broadly.
