Joshua Dillon has a history of building things that other engineers rely on without knowing his name. At Google Research, he created TensorFlow Probability, the statistical framework embedded in thousands of machine learning pipelines. Later, he contributed to Gemini and VideoPoet. Dan Kondratyuk, his co-founder, was first author on VideoPoet, a significant contribution to video generation research.
Now the two are making a different kind of bet. Results they posted recently suggest they've cracked a reasoning benchmark at a per-task cost roughly 10,000 times lower than the current state of the art. Not through bigger models. Not by throwing more GPUs at the problem. Instead, through autonomous AI teams that rewrite their own code, run hundreds of experiments in days, and surface algorithmic shortcuts that expensive brute-force search misses entirely.
If the claims hold, the implications stretch well beyond a single benchmark. The economics of scientific R&D—who can afford to do it, how fast breakthroughs arrive, what kinds of labs dominate—may be entering a sharp inflection. And it's happening at a moment when enterprises are racing to deploy autonomous systems faster than they're building the governance, verification, and infrastructure to manage them.
The question, in other words, isn't whether autonomous research agents will arrive. It's whether the rest of the world will keep pace.
Two Paths Diverging
Dillon and Kondratyuk launched rekursiv.ai out of Y Combinator with a positioning statement that reads less like a startup pitch and more like a manifesto: "Scale AI scientists whose own breakthroughs accelerate the next." The premise is that discovery is idea-bounded, not compute-bounded—a direct challenge to the prevailing wisdom in AI research, where the default response to hard problems has been to add parameters and pour in more compute.
Their evidence, detailed in a pair of technical posts, centers on two benchmarks. On ARC-AGI-1, a widely tracked reasoning test, they reported accuracy between 71.4% and 75.5%, depending on whether you're looking at a single model or an ensemble. Computational costs? Between 16 and 11,600 times lower than competing approaches at similar accuracy levels. The company baselined costs at $1.50 per H100 GPU-hour—a standard cloud rate that makes the math straightforward.
On Sudoku, they claim something more eyebrow-raising: the first 100% accuracy using only neural networks trained on input-output pairs, no symbolic solvers. They got there through a technique they're calling Hypothesis-Pinning Search, which emerged from an autonomous process that tested 684 experimental configurations over four days and generated 125,000 lines of code—though the technique proved specific to the Sudoku domain and didn't transfer directly to other problems.
Contrast that with the prevailing approach. A paper published in late June reported 72.9% accuracy on ARC-AGI-2's semi-private evaluation set—slightly better than rekursiv.ai's public-set results—at a cost of $38.99 per task. That's the test-time search paradigm: lean on heavy inference budgets, let the model reason its way through possible solutions, verify outputs, and bill accordingly. It works. It's just expensive, maybe prohibitively so if you're trying to run thousands of experiments rather than score a leaderboard once.
Anthropic demonstrated a version of the same tension in research it released earlier this year on what it calls Automated Alignment Researchers. The system showed strong productivity gains on math benchmarks—0.94 compared to human researchers on certain problem sets—and more modest improvements on coding tasks. But Anthropic was forthright about the costs: around $18,000 in tokens just for training runs, plus $22 per hour of active researcher time. Useful for narrow, high-value problems. Not obviously scalable to the kind of open-ended exploration rekursiv.ai is describing.
NVIDIA's ASPIRE robotics framework, detailed in a recent publication, took yet another tack—open-ended skill discovery loops that reduced token costs through compounding over time, but still required iterative human refinement at key decision points. A separate July paper on cost-effective agent orchestration achieved 67.25% on ARC-AGI-1's public set at $0.62 per task, illustrating the wide spread between inference-heavy strategies and algorithmically lean ones.
The rekursiv.ai results, if they generalize, suggest a third way: small models, tight feedback loops, autonomous experimentation, and a willingness to let the system rewrite its own approach until it finds something that works. Whether that proves out under external scrutiny is the next test.
Deployment Outpacing Governance
The timing of rekursiv.ai's emergence is not coincidental. According to Gartner, only 17% of organizations had deployed AI agents as of mid-year, but more than 60% expect to within the next two years. Forrester's assessment was blunter: companies are chasing, few are catching. The gap, analysts say, is partly technical complexity but increasingly about governance—how do you manage autonomous systems that iterate faster than your compliance teams can review?
Deloitte found that 80% of organizations lack mature governance frameworks for agentic AI, yet deployment is accelerating anyway. The Cloud Security Alliance reported something more unsettling: 54% of enterprises already have between one and 100 unsanctioned "shadow agents" running inside their networks, often without identity and access management oversight. These aren't rogue employees experimenting on weekends. They're business units solving problems with off-the-shelf tools, moving faster than central IT can track.
Regulatory bodies are scrambling to catch up, though the pace varies wildly by jurisdiction. The EU AI Act's transparency requirements became enforceable in early August, though high-risk obligations don't fully apply until later in the decade. China published national standards for AI agent interoperability over the summer and issued interim measures on anthropomorphic AI services. The UK is taking a sectoral approach, issuing guidance industry by industry rather than sweeping legislation.
In the U.S., the picture is more fragmented. The White House Office of Science and Technology Policy released a report in late July titled "Science: A New Golden Age," sketching a vision for AI-native scientific institutions and autonomous labs. Two days later, the National Science Foundation announced a $400 million network of AI-programmable cloud laboratories, part of what the agency is calling the Genesis Mission. Public-sector investment is ramping up. But federal agencies are also issuing reminders—NIH and CDC both put out guidance earlier this summer—warning against integrity breaches in AI-assisted research, a polite way of saying that automation without verification invites fabricated data and irreproducible results.
The gap between deployment velocity and governance maturity remains uncomfortably wide. And the companies building these systems are moving ahead regardless.
The Sudoku Reveal

rekursiv.ai's Sudoku work offers a window into how autonomous teams actually operate, and it's stranger than you might expect.
The company started with a baseline transformer achieving 36% accuracy—respectable for a neural approach to a problem that's trivial for symbolic solvers but notoriously hard for learned models. Over four days, the AI system proposed architectural changes: QK normalization, the Muon optimizer, tweaks to attention mechanisms. It invented Hypothesis-Pinning Search, a technique that showed strong results specifically in the Sudoku domain but didn't transfer to other problems like ARC. It tested 684 configurations. Ensembling brought accuracy to 100% first; then single-model methods caught up.
The entire process generated 125,000 lines of code. No human specified the algorithmic path. That's the provocative part. It's one thing to automate hyperparameter tuning or neural architecture search within a constrained design space. It's another to hand the system a problem and watch it invent novel search techniques.
The ARC-AGI work followed a similar arc. rekursiv.ai began from what it calls a Tiny Recursive Model scoring 44.9% and layered on feedback-and-repair loops, data augmentations, and ensemble methods to push into the low-to-mid 70s on ARC-AGI-1. Notably, they avoided test-time training—the technique powering many of the highest-scoring but most compute-intensive leaderboard entries. The company acknowledged in its post that it expects performance on the semi-private evaluation to drop by four to six percentage points, a refreshingly honest caveat. Public benchmark results don't always hold when moved to harder hidden tests, and rekursiv.ai hasn't submitted to that leaderboard yet.
On ARC-AGI-2, their reported accuracy was 17.5%, starting from a baseline around 7.8%. That's a meaningful jump, though still far from the frontier. What matters isn't the absolute number—it's the cost per point of improvement and whether the approach scales to harder problems.
Other teams are pursuing variations on the same theme, with mixed results. A preprint released earlier this year on AutoResearchClaw proposed a multi-agent pipeline with seven human-in-the-loop modes and cross-run evolution to prevent fabricated citations and hallucinated outputs. The system reported a 54.7% improvement over an earlier version of the AI Scientist framework on an internal benchmark of 25 research topics. The emphasis on verification is telling. As autonomous agents generate hypotheses faster than humans can validate them, the productivity multiplier depends entirely on building verification systems that scale with discovery.
NASA's vision paper on Accelerated Knowledge Discovery, published this spring in Earth & Space Science, framed the same tension from an institutional perspective. The authors argued for a new research paradigm combining AI models, autonomous agentic systems, and human-AI collaboration. But they flagged verification as the bottleneck: if agents can propose ten experiments in the time it takes a human to validate one, where's the win?
Money Talks
Venture investors are pricing in the shift, even if the business models remain hazy. Cognition Labs, maker of the Devin coding agent, raised $1 billion earlier this year at a $25 billion pre-money valuation—a number that raised eyebrows even in a frothy market. Sail Research, building infrastructure for long-horizon agents, closed an $80 million combined seed and Series A in late June at a $450 million valuation.
Smaller rounds tell a similar story. Enjamb raised $650,000 in a pre-seed round in the spring. Autoscience closed a $14 million seed in March. The pattern cuts across the stack: infrastructure for agent orchestration, verification tooling, domain-specific agent platforms. Investors are betting that autonomous systems will command enterprise budgets, even as ROI data remains patchy.
rekursiv.ai has disclosed no funding yet, and LinkedIn lists the company at two to ten employees—though such headcounts often lag reality by months. The company's go-to-market strategy is opaque. It describes its product as a "cockpit for autonomous research" where users dispatch fleets of AI scientists and trace claims back to evidence. Access is gated; there are no public customer case studies or customer logos. The company has open-sourced two components under Apache 2.0 licenses: configgle, a library for type-safe experiment configurations, and sagent, a CLI for self-mutating, multi-provider coding agents. Both repositories were active as of early August. Whether that open-core approach gains traction will depend on how quickly the broader ecosystem adopts verifiable, reproducible autonomous experimentation.
KPMG's survey of more than 2,000 executives earlier this year found that only 8% report established returns across the enterprise from AI investments, despite scaling programs. Microsoft's Work Trend Index noted uneven sectoral adoption and emphasized that management's role in rollout success remains decisive. Forrester's "few are catching" line captures the gap: enthusiasm is high, but value capture is concentrated among companies that solve verification, orchestration, and cost management simultaneously.
The capital is flowing because the potential multiplier is enormous. Anthropic sketched the outline in an essay published over the summer titled "When AI builds itself." The piece predicted productivity gains where 100-person organizations output work equivalent to 10,000 or even 100,000 people. But it also stressed misuse risks—scaled manipulation, surveillance—and argued that verification systems need to mature in parallel with discovery systems, not lag years behind.
That's the tension underneath rekursiv.ai's claims. If autonomous teams can generate breakthroughs at negligible cost, who verifies the breakthroughs? And what happens when verification itself becomes the bottleneck?
The Road Ahead

The next year will likely resolve whether rekursiv.ai's cost advantages are generalizable or domain-specific. The Sudoku post acknowledged that Hypothesis-Pinning Search didn't transfer directly to ARC. The absence of semi-private ARC submissions or external replications means the claims rest on self-reported public-set results, which is standard for early-stage work but leaves room for skepticism.
Still, the broader pattern is showing up across domains. NVIDIA's ASPIRE framework in robotics, Anthropic's Automated Alignment Researchers, and now rekursiv.ai's abstract reasoning results all point in the same direction: small teams using self-improving loops to find algorithmic shortcuts that expensive search misses. If that pattern holds across a wider set of problems, the R&D playbook shifts from "buy more compute" to "build better discovery loops."
The technical frontier is moving toward ARC-AGI-3, an agentic, interactive benchmark introduced earlier this year where frontier models currently score below 1% while humans hit 100%. The competition includes milestones running through late September, and leaderboards are shifting week to week. What makes ARC compelling for developers is that it resists brute-force scaling. Improvements come from discovering novel search algorithms or architectural innovations, not just adding parameters. If rekursiv.ai or its peers can demonstrate similar cost reductions on ARC-AGI-3—where the task is fundamentally harder—the "idea-bounded, not compute-bounded" thesis becomes much harder to dismiss.
Bento Labs reported earlier this year that adding what it called a "self-learning layer" improved ARC-AGI-3 scores by 2.6 times at constant budget, hinting at scaffold-level improvements that don't require retraining base models. That kind of architectural innovation is exactly what rekursiv.ai's autonomous loops are designed to surface.
The regulatory picture, meanwhile, remains unsettled. The EU's transparency rules are now enforceable, but high-risk obligations phase in over the next couple of years. China's interoperability standards and interim measures suggest a push toward governance-by-architecture—technical standards that constrain what systems can do, rather than process-heavy compliance regimes. The U.S. is threading a line between public-sector investment (NSF's $400 million lab network, OSTP's vision reports) and scattered agency-level guidance on research integrity. The gap between deployment speed and governance maturity is wide, and Deloitte's finding that four out of five enterprises lack agentic AI governance means that gap will produce compliance headaches before it narrows.
For founders and CTOs, the open question is whether autonomous research platforms coalesce into a distinct category or get absorbed into existing MLOps and experimentation tooling. rekursiv.ai is betting on the former, positioning itself as a new kind of research infrastructure. Whether that bet pays off depends on solving problems that go well beyond the technical: trust, reproducibility, attribution, and the question of what happens when the system inventing the next experiment is also the system that wrote the code for the last one.
Perhaps the most unsettling possibility is that rekursiv.ai is right about the cost curves but wrong about the timelines. If autonomous research agents can compress years of experimental work into days, the question isn't just whether humans can keep up. It's whether the institutions that fund, regulate, and verify scientific work can adapt fast enough to matter. Because if they can't, the breakthroughs will happen anyway—just somewhere outside the system.
