For roughly $500, a small team of artificial intelligence agents spent a month running nearly a thousand experiments—and discovered an algorithm their own creator admits he probably couldn't have devised himself.
That's the claim, at least, from rekursiv.ai, a startup whose July 16, 2026 results suggest something potentially more consequential than another incremental benchmark victory: that AI systems might soon conduct research on themselves at a fraction of what today's frontier models cost to run. The implications, if the numbers hold up, cut deeper than academic bragging rights. This is about who—or what—discovers the next generation of algorithms, and at what price.
The headline figures are striking enough to warrant skepticism. Rekursiv.ai's autonomous scientist teams achieved between 71.4% and 75.5% accuracy on ARC-AGI-1, a notoriously punishing abstract reasoning benchmark, while burning through anywhere from 16 to 11,600 times less compute than comparable large language models. One configuration hit 59.8% accuracy at $0.000035 per task. OpenAI's o4-mini (High) scored 58.7% at $0.406 per task—roughly 11,600 times more expensive. At the upper end, their system reached 75.5% for $0.058 per task, versus Claude Opus 4.5 (Thinking, 32K) at 75.8% for $0.950. About sixteen times pricier.
Those calculations assume $1.50 per GPU-hour on H100 hardware and cover inference costs only, not training overhead. Whether that claim survives independent scrutiny is another matter entirely.
The Method Behind the Numbers
What distinguishes this work—at least in theory—is how these results emerged. The rekursiv.ai team didn't lean on test-time training, the fashionable technique where models spend extra compute "thinking" through problems at inference. Instead, they ran 235 full evaluations spanning 15 research directions using what they call autonomous research campaigns: AI systems proposing hypotheses, designing experiments, executing them, and interpreting the outcomes.
The discoveries came from the search itself. Their autonomous teams surfaced a cascade of training and evaluation tweaks—optimizer adjustments, normalization techniques, architecture simplifications, ensemble methods. These bubbled up from what the company describes as a multi-agent research loop driving idea exploration. One technique, dubbed Hypothesis-Pinning Search (HPS), helped them nail 100% accuracy on Sudoku-Extreme, a companion result published the same day. That involved a neural network trained exclusively on input-output pairs across 99,768 hard puzzles.
The Sudoku work matters because it established a pattern of transferability. Ideas discovered there—optimizer recipes, feedback-and-repair mechanisms, ensemble narrowing of the generation-selection gap—carried over to ARC-AGI-1 and partially to ARC-AGI-2. The system wasn't merely solving isolated puzzles. It was, ostensibly, building something like research intuition.
On ARC-AGI-2, the 2025-released version of the benchmark, rekursiv.ai reported 17.5% accuracy. For context, the top private set score during the ARC Prize 2025 competition hit 24.03%, achieved by team "NVARC" at around $0.20 per task. That competition ran from late March through early November 2025, with the technical report landing January 15, 2026. Rekursiv.ai's semi-private results aren't yet available, and they caution that they expect a 4-to-6 percentage point drop when evaluated on held-out data. A dose of realism, perhaps, or hedging.
Four AI Scientists, One RTX 5090, and $500
Joshua V. Dillon, one of rekursiv.ai's founders, framed the effort in a March 23, 2026 blog post as tackling a gap in how AI systems approach the scientific method. He described a month-long autonomous campaign costing roughly $500—four "AI scientists," an RTX 5090 GPU, and a Claude Max subscription—that churned through 984 experiments and yielded an algorithm he claims he "wouldn't have devised."
The architecture involves four agent roles. The Scientist proposes hypotheses. The Analyst interprets results. The Engineer implements and runs experiments. The Reviewer evaluates quality and decides what happens next. In that March demonstration, the system achieved 97% accuracy on Sudoku-Extreme versus an earlier 85%, using something like 167 times less training compute. Big claims, certainly.
Dillon's pedigree lends weight. He spent years as a Staff Research Scientist at Google Research and DeepMind, where he created TensorFlow Probability in 2017 and contributed to Gemini and VideoPoet. His co-founder, Dan Kondratyuk, led video generation models at Luma AI and was first author on VideoPoet, which won the ICML 2024 Best Paper award. Both bring serious technical chops to the table, though that doesn't guarantee their autonomous research thesis will pan out at scale.
Rekursiv.ai has released two tools under Apache-2.0 license: configgle, for type-safe experiment configurations, and sagent, a self-mutating agent framework supporting multiple providers. The company bills itself as building a "cockpit for autonomous research" with fleets of scientists dispatching experiments and tracking progress via leaderboards. Whether enterprises will trust such systems with mission-critical R&D is another question altogether.
What ARC-AGI Actually Measures—and Why It Matters

ARC-AGI (Abstraction and Reasoning Corpus for Artificial General Intelligence) tests a narrow but revealing capability: solving novel visual reasoning puzzles with minimal examples. François Chollet designed it to measure human-like skill acquisition efficiency rather than memorized knowledge. The benchmark keeps evolving. ARC-AGI-2, released in 2025, introduced harder puzzles. ARC-AGI-3, released in early 2026, went further still—it's an interactive agentic benchmark where human testers score 100% and frontier AI models scored "below 1%" as of March 2026. The focus shifted to exploration, goal inference, planning, and memory, with efficiency scored relative to human action counts.
Mike Knoop, ARC Prize co-organizer, wrote in a February 7, 2026 thread that AGI should be understood as human-level skill acquisition efficiency, not merely raw performance. ARC-AGI-3 was designed to measure action efficiency rather than rewarding systems for "thinking longer" without constraint. This emphasis on learning efficiency aligns with what rekursiv.ai is demonstrating—finding superior algorithms through targeted search rather than brute-force inference scaling.
The January 2026 ARC Prize technical report highlighted "refinement loops" as 2025's defining trend: iterative program optimization, evolutionary search, application-layer refinements. Small, zero-pretraining networks competed via recursive reasoning. The report positioned ARC-AGI-3 as the next frontier for interactive agentic intelligence. Rekursiv.ai's results fit that trajectory. Their systems sidestep expensive test-time compute by discovering better training recipes and architectures upfront. At least in principle.
Recent academic work buttresses the cost-efficiency angle. A July 7, 2026 preprint, "Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1," analyzed test-time compute versus benchmark-specific training and tools. Another line of research—Tiny Recursive Models (TRM), introduced in an October 2025 preprint, and Probabilistic TRM (PTRM), published May 2026—showed that small, domain-specific recursive models can rival larger LLMs on ARC-AGI-1. The lesson, it seems, is that specialization and recursive refinement can compete with sheer scale.
A Cluster of Competitors
Rekursiv.ai isn't working in a vacuum. A cluster of startups and labs is converging on autonomous research systems. Medra announced on June 24, 2026 that it launched an "AI Experimentalist" alongside a DARPA-funded project. The system translates high-level research objectives into executable plans. Medra also operates what it calls the "Largest Autonomous Lab in the US," scaling to 38,000 square feet with hundreds of robots, and hiring for positions like "Physical AI Scientist." Earlier, on April 23, 2026, the company detailed its Lab 001 facility.
CoreWeave launched ARIA on June 29, 2026—an AI research and iteration agent leveraging Weights & Biases Weave and CoreWeave infrastructure to accelerate machine learning research loops. Aster, a Y Combinator Spring 2026 company, positions itself as the "first autonomous research lab" orchestrating thousands of AI research agents in parallel. Propab AI announced plans for a V1 platform launching Fall 2026, targeting autonomous AI research for scientific use cases.
Outside startups, autonomous laboratories are cropping up in life sciences. A June 5, 2026 GPB/NPR feature covered "robot scientist" workflows at companies like Ginkgo Bioworks, citing an internal case where protein synthesis costs dropped roughly 40% compared to human-led workflows. The narrative spans computation and wet lab work alike.
DeepMind's earlier breakthroughs established a precedent. AlphaTensor, published in Nature in October 2022, discovered novel matrix multiplication algorithms. AlphaDev, published in Nature in June 2023, found faster sorting routines that made it into the C++ standard library. These were AI-discovered algorithms with measurable real-world impact. Rekursiv.ai's work suggests that paradigm might now run continuously and cheaply. Might.
The funding landscape reflects the momentum—or the hype, depending on your vantage point. Cognition, maker of the Devin coding agent, raised $1 billion at a $26 billion post-money valuation in late May 2026. A separate company, Recursive Superintelligence (distinct from rekursiv.ai), emerged from stealth with over $650 million at roughly $4.65 billion post-money in mid-May 2026, according to reports. Bloomberg had covered earlier funding talks in January 2026 at a $4 billion valuation. The "recursive self-improvement" pitch clearly draws investor attention, even across different organizations.
Gartner forecasted worldwide AI spending at $2.59 trillion in 2026, up 47% year-over-year, with agentic workflows cited as a key driver. In May 2026, Gartner placed agentic AI at the Peak of Inflated Expectations in its Hype Cycle, noting that 17% of organizations have deployed agents and over 60% plan to within two years, based on CIO surveys. IDC projected that by 2027, roughly 50% of enterprises will use AI agents to redefine human-machine collaboration, with agentic automation enhancing over 40% of enterprise applications. McKinsey, in reports from April and July 2026, observed infrastructure and sales functions shifting to agentic orchestration.
Those market projections warrant a grain of salt. Grand View Research lists an AI agents market size of $7.63 billion in 2025, forecasting $182.97 billion by 2033 at a 49.6% compound annual growth rate. The methodology is proprietary and hard to verify across regions and verticals. Analyst forecasts often reflect aspiration as much as reality. Still, the enterprise appetite appears genuine enough.
The Underlying Bet

Rekursiv.ai's July 16 blog post framed progress as "bounded by ideas, not compute." That's the philosophical wager underlying autonomous research systems. If better algorithms matter more than bigger models—or at least matter enough to shift the Pareto frontier—then tools that search efficiently for those algorithms become strategic assets.
The regulatory environment is taking shape, for better or worse. The EU AI Act's general-purpose AI (GPAI) obligations started August 2, 2026, requiring documentation, training data summaries, copyright policies, and systemic risk notifications for certain providers. Staged enforcement for high-risk systems follows in late 2027 and 2028. NIST's AI Risk Management Framework continues evolving, with a generative AI profile published in July 2024 and a critical infrastructure concept note in April 2026. DARPA's 2026 solicitations, including programs focused on compositional learning and reasoning, signal defense-sector interest in verifiable autonomy.
ARC-AGI-3's interactive format will likely favor systems with efficient exploration, memory, and verifiers over pure inference-time scaling. Rekursiv.ai notes they focus on ensemble selection to close the generation-selection gap—the difference between what a model can produce and what it can recognize as correct. As benchmarks reward learning efficiency, approaches that iterate on training and architecture will have structural advantages over those that simply scale inference budgets. In theory, anyway.
The Replication Question

The immediate challenge is replication. Rekursiv.ai's results are self-published on their blog. Independent validation on semi-private and private test splits would clarify how well their methods generalize. The company acknowledges they haven't submitted semi-private results yet and expect a 4-to-6 percentage point drop from public performance. If the cost-efficiency claims survive scrutiny, the implications extend well beyond abstract reasoning puzzles.
For CTOs evaluating autonomous R&D tools, the rekursiv.ai results suggest a new calculus. A four-agent system running on modest hardware might explore design space faster than human researchers—or at least explore different corners of that space. For AI infrastructure investors, the shift toward training-time efficiency rather than inference-time scaling could reshape compute demand patterns. If self-improving systems can uncover 100× or 10,000× cost reductions by discovering superior algorithms, the returns to raw compute scale start looking decidedly less linear.
Perhaps the most revealing detail is this: rekursiv.ai spent roughly $500 over 30 days in March to run 984 experiments that produced an algorithm their own founder said he wouldn't have invented. The system wasn't superintelligent. It was relentless, systematic, and cheap enough to explore ideas a human researcher might dismiss as improbable or uninteresting.
That's a different kind of advantage. Whether it scales beyond narrow benchmarks is an open question. The early data points, though, suggest this is a direction worth watching closely—even if some of the grander claims eventually deflate. After all, in an industry prone to hyperbole, a $500 experiment that genuinely surprises its own creator feels, at the very least, like an honest place to start.
