The pitch sounds almost absurd: compress a decade of research into a week. Run experiments that would normally take months in the span of a few hours. Do all of this with just two people and an army of artificial intelligence agents working in parallel.
Yet that's precisely what Aster Lab claims to be building—and the early results, however preliminary, are difficult to dismiss out of hand.
The San Francisco startup emerged from Y Combinator earlier this year calling itself "the first autonomous research lab." What that means in practice: a system designed to orchestrate thousands of AI research agents simultaneously, tackling everything from protein fitness prediction to GPU kernel optimization. It's an ambitious framing, perhaps more so than founders Emmett Bicker and Olivia Long initially expected.
The timing, at least, makes sense. Global R&D spending was estimated at roughly $3.8 trillion in 2024, and enterprises across sectors are racing to automate discovery workflows. A new category is taking shape—autonomous research systems that don't merely assist scientists but attempt to run experiments from hypothesis to conclusion. Market projections suggest the segment dedicated to agentic AI in scientific discovery grew from $260 million in 2025 to roughly $400 million in 2026, a 57% jump that reads like more than wishful thinking.
What sets Aster apart isn't just scale, though running 1,000 concurrent agents on commodity T4 GPUs is hardly trivial. It's the scope of their ambition. Competitors like Recursive have focused on automated AI research; FutureHouse has demonstrated end-to-end drug discovery. Aster, by contrast, is positioning itself as something more general-purpose: a research automation platform that can pivot across domains.
Their published results, released in June, span machine learning benchmarks, protein engineering, and even a mechanistic-interpretability project that compiled a language model into logic gates. Whether the breadth holds up under scrutiny remains an open question.
Planner, Worker, Subagent: The Hierarchy of Scale
Aster's architecture operates on three tiers. Planner agents decompose research questions. Worker agents execute experiments. Subagents handle granular tasks. The company laid out the framework in a June 8 research post titled "Scaling Autonomous Research to Thousands of Agents."
The most eye-catching demonstration involved ProteinGym's DMS substitutions benchmark. Aster reported a Spearman correlation of 0.524 to 0.526, achieved in roughly 30 minutes using approximately 1,000 concurrent agents on T4 GPUs. That figure, if it holds, would edge past the prior top score of 0.518 held by AIDO Protein-RAG. But there's a catch: as of mid-June, the official ProteinGym leaderboard hadn't yet reflected the submission. Leaderboard lag is common enough in academic benchmarks, yet until the numbers clear official review, they remain provisional at best.
Critically, Aster achieved these results through what they describe as inference-time changes—no model retraining required. If true, it's a meaningful data point. Inference-time gains suggest the architecture itself is doing heavy lifting, not just better-trained models.
On the NanoChat benchmark—a minimal language model training challenge created by Andrej Karpathy—Aster reported a validation bits-per-byte of 0.9098 after roughly 50 hours on 30 Blackwell B200 GPUs. That compares to Recursive's 0.9109, achieved in 40 hours, and a community best of 0.9372. Aster acknowledges high variance in these runs and frames the results as broadly comparable rather than definitively superior. A refreshingly honest admission, and one that suggests they're not yet claiming victory so much as demonstrating viability.
Elsewhere, the company's June research reports progress on optimizer search workflows, where 10 parallel agents running across 6 iterations completed experiments in approximately 2 hours. A separate project claims a 36.2% sampling-efficiency gain from what they describe as a "tessellated MAP-Elites archive." The language is technical, the implications less immediately clear.
Two Founders, Ambitious Timelines
Emmett Bicker, Aster's CEO, arrived from Magic, where he worked on long-context coding LLMs. His co-founder Olivia Long appears credited on the company's research posts but has kept a lower public profile. The Y Combinator directory lists the team size as two—lean even by startup standards—which raises an obvious question: how much of their operation is already agent-mediated?
Bicker's stated ambition, per the YC profile, is to "automate ten years of research over a single week." The company's public positioning describes a mission to compress research timelines from "years to hours" and claims progress toward "millions of tokens a second" experimentation throughput.
Bold talk. Whether it's achievable is another matter entirely.
The YC listing shows the company as "Active" but includes no disclosed public funding round, suggesting they're operating on standard YC batch funding or bootstrapping. For a compute-intensive operation running thousands of GPU hours, that seems like a narrow runway—unless cloud credits or partnerships are filling the gap.
The Broader Wave: Co-Scientist, Robin, and the Race for Autonomous Discovery
Aster is hardly alone. In May, Google opened researcher registration for "Co-Scientist," a Gemini-powered tool for hypothesis generation backed by a Nature paper describing multi-agent ideation architectures and lab validation. OpenAI updated its "Deep Research" feature in February, framing multi-hour web research as analyst-level agent work, with integrations via the Model Context Protocol aimed at improving reliability.
FutureHouse, a non-profit, announced in May that its "Robin" multi-agent system completed what it called the first end-to-end scientific discovery—from ideation through candidate ranking to wet-lab validation. The breakthrough involved actual molecules synthesized and tested, not just computational predictions. A milestone, if it holds up.
Sakana AI launched its "RSI Lab" and in June released "Fugu," a model that internalizes multi-agent orchestration—an attempt to move beyond external frameworks toward learned coordination. Recursive, meanwhile, continues to focus squarely on automated AI research, reporting NanoChat results and GPU kernel optimizations achieved through autonomous program search.
The physical lab automation side is converging too. Ginkgo Bioworks doubled down on its "Nebula" autonomous lab in the first quarter, opening cloud lab access via web browser and announcing a ProQR partnership in April to scale RNA editing discovery. Benchling released Automation in late May, a hardware-agnostic orchestration layer connecting instruments from partners including HighRes, Automata, and Opentrons. At the SLAS conference in February, Opentrons and HighRes demonstrated live "agent-to-agent" lab workflows—a signal that the industry is shifting from human-mediated automation to AI-directed physical experimentation.
Infrastructure: Enabler and Bottleneck

The surge in autonomous research systems tracks closely with GPU availability. NVIDIA's Blackwell B200 chips, which began deploying at major clouds in early 2026, offer substantial throughput and cost-per-token improvements over the prior Hopper generation. Aster's 30-GPU NanoChat run and Recursive's comparable experiments became feasible only as B200 capacity scaled.
Looking ahead, NVIDIA's Rubin architecture and GB300 chips, slated for release in the second half of this year, promise 5× inference performance and 10× lower cost per token versus Blackwell. If those numbers materialize in production—a big if—the economics shift materially. What costs thousands of dollars today might cost hundreds. Experiments that took hours might complete in minutes.
Hardware is only part of the equation, though. Framework consolidation is underway. Microsoft retired AutoGen to maintenance mode in April, launching the Microsoft Agent Framework 1.0 as a unified runtime combining AutoGen and Semantic Kernel. At BUILD in May, the company positioned standardized agent-to-agent protocols and Model Context Protocol support as enterprise plumbing—the unglamorous but necessary infrastructure layer.
The open-source ecosystem—LangGraph, CrewAI, and others—continues its rapid iteration. But a June disclosure of "AutoJack," a remote code execution vulnerability in AutoGen Studio, underscored the security challenges lurking beneath the hype. A joint CISA and Five Eyes guidance on agentic AI, released April 30, reflects mounting regulatory attention. The era of casual experimentation is ending.
What Still Doesn't Work
For all the momentum, gaps persist. And not small ones.
A January paper titled "Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision" highlighted brittleness in iterative editing tasks—a reminder that sequential reasoning remains harder than parallel search. AstaBench, an agent benchmarking suite published last October, concluded that research assistance remains "far from solved."
Reproducibility is another issue. Aster's own June results acknowledge high variance in the NanoChat benchmark, noting that comparability matters more than outright superiority. Multiple trials, seeded runs, and standardized evaluation harnesses—work exemplified by efforts like SWE-bench-Live and AgentCanary—will be necessary before the field moves beyond claims and into verifiable science. Right now, we're still in the claims phase.
Enterprise governance looms large. Gartner warned in May that 40% of enterprises might roll back autonomous agents by 2027 due to governance shortfalls. Authorization, oversight, and audit trails aren't research problems; they're operational necessities. And they're maturing more slowly than the agents themselves, which is a problem when systems start making consequential decisions without clear attribution.
The regulatory picture is sharpening as well. The EU AI Act's general applicability date—August 2 of this year—means tool-use in regulated domains may trigger compliance obligations, even as research exemptions provide some breathing room. In the U.S., NIST launched an AI Agent Standards Initiative in February to drive interoperability and safety. FDA guidance activity around AI in drug development continues to expand, albeit without the clarity that enterprises crave.
Reading Between the Benchmark Lines
Aster's early results are impressive on paper. Context, though, matters.
The ProteinGym leaderboard snapshot from June still lists AIDO Protein-RAG at 0.518, raising the question of whether Aster's claimed 0.524 to 0.526 represents a true advance or a difference in evaluation protocol. Leaderboard lag is common, as noted earlier. But until submissions clear official review, the numbers remain preliminary—suggestive, perhaps, but not definitive.
The NanoChat result of 0.9098 sits between Recursive's 0.9109 and a community best of 0.9372. Aster frames this as rough parity, which seems fair. It also suggests that massive agent parallelism hasn't yet delivered order-of-magnitude improvements over more focused approaches. Not yet, anyway.
What's more interesting than the specific numbers is the workflow Aster is validating: planner-worker-subagent hierarchies that can spin up hundreds or thousands of agents for parallel exploration, then consolidate results. If the architecture scales and generalizes, it becomes a platform—not just for reproducing known results faster, but for searching design spaces too large for human-directed experimentation. That's the real prize.
The Next Twelve Months (and What to Watch)

Several trends seem likely to collide over the coming year.
First, compute costs will drop as next-generation hardware reaches hyperscale deployment, making large-scale agent swarms economically viable for more teams. Second, wet-lab integration will mature as orchestration layers like Benchling Automation and autonomous platforms like Ginkgo's Nebula close the loop from digital hypothesis to physical test. Expect more pharma partnerships following the ProQR-Ginkgo RNA editing collaboration announced in April.
Third, benchmark realism will improve. The expansion of contamination-resistant evaluations and live benchmarks will recalibrate performance narratives, while multi-modal and desktop benchmarks gain traction. Fourth, enterprises will stress-test agentic deployments in production, and the survivors will be systems with mature authorization, audit, and governance controls. The others won't disappear, exactly, but they'll retreat to research labs.
Finally—and this is the hardest to predict—watch for signals around whether these systems can move beyond optimization and into genuine discovery. FutureHouse's Robin demonstrated end-to-end drug candidate identification. Can similar architectures find novel algorithms, materials, or biological mechanisms that weren't already implicit in training data? That's the test. Not whether agents can search faster, but whether they can search differently.
Aster Lab is two people building infrastructure for a future where research happens at machine speed. Whether Bicker and his team succeed in compressing years into hours remains an open question. But the broader industry is betting heavily that someone will—and that the someone might as well be a startup operating out of San Francisco with a Y Combinator badge and an improbable amount of ambition.
For now, the benchmarks are intriguing. The architecture is plausible. And the bet, however audacious, is at least worth watching.
