A speedrun record set on a Monday. A startup launched on Tuesday. And a claim, ambitious to the point of audacity, that artificial intelligence could now discover better ways to train itself—faster than any human team working the problem.
That was Emmett Bicker's opening gambit in early February when he unveiled Aster Lab, a one-person venture emerging from Y Combinator's Spring batch with a premise that sounded either visionary or wildly premature, depending on who you asked. In a preprint uploaded to arXiv days later, Bicker claimed that his agentic workflows achieved results "over 20x faster" than existing methods across a sprawl of tasks: GPU kernel optimization, neural activity prediction, even obscure mathematical challenges. Whether those claims survive peer review and independent verification remains an open question. But the broader ambition behind Aster—that AI research itself could be automated, that the loop could close without human researchers in the driver's seat—reflects something larger happening in the field right now.
The spring brought a curious convergence. Industry titans who rarely sing from the same hymnal suddenly found common language: we've entered the "agentic era," they declared, almost in unison. Jensen Huang told a packed Dell Tech World audience in May that "for the very first time" we had truly useful AI. Satya Nadella spent the early months of the year evangelizing "agents as the new apps," describing a future where Windows itself would be reimagined for autonomous systems. The rhetoric felt coordinated, perhaps because the underlying economics had become impossible to ignore.
Gartner's numbers, issued in August 2025, forecast that 40% of enterprise applications would feature task-specific agents by year-end—a leap from less than 5% twelve months prior. By May, the firm projected worldwide AI spending would hit $2.59 trillion by the end of 2026, up 47% year over year. That represents an enormous amount of compute power funneling toward systems designed to operate, at least in theory, without constant human supervision.
Research automation occupies a different stratum of difficulty than customer service chatbots or code completion tools. Here, feedback loops stretch longer. The hypothesis space yawns vast and treacherous. Getting it wrong doesn't merely frustrate users—it burns expensive GPU time and potentially sends researchers chasing dead ends for weeks. Multiple groups have been circling this problem from different angles, with varying degrees of success and, crucially, transparency.
When Infrastructure Meets Ambition
Three forces converged to make autonomous research systems plausible now, when they seemed fanciful two years earlier.
Foundation models crossed a capability threshold. The Lion optimizer, discovered through program search back in 2023, proved that non-human approaches could yield components competitive with decades of human-designed alternatives in deep learning. That was proof of concept, a data point. By late 2025 and into 2026, models grew sophisticated enough to maintain context across longer experimental runs, reason about failure modes, and iterate without getting trapped in obvious local minima—most of the time.
Infrastructure caught up, too. The Model Context Protocol achieved widespread adoption, with reports suggesting over 10,000 active servers and somewhere in the neighborhood of 97 million monthly SDK downloads by spring, though those figures remain difficult to verify independently and security vulnerabilities were disclosed in April. Frameworks like Microsoft's AutoGen and LangGraph provided the orchestration layer that made multi-agent collaboration practical rather than theoretical. Zendesk adopted MCP in May. It was becoming standard plumbing, the kind of technology that stops making headlines once everyone's using it.
The third force? Economics. Training runs are expensive. Researcher time is expensive. If you can compress the discovery cycle from months to weeks, or weeks to days, you potentially unlock entire categories of experiments that were previously impractical—too costly, too slow, too speculative for finite research budgets. Google DeepMind's AlphaEvolve, introduced the previous May, demonstrated this with algorithm discovery and systems optimization. Andrej Karpathy's Autoresearch project, open-sourced in March, showed it could autonomously edit training code, run short experiments, and keep winners over multi-day stretches. Hundreds of changes executed without human intervention. Not perfectly, but functionally.
The Solo Act
Aster arrived in this context as a peculiar specimen. Bicker left his role researching long-context coding models at Magic to pursue the idea full-time. The company's homepage, as of late May, was actively recruiting "one extraordinarily talented person"—singular. For an enterprise positioning itself as "The First AI-Native AI Research Lab," it was an unusually lean operation.
The February preprint outlined Aster's approach across five domains: the Erdős minimum overlap problem from mathematics, GPU kernel optimization, single-cell denoising, neural activity prediction, and the NanoGPT Speedrun Competition. The paper claimed state-of-the-art or near-state-of-the-art results in each. Aster's homepage displayed artifacts like "SecantPolar," described as a direction-aware, geometry-aware optimizer, though detailed technical reports remained unpublished. Secondary coverage mentioned "PulseDelta," an architecture involving latent paths and grouped query attention, but specifics were scarce. The work existed in that liminal space between preprint and replication, between claim and confirmation.
What stood out was the framing. Aster wasn't positioning itself as infrastructure for researchers or a productivity tool for labs. It presented as an autonomous research entity—one that happened to be implemented in software rather than flesh. The homepage featured a "Prompt Aster" interface and visualizations of multi-step "Aster Runs" moving from hypothesis generation through testing to results. It felt closer to collaborating with an opinionously colleague than deploying software.
The pattern repeated elsewhere, with variations. DeepMind's AlphaEvolve used evolutionary coding agents for algorithm discovery. Stanford and industry researchers developed Astra, a multi-agent system that autonomously optimizes CUDA kernels with measurable speedups. KernelSkill, published in March, showed similar results. These weren't assistants; they were systems operating with significant autonomy, constrained primarily by compute budgets and the quality of their evaluation functions.
Karpathy's Autoresearch perhaps best captured the frontier's current state—open source, fully transparent, documented with warts and all. Community forks had applied it to hyperparameter and architecture exploration. It worked, within limits. It also occasionally pursued nonsensical directions before self-correcting, like a graduate student having an off week. The technology was real but imperfect, useful but not yet reliable.
The Messy Middle

Aster's funding situation illustrated both momentum and uncertainty. CBInsights listed a $500,000 convertible note raised roughly three weeks before late May, with Y Combinator as the investor, though Aster hadn't issued a press release confirming this. The company's YC page still showed a team size of one. For a field moving as quickly as agentic AI research, the usual startup signaling felt secondary to technical questions: Can these systems replicate their claimed results? Can others build on them? Can the work survive scrutiny? Until peer review or independent replication occurs, the results should be treated as unverified claims rather than established findings.
Those questions matter more than they might initially appear. Sakana AI's "AI Scientist" project, announced in 2024 and promoted through 2025, faced significant skepticism from independent analysts who questioned its claims and highlighted safety issues. The difference between a preprint and a reproduced result is the difference between an interesting idea and a reliable tool. Aster's February preprint was single-author, self-published on arXiv, not yet peer-reviewed. That doesn't make it wrong. It makes it unconfirmed.
The regulatory environment added another layer of complexity. The EU AI Act was set to reach general application on August 2, with staggered obligations continuing through 2027. While the U.S. had rescinded Executive Order 14110 back in January 2025—creating domestic ambiguity—European enterprises would need to document autonomous agent workflows for high-risk applications. Gartner warned in April that Fortune 500 companies could operate 150,000-plus agents by 2028, describing "agent sprawl" as a governance challenge waiting to happen.
McKinsey's surveys from spring showed agentic AI moving from experimentation to scaled deployment across industries, but trust and governance gaps persisted. The technology had outpaced the institutional frameworks needed to deploy it safely at scale. This lag wasn't necessarily catastrophic—it might prevent premature lock-in to approaches that wouldn't age well—but it created strategic uncertainty for everyone from solo founders to established labs trying to plan beyond the next funding round.
What Happens When the Loop Closes?

What Aster represents, beyond its specific technical contributions, is a bet on direction. If autonomous research systems can truly close the loop—not just assisting human researchers but independently proposing, testing, and validating hypotheses—then the bottleneck in AI development shifts from researcher time to compute and problem formulation. That's a different world, one where the pace of discovery accelerates in ways that are difficult to predict.
Whether it's a better world depends partly on technical maturity, partly on how the field handles interpretability and safety, and partly on questions we probably haven't thought to ask yet. The distinction between AI as tool and AI as autonomous research entity might seem semantic until you start running hundreds of experiments overnight without human oversight. Then it becomes rather concrete.
The NanoGPT speedrun that Bicker set the day before launch? It was both technical achievement and statement of intent. A demonstration that the founder could compete at the frontier, certainly. But also a signal about what Aster aspired to become: not infrastructure for human-led research, but a research entity that happened to run on silicon instead of neurons.
The longer-term implications remain genuinely unclear. Which might be the most honest thing you can say about any genuinely new technology.
