Emmett Bicker wasn't looking for SecantPolar. His research agent found it.
The optimizer—described as "direction-aware" and "geometry-aware" in technical terms that make most people's eyes glaze over—represents something larger than another incremental tweak to machine learning architecture. It's an artifact from a loop that feels almost unsettlingly circular: artificial intelligence designing better artificial intelligence, at a pace that leaves human researchers scrambling to keep up.
Bicker runs Aster Lab, a one-person operation out of San Francisco that emerged from Y Combinator's Spring 2026 batch with an audacious pitch: the first genuinely AI-native research lab. Not humans using AI tools. AI doing the research, with humans mostly watching. A preprint published in February claims the company's agentic system can deliver results more than 20 times faster than conventional methods across an improbable span of disciplines—mathematics, GPU kernel design, biology, neuroscience, even the training of language models themselves—though this speedup has not yet been independently verified.
The timing, perhaps, is everything. Gartner projected this May that worldwide AI spending would reach $2.59 trillion this year, a 47% jump that reflects more than just hype. AI-optimized server spending is expected to triple over five years. And agentic workflows—systems that operate with increasing autonomy—are driving much of that surge. Another Gartner analysis found that while only 17% of organizations have deployed agents today, over 60% expect to within two years. No other emerging technology shows that kind of adoption trajectory.
But here's the uncomfortable part: three months after that preprint, SecantPolar's optimizer details and validation remain unverified through public technical papers or independent studies. No ablation studies validate the claims. The optimizer appears on Aster's homepage as a headline achievement, yet the evidence trail grows thin quickly.
Which raises a question the industry is increasingly grappling with—how do you evaluate discoveries made by machines, at machine speed, when the usual peer review and replication cycles can't keep pace?
A Proof of Concept, or Just Proof of Hype?
Aster's website invites visitors to prompt the system with a simple challenge: "discover a new LLM architecture." SecantPolar is displayed as the answer. A second architecture, PulseDelta, surfaces in scattered third-party write-ups but lacks primary documentation confirming its existence. The documentation gaps aren't trivial. Independent researchers need details to replicate, validate, or—just as important—identify where things might fail.
Bicker came to this work young, studying machine learning at 15 before moving on to long-context coding models at a startup called Magic. CB Insights lists Aster's funding at $500,000, likely the standard Y Combinator package, though this has not been independently confirmed by the company. LinkedIn shows one employee. The February preprint describes applications ranging from the Erdős minimum overlap problem in mathematics to optimizing GPU memory patterns and improving single-cell denoising in biology—a scope so broad it invites skepticism.
The NanoGPT Speedrun has become an unexpected testing ground for these systems. It's a competitive benchmark: train a small language model as fast as possible. METR, an AI safety research group, noted in an April analysis that Aster "improved kernel memory access patterns" in record-breaking runs. A February 2 record was credited to "EmmettBicker & AI System Aster," according to a May feature in Founderland. But that phrasing—the ampersand doing all the work—captures the ambiguity. How much was human insight? How much was the agent operating on its own?
The distinction matters more than it might seem.
Compute Scarcity, Meet Algorithmic Discovery
Resource constraints are making agent-driven research attractive in ways that transcend pure novelty. IDC warned earlier this year that memory shortages are pressuring both PC shipments and data center infrastructure, even as the total market value climbs toward $274 billion. Teams that can discover more efficient optimizers or novel architectures without burning through massive trial-and-error cycles gain leverage in an environment where compute is increasingly precious.
The idea of machines discovering algorithms isn't new. DeepMind's AlphaTensor found new matrix multiplication methods via reinforcement learning back in October 2022. AlphaDev, published in Nature the following June, uncovered faster sorting algorithms. FunSearch, released in January 2024, applied large language models to combinatorial math and delivered fresh results on problems like the cap set problem. GNoME, announced in November 2023, predicted 2.2 million crystal structures—around 380,000 deemed stable—a materials science breakthrough powered by deep learning at scale.
Sakana AI, operating out of Tokyo, has been iterating on end-to-end agent pipelines since its "AI Scientist" v1 dropped in August 2024. Allen Institute for AI launched Asta in 2026, building an entire ecosystem of agents, benchmarks, and resources aimed at accelerating scientific research. And Yann LeCun's AMI Labs pulled in $1.03 billion this March to pursue "world models" as an alternative to pure language model scaling—a signal that capital is flowing toward varied approaches to intelligence amplification, not just one dominant paradigm.
Still, precedent and promise don't equal proof.
The Transparency Problem

Aster's preprint claims the system can "find SOTA results in minutes to hours," automating hypothesis generation, code implementation, and iterative refinement. The workflow mirrors what DeepMind demonstrated with FunSearch: a language model generates candidate programs, an evaluator scores them, the loop continues until performance plateaus or a target is hit. Aster extends this concept across more domains—neuroscience benchmarks like ZAPBench, GPU kernel optimization, biological data processing.
Yet as of late May, no ablation studies, benchmark comparisons, or open-source code were publicly accessible. The company's site mentions a web interface and API, though access details remain unclear. Independent researchers have started asking whether claimed speedups account for human oversight, prompt engineering, or debugging time—factors that blur the line between "autonomous discovery" and "well-tooled human-AI collaboration."
Sakana's AI Scientist faced similar scrutiny. While the open-sourced pipeline demonstrated that agents could draft papers and run experiments, reviewers flagged issues with novelty, rigor, and the occasional generation of figures that made no sense. The promise feels real. The execution remains uneven.
Intology's Locus agent claimed a NanoGPT speedrun world record this January, posting proof to Reddit with a GitHub pull request. PrimeIntellect published a post-mortem in mid-May analyzing the "agent frontier vs human awareness" in these contests, noting that human competitors often discover optimizations only after agents surface them. It's a reminder that the feedback loop runs both ways—machines learning from humans, humans learning from machines, the boundary increasingly porous.
What Happens When Validation Can't Keep Up?
The next chapter hinges on whether these systems can withstand scrutiny. Learned optimizers like VeLO and μLO, published between 2022 and 2024, showed promise but struggled with cross-task generalization. Some studies questioned whether gains held up outside narrow benchmarks. Aster and similar systems face the same test: Do agent-discovered optimizers transfer? Can they be reproduced independently? Do they introduce subtle failure modes that only emerge at scale?
Regulatory pressures, meanwhile, are mounting. The EU AI Act's main obligations take effect August 2, with high-risk classification guidance under public consultation as of May 20. Agents deployed in hiring, healthcare, or critical infrastructure will face compliance requirements. In the U.S., Executive Order 14110 from October 2023 requires reporting on dual-use foundation models, and NIST's AI Risk Management Framework remains the voluntary standard many enterprises reference.
Benchmarks are proliferating to stress-test what agents can actually do. YC-Bench, released April 1, simulates a startup's year-long lifecycle to evaluate long-term planning. AssistantBench and AgentBench exposed persistent gaps in navigation and real-time task execution. Agent-SafetyBench and Agent Security Bench revealed vulnerability rates that should give anyone deploying these systems pause—high attack success rates across many operational stages.
The market, for its part, expects rapid growth. IDC's FutureScape 2026, published last November, projected that agentic systems will account for nearly half of AI spending by 2029. Analyst estimates in April pegged the agentic AI market at $7.8 billion to $11 billion this year, with long-range forecasts reaching $139 billion to $199 billion by 2034. McKinsey's June 2023 estimate of $2.6 trillion to $4.4 trillion in annual economic potential from generative AI still anchors many corporate strategies, even as newer data refines the picture.
For founders and technical leaders, the calculus is shifting in ways that feel both exhilarating and precarious. If agents can compress months of optimizer tuning into hours, R&D budgets stretch further. If they can explore architectural search spaces humans wouldn't think to probe, novel capabilities emerge. But if those discoveries lack rigorous documentation or independent replication, the value proposition collapses into hype.
Signal, Noise, and What Comes Next

Aster Lab's SecantPolar remains more artifact than validated breakthrough. A signpost of what might be possible, not proof of what already is. The preprint is three months old. The company has one employee. The technical details are thin, the independent verification thinner.
Yet the broader trend it represents—AI systems iteratively improving themselves, surfacing optimizers and architectures at machine speed—is undeniable. The pattern is accelerating across labs in San Francisco, Tokyo, Seattle, Paris. The question isn't whether this approach will mature.
It's how quickly the industry can separate signal from noise. And whether the agents discovering AI can withstand the same scrutiny that human researchers have always faced—the demand for replication, for documentation, for proof that what looks like a breakthrough today won't turn out to be a mirage tomorrow.
Because if they can't, we're building on sand. And the foundation for the next generation of AI will be as fragile as the hype cycles that preceded it.
