The pipette sat untouched. For seventeen days straight at Berkeley Lab, robots synthesized novel materials—more than two per day—while software designed experiments, analyzed results, and proposed the next round of tests. No graduate student pulled an all-nighter. No principal investigator reviewed protocols at 2 a.m.
Then, in February 2025, Google Research unveiled something that felt less like a tool and more like a colleague: a multi-agent system that generates hypotheses, designs experiments, debates itself, and refines protocols before any human intervention. Built on Gemini 2.0, the "AI co-scientist" validated itself by recreating published findings in acute myeloid leukemia drug repurposing, then pivoted to propose fresh hypotheses around antimicrobial resistance. Berkeley's autonomous A-Lab had already demonstrated the concept with materials synthesis. Google's system suggested the paradigm works for biology, too.
The timeline compression is jarring. What once consumed months now happens in weeks. Sometimes days.
The question facing research institutions isn't whether AI will reshape scientific R&D—that much seems settled. It's how quickly labs can retool their people, processes, and infrastructure. And whether the science itself remains reproducible when machines start doing the thinking.
From Narrow Tools to End-to-End Autonomy
The shift is well underway, though unevenly distributed. DeepMind's GNoME model predicted 2.2 million new crystal structures in November 2023; roughly 380,000 were flagged as stable, and Berkeley's A-Lab validated dozens autonomously through synthesis and characterization. That pipeline—hypothesis generation, experimental design, autonomous execution, analysis, reporting—represents a fundamentally different R&D cadence.
Materials science got there first. Life sciences is catching up fast, driven by money and impatience.
McKinsey estimates generative AI could unlock $60 billion to $110 billion annually across pharma and medtech. BCG's modeling goes further: AI-first biopharma companies can compress early discovery from the typical four-to-five-year slog down to roughly eight months, with potential revenue uplifts of 5 to 15 percent at scale. The caveat, repeated in multiple surveys from 2024 and 2025, is that most firms remain stuck in what consultants politely call "pilot purgatory." Scaling gaps aren't algorithmic—they're organizational. A company can license the best model in the world and still fail if its scientists don't trust it, or if its data sits in siloed spreadsheets.
Isomorphic Labs, Alphabet's drug discovery spinout, signed multi-target collaborations with Eli Lilly and Novartis in early 2024, with total deal values approaching $3 billion. The partnerships hinge on next-generation AlphaFold capabilities and the promise of "biology-wide" modeling—a term that sounds like marketing until you see what the system can do with protein structures. In May 2024, Recursion Pharmaceuticals completed BioHive-2, a 63-DGX H100 supercomputer that ranked 35th on the TOP500 list. It's the largest pharma-owned compute cluster on record. Recursion is using it to train foundation models and agentic systems for biology.
Compute, it turns out, is becoming as much a competitive moat as compound libraries. Perhaps more.
Under the Hood: How AI Co-Scientists Actually Work
Google's architecture illustrates where the field is heading, though few competitors have disclosed this level of technical detail. The system deploys six specialized agents—Generation, Reflection, Ranking, Evolution, Proximity, and Meta-review—coordinated by a Supervisor agent. Each handles a discrete function: one generates hypotheses from literature, another critiques them, a third ranks by feasibility. The system runs self-play debates and uses Elo-style auto-evaluation to refine outputs before surfacing them to human experts.
Google emphasizes what it calls "test-time compute scaling," meaning the model invests more computational resources when the problem demands it, rather than relying solely on training-time scale. It's the difference between a sprinter and a marathoner: some problems need a quick answer, others need the machine to think longer.
Edison Scientific's Kosmos, announced in November 2025 and backed by a $70 million seed round, takes a different approach. The system—spun out of FutureHouse, the research nonprofit funded by Eric Schmidt—uses structured world models to maintain coherence across tens of millions of tokens. In technical benchmarks, Kosmos reportedly reads around 1,500 papers per run and executes roughly 42,000 lines of code, producing outputs the company claims are equivalent to months of human effort. An independent audit found statement accuracy near 79.4 percent. The team has published seven claimed discoveries spanning metabolomics, neuroscience, materials science, and genetics, though peer validation is ongoing.
Whether those discoveries hold up matters more than the benchmark scores. The field has seen too many hyped papers crumble under replication attempts.
Autonomous and cloud labs are maturing in parallel. Emerald Cloud Lab operates software-controlled life science facilities accessible remotely; Carnegie Mellon launched the first university cloud lab in partnership with ECL, positioning AI-driven protocol generation as a core use case. Arctoris acquired Eli Lilly's automated Life Science Studio in 2024, relocating the platform to Oxford to expand its Ulysses end-to-end biology and machine learning offering. Opentrons has pushed no-code bench automation—its Flex robot now supports over 100 protocols, with AI-assisted protocol generation tools rolling out in 2024.
Lab informatics platforms are embedding agents directly into workflows. Benchling launched "Benchling AI" in October 2025, integrating NVIDIA NIM and BioNeMo models into its electronic lab notebook and LIMS infrastructure. Sanofi deployed the system across governed R&D data earlier in 2024. The pitch is straightforward: rather than bolt AI onto legacy systems, make the ELN itself the orchestration layer for foundation models, domain models, and agentic reasoning.
It sounds clean in a press release. Implementation is messier.
Three Converging Forces

The data bottleneck is easing, though not fast enough for everyone's taste. The Materials Project, Berkeley Lab's open database of computed properties, now includes roughly 380,000 likely-stable materials from GNoME predictions, feeding both human researchers and autonomous synthesis loops. In life sciences, platforms like Scite have indexed 1.4 billion citation contexts, enabling AI agents to trace claims back to original evidence and flag retractions or contradictions. NIH's Data Management and Sharing policy, effective since January 2023, mandates DMS plans for funded research—creating a compliance tailwind for structured, machine-readable datasets.
The irony: scientists who spent careers resisting standardized data formats are now racing to clean up their lab notebooks because the AI won't work otherwise.
Hardware economics are shifting in parallel. Recursion's decision to build BioHive-2 in-house signals that foundation model training for biology is no longer the exclusive domain of tech giants. Smaller biotechs can't match that scale, but partnerships with cloud labs and modular automation vendors—Opentrons, for instance—lower the barrier to closed-loop experimentation. The University of Toronto's Acceleration Consortium, backed by a $200 million grant in 2023, projects self-driving labs can cut materials discovery timelines from roughly 20 years and $100 million to one year and $1 million.
If that ratio holds, the calculus around vertical integration changes quickly. So does the talent war: do you hire more chemists or more machine learning engineers?
Regulatory clarity is improving, though unevenly. The EU AI Act exempts research and development activities from many provisions, but deployment in the EU triggers obligations. In the U.S., the Biden administration's Executive Order 14110 on AI was revoked in January 2025, replaced by an order titled "Removing Barriers to American Leadership in Artificial Intelligence." The FDA issued draft guidance in 2025 on AI and machine learning in drug development, noting it had received over 500 AI-containing submissions between 2016 and 2023. NIST's AI Risk Management Framework, released in 2023 with a Generative AI Profile in 2024, is being adopted as a de facto standard by procurement teams, even as revisions remain in progress.
Translation: companies operating globally will default to the strictest standard—likely a blend of EU compliance and NIST RMF alignment—because fragmented regulatory strategies don't scale. And lawyers are expensive.
What Deployment Actually Looks Like
Berkeley's A-Lab integrated computational targets from the Materials Project with real-time synthesis and characterization. The work sparked peer debate—some questioned whether the discoveries were genuinely novel or rediscoveries of known phases—but the paradigm held: a machine proposed candidates, designed experiments, executed them, and validated results without waiting for a human scheduler.
Google's co-scientist isn't publicly available yet; the company is running a trusted tester program with academic and industry partners. What's notable is the architecture's modularity. The Reflection agent critiques the Generation agent's output, the Ranking agent scores feasibility, and the Evolution agent iterates on weak proposals. It's less "ask a chatbot" and more "simulate a research team." One where nobody takes vacation or argues about authorship order.
Benchling's enterprise deployments show what integration looks like at scale. Sanofi uses Benchling AI for data entry, protocol composition, and deep literature searches—tasks that previously consumed junior scientist time. The system pulls from governed internal datasets, applies domain-specific models via NVIDIA NIM, and logs everything for audit trails.
That last piece matters. Pharma R&D operates under 21 CFR Part 11 and ALCOA+ data integrity principles. AI outputs need provenance chains, not black boxes. When the FDA auditor shows up, "the algorithm said so" isn't a satisfying answer.
The Reproducibility Problem Nobody Wants to Talk About

A 2023 Nature editorial warned that AI-generated scientific content risks embedding errors, fake citations, and subtle biases at scale. The concern isn't hypothetical. A 2026 study analyzing 41.3 million papers found that AI-augmented researchers publish and cite more, but collective research diversity contracts—data-rich areas get preferential attention while sparser domains stagnate.
The worry isn't that AI hallucinates, though it does. It's that well-funded labs will chase machine-legible problems, leaving harder questions underexplored. If your autonomous lab can synthesize battery materials efficiently but struggles with complex organic molecules, guess which research program gets expanded?
Validation infrastructure is nascent. HeurekaBench, a benchmark for end-to-end AI co-scientists in single-cell biology, was proposed in January 2026 to address this gap. The field needs shared tasks, datasets, and wet-lab validation protocols to compare Google's system against Kosmos against open-source alternatives. Without that, every vendor claims state-of-the-art on proprietary benchmarks, and customers have no neutral ground truth.
Regulatory flux adds uncertainty. The EU AI Act's research exemptions don't extend to deployment, so a model trained in an academic lab faces compliance obligations when a company uses it for submissions. The U.S. revocation of Executive Order 14110 created jurisdictional variance: some states may adopt stricter standards, others may defer to federal guidance that's now in flux. FDA's draft guidance signals receptivity to AI in drug development, but the agency has yet to clarify how it will handle fully autonomous experimental design—especially when the machine proposes a protocol no human reviewed in advance.
IP questions remain unresolved. If an AI agent generates a hypothesis that leads to a patentable compound, who owns it? The lab that ran the agent, the vendor that built the model, or the institution that provided the training data? Procurement teams are demanding answers. Legal frameworks are lagging.
What Comes Next—and What Could Go Wrong

The near-term trajectory seems clear enough. More pilots will scale, more cloud labs will come online, and more ELN/LIMS platforms will embed agentic capabilities. BCG projects that end-to-end AI adoption can cut preclinical timelines by 30 to 50 percent and reduce costs by 25 to 50 percent, but only for firms that retool people and processes—not just plug in a model.
The bottleneck, again, is organizational. You can't automate culture change.
Longer term, three to five years out, we may see autonomous labs that anticipate scale-up challenges during early discovery, looping in manufacturing constraints before a candidate hits IND-enabling studies. Materials science is already heading there; ACS Omega outlined a roadmap in 2025 for "factory-to-lab integration" in self-driving labs. Pharma will follow if the economics hold and the regulators don't flinch.
Benchmarking will mature. Expect consortia to publish head-to-head comparisons of Google's co-scientist, Kosmos, and open alternatives on shared datasets with wet-lab validation. When procurement teams can point to a leaderboard that tracks reproducibility, citation accuracy, and experimental success rates—not just speed—adoption will accelerate. Maybe.
Policy will stabilize, unevenly. NIST is revising the AI Risk Management Framework, and the FDA will finalize guidance on AI in drug development. The EU's AI Act implementation timelines will clarify what "deployment" means for research tools.
The University of Toronto's Acceleration Consortium framed the goal as cutting discovery timelines by an order of magnitude. The data suggest we're partway there, though whether materials science success translates to complex biology remains an open question.
Whether we maintain scientific rigor at machine speed is the experiment still running. No control group, no stopping criteria, and the entire research enterprise as the test subject.
The pipette still sits on the bench. Increasingly, though, nobody's picking it up.
