For 17 days in 2023, nobody checked on the robots at Lawrence Berkeley National Laboratory. The A-Lab—a self-contained system melding machine learning, robotic arms, and X-ray diffraction—worked through a shopping list of 58 predicted inorganic materials. When researchers finally looked at the results, 41 new compounds existed that hadn't before.
No adjustments mid-experiment. No human troubleshooting. Just a machine making decisions about what to synthesize next based on what it had learned from what came before.
This wasn't a carefully stage-managed demo for a press release. It was Tuesday.
Across the country at Carnegie Mellon, another system called Coscientist was doing something arguably stranger. Built around GPT-4, it could read scientific papers, design experiments, write the code to run them, then physically execute the chemistry. The research landed in Nature—a journal not known for publishing vaporware.
The common thread? These systems close loops. Hypothesis to validation, no humans in between. And that represents something more consequential than faster pipetting or better literature searches. It's a restructuring of how discovery itself gets done.
The numbers tell part of the story. Lab automation hit $8.27 billion in 2024, per Grand View Research, with forecasts stretching toward $18.39 billion by 2033. But the hardware spend misses the deeper shift: machines that can plan experiments, analyze what went wrong, and adjust their approach. Not tools. Colleagues, perhaps more than the people building them initially expected.
Where the Rubber Meets the Reagent
Adoption patterns are all over the map. Benchling's 2026 Biotech AI Report—surveying actual users, not aspirational adopters—found that 76% deploy AI for literature review. Seventy-one percent use it for protein structure prediction. Sixty-six percent for scientific reporting. Target identification, a higher-stakes activity where mistakes compound quickly, clocks in at 58%.
Those aren't fringe tasks. They're the backbone of research work.
The divide between pharmaceutical giants and smaller biotechs remains brutal. Large pharma companies adopt AI at roughly three times the rate of their smaller competitors—67% versus 23%, according to Benchling's 2024 State of Tech in Biopharma study. For the big players, AI has become a top-two investment priority. They have the capital. They have the data infrastructure. They can afford to be wrong a few times while figuring it out.
McKinsey pegs the potential value of generative AI in pharma at $60 billion to $110 billion annually. Broader analyses suggest $360 billion to $560 billion across R&D sectors. Those estimates explain why one-third of researchers now use ChatGPT for work, per a 2024 Elsevier survey, and why 74% believe the technology will fundamentally alter their field.
Yet a 2023 Nature poll revealed something telling: 80% of researchers had tried AI chatbots. Trying something is easy. Actually integrating these tools into validated workflows—where reproducibility, traceability, and regulatory compliance aren't optional—that's where most efforts still stall.
Proteins, Predictions, and the Nobel Committee's Stamp
When DeepMind released AlphaFold 3 in May 2024, expanding protein structure prediction to include ligand and nucleic acid complexes, the scientific community understood this wasn't incremental progress. The Nobel Committee agreed. Demis Hassabis and John Jumper received the prize later that year.
The server opened for non-commercial research, though debates persist about whether it's truly as "open" as AlphaFold 2 was. Openness, it turns out, means different things when commercial interests sharpen.
Self-driving labs are moving beyond proof-of-concept. These closed-loop systems—selecting experiments, executing them robotically, analyzing results, then deciding what to try next—are appearing in multiple institutions. Berkeley's A-Lab isn't an outlier anymore. Reviews in journals like Digital Discovery now focus on data fusion standards and benchmarking frameworks, the kind of infrastructure conversations that signal a field maturing past the novelty phase.
Carnegie Mellon's Coscientist represents a different architecture: the orchestrating agent. Instead of replacing a specific instrument, it coordinates across literature databases, planning tools, code environments, and physical hardware. It reads, reasons, plans, acts—all through natural language interfaces and API calls.
What these systems share is operational autonomy. They don't assist; they decide. That shift—from copilot to colleague—changes the economics of research labor in ways we're only starting to map.
The Infrastructure Land Grab

NVIDIA saw this coming earlier than most. Its BioNeMo platform is positioning itself as foundational infrastructure for scientific AI. In January 2026, the company announced partnerships with Eli Lilly for a co-innovation lab and Thermo Fisher for autonomous lab systems. The pitch: "agentic physical AI" connecting foundation models to real lab workflows. NVIDIA's Isaac and Cosmos platforms feed execution data back into model training, tightening the learning loop.
Opentrons—10,000-plus lab robots deployed globally—partnered with NVIDIA in February 2026 to embed what they're calling "physical AI" capabilities. OpentronsAI generates protocols from natural language prompts, learning from each run. It's not just automation. It's systems that get better through use.
Benchling threw down a marker at its BenTalk 2025 event, declaring itself the "command center for scientific AI." Integrations with NVIDIA NIM microservices, including BioNeMo and OpenFold2, followed. Enterprise clients scaled up: Sanofi (1,500 scientists), Moderna, Eli Lilly running TuneLab models in-house. The platform is becoming the data substrate where agents operate, with governance layers baked in from the start.
Thermo Fisher's October 2025 partnership with OpenAI targets R&D and clinical operations across the board. These aren't pilots. They're operating system plays, positioning AI as the layer that scientific work runs on top of.
Emerald Cloud Lab offers yet another model: over 200 instrument types, available 24/7, remotely controlled. Experiments as API calls. Arctoris acquired Eli Lilly's remote robotic Science Studio—originally built with Strateos—to scale its Ulysses platform for automated wet-lab biology. The trend is unmistakable: labs are becoming programmable, networked, increasingly agent-ready.
CoreWeave, which went public in 2025, acquired Weights & Biases and locked in multi-year contracts with OpenAI. It's deploying early GB200 and GB300 clusters—the GPU infrastructure these models actually run on. Emerging marketplaces like TensorPool, Prime Intellect, and Shadeform offer fractional compute access, lowering barriers for smaller teams. The stack is both consolidating and fragmenting into specialized layers.
The Easy Wins and the Hard Slogs

Literature triage works. So does procedure drafting, protein structure prediction, scientific reporting. Benchling's 2026 data shows these as established use cases, not experiments. Target identification, at 68% adoption, has crossed into production territory.
The harder problems? Generative design in regulated environments where validation burden is crushing. Biomarker discovery with scattered, multimodal data. ADME-tox modeling, where datasets fragment and error costs spike. Orchestration becomes critical here—stitching together partial datasets, reasoning across modalities, maintaining provenance through every step.
The shift from copilots to full orchestration is accelerating faster than many anticipated. Multi-agent systems that coordinate experiments end-to-end are moving from research prototypes to production tools. AutoGen v0.4, LangGraph prebuilts, OpenAI's Agents SDK (replacing the earlier Swarm project)—these frameworks handle coordination: which agent does what, how information flows, when to escalate to humans.
Synthetic Sciences, which emerged from Y Combinator's Winter 2026 batch, offers a glimpse of where this trajectory leads. Operating under Inkvell Inc., the company built multi-mode agents that scope hypotheses, search PubMed and bioRxiv, provision GPUs, train and evaluate models, then generate publication-ready LaTeX. The company claims state-of-the-art performance on BixBench biology questions. Pricing starts at $50 monthly for Plus, $200 for Pro, with custom enterprise deals.
The founders are teenagers. They raised $1.4 million in pre-seed funding from Z Fellows, Pioneer Fund, a16z Scout, and others.
Their system integrates across 12-plus services: Tinker, Modal, Prime Intellect for compute orchestration; Pinecone for vector search; LangSmith for tracing; Hugging Face and Weights & Biases for model management; Together, Groq, Fireworks for inference. The architecture reflects a broader industry pattern: agents as orchestration layers, not monolithic systems.
When the Regulators Wake Up
Regulation is tightening, sometimes awkwardly. The EU AI Act's core obligations take effect August 2, 2026. Penalties reach €35 million or 7% of global revenue, whichever stings more. High-risk AI systems embedded in research workflows face documentation, logging, and human oversight requirements that many labs aren't remotely prepared for. Debates about grace periods surfaced in late 2025, but the direction is set.
The NIH issued guidance in July 2025: grant applications "substantially developed by AI" will not be considered. Peer reviewers remain prohibited from using generative AI tools, a ban standing since June 2023. Major publishers—Science, Nature, JAMA—restrict AI authorship and require disclosure. Recent editorials in February 2026 reinforced the position: AI cannot be a co-author. Humans must own the work.
Which raises interesting questions about what "ownership" means when the AI did most of the hypothesis generation, all of the literature review, the experimental design, the execution, and drafted the manuscript.
The National Academies called in May 2024 for human accountability frameworks in AI-assisted research, proposing a Strategic Council on Responsible AI in Science. Ethics scholarship emphasizes standardized disclosures and responsible use guidelines. The regulatory apparatus is catching up, but it's running behind.
NIST's AI Risk Management Framework, including the Generative AI Profile released in July 2024, offers voluntary guidance on risk management, testing, evaluation, validation. The AI Risk and Compliance resources provide implementation support. These aren't binding. But they shape best practices and foreshadow future regulation.
The practical challenge: co-scientist outputs must be traceable, auditable, human-verified. Data lineage matters. Provenance matters. The gap between "AI helped" and "AI did it" is collapsing faster than governance systems can adapt.
What Comes Next (And Who Decides)

Demis Hassabis, in his Nobel interview, called AI for science "the ultimate tool." He acknowledged gaps in continual learning, long-term planning, consistency. Meta's Yann LeCun advocates for open-source trajectories and systems that reason, remember, understand the physical world. Both see outsized impact in biomedical and materials discovery.
The next 18 months will likely accelerate what's already visible. Agentic and physical AI integrations are expanding across Opentrons/NVIDIA partnerships, Thermo Fisher initiatives, BioNeMo microservices adoption. Benchling's ecosystem embeds agents directly into electronic lab notebooks, LIMS, and scientific data management systems, with governed data access built in.
Self-driving labs will mature with standardized benchmarks and metadata ontologies. Hybrid human-AI orchestration will move from novel to normal in materials and chemistry workflows. EU AI Act pressure will force reproducibility, logging, and data governance into scientific AI stacks—potentially accelerating consolidation between instrumentation vendors, orchestration software providers, and AI model developers.
"Venue experiments" like the Agents4Science workshop signal that the scientific community is attempting to evaluate AI-authored research in controlled settings. Cross-domain validation remains immature, but the need is recognized.
The technical barriers are narrowing. Models improve. Tooling matures. Infrastructure consolidates. The real constraints are organizational and regulatory. BCG and McKinsey studies show that only a minority of firms capture real value from AI. Success requires leadership commitment, reimagined workflows, robust data foundations, scaled operating models. The firms that solve governance, data plumbing, and change management will pull decisively ahead.
The lab of 2030 won't look unrecognizable. Same benches, probably. Same glassware. But the decision-making, experiment design, hypothesis generation will increasingly route through autonomous agents reasoning across literature, data, physical constraints.
Not replacing scientists, exactly. Expanding what's tractable. The question isn't whether this happens—that ship sailed somewhere between AlphaFold 2 and the A-Lab's 17-day sprint. The question is who builds the infrastructure, who sets the standards, and who captures the value as the loop finally closes.
