Three years back, if you wanted artificial intelligence to assist with research, you mostly got a glorified literature summarizer—decent enough for skimming papers, but hardly revolutionary. Now? You can, in theory, point an AI agent at a research hypothesis and step away while it plans experiments, coordinates computing resources, drafts protocols, even writes up findings. The whole thing, supposedly, while you're asleep.
Whether that vision is oversold remains an open question. But the shift underway isn't purely about what the technology can do—it's about what investors think it will do. When venture capitalists write million-dollar checks to teenage founders promising "AI co-scientists," and when self-driving labs claim to boost data throughput by a factor of ten, you know something fundamental about the research playbook is being rewritten. Or at least someone's trying awfully hard to rewrite it.
From Literature Search to Wet-Lab Work
Google made waves in early 2025 with an AI system that moved past the usual parlor tricks of summarizing academic papers. The tool—unveiled in February—generated its own hypotheses for repurposing drugs against acute myeloid leukemia, then validated those ideas in actual cell lines. It identified epigenetic targets for liver fibrosis and tested them in human hepatic organoids. Perhaps most striking, it independently rediscovered a mechanism of antimicrobial resistance gene transfer, timing its findings alongside a separate human research team grinding through the same question. These weren't simulations. They were wet-lab results, published with a "Trusted Tester" program for select organizations.
By early 2026—or so the pitches went—the market had grown crowded. Synthetic Sciences, a Y Combinator-backed startup that used to go by InkVell, launched with four agent "modes" it claimed could handle everything from hypothesis framing to manuscript drafting. The company's Biology mode reportedly scored 92% on something called BixBench Verified, a benchmark for bioinformatics agents, according to a Y Combinator social media post. That figure hasn't surfaced in peer-reviewed literature, which tells you something about how fast this sector moves—and how much runs on momentum rather than rigorous validation. Pricing starts at $50 a month for individuals, which feels almost quaint given the ambitions on display.
The two co-founders, Ishaan Gangwani and Aayam Bansal, joined Y Combinator at seventeen or eighteen after raising a $1.4 million pre-seed round. Their investor roster reads like a contact list from an Andreessen Horowitz networking event: Y Combinator, Pioneer Fund, Amplo VC, Pareto Holdings, Charlie Songhurst, Walter Kortschak of Firestreak Ventures, a16z scout Yana Welinder. In a LinkedIn post from March 2026, Bansal outlined plans to move from renting frontier APIs to owning post-trained models through the company's "Flywheel mode"—a play to improve margins, speed, and control.
They're not alone. Elicit rolled out Research Agent workflows late last year. Perplexity's Comet agentic browser flagged "Learning & Research" among its top use cases in internal studies. A January 2026 paper in Nature Computational Science introduced SciSciGPT, a multi-agent system for "science of science" workflows. The pattern repeats: agents graduating from assistants to something closer to collaborators, though "collaborator" may be generous depending on whom you ask.
The Convergence

Three forces are coming together, and they're making the case—at least on paper—for a new kind of research infrastructure.
First, compute. Multi-cloud GPU marketplaces like Prime Intellect now aggregate resources from providers like Nebius and Akash, making it cheaper to spin up the persistent, stateful sandboxes that long-running research agents need. The National AI Research Resource Pilot handed out its first 35 grants in mid-2024, broadening access to compute and datasets for U.S.-based researchers. Synthetic Sciences integrates directly with GitHub, Hugging Face, Weights & Biases, Modal—essentially betting that the research stack is standardizing fast enough that agentic orchestration becomes feasible.
Second, the models themselves. DeepSeek-R1, a 671-billion-parameter reasoning model released in January 2025, became a popular baseline for fine-tuning on scientific tasks. AlphaFold 3, announced in mid-2024, remains foundational to AI-in-biology narratives. Isomorphic Labs' deals with Eli Lilly and Novartis—valued at nearly $3 billion when announced in early 2024—are reportedly moving into first trials. NVIDIA's internal surveys suggest that around three-quarters of pharma and biotech companies are now using AI actively, with generative AI the top workload.
Third, self-driving labs are no longer just demos. Researchers at NC State, UNC, and Tec de Monterrey published work in Nature Chemical Engineering in mid-2025 showing a dynamic-flow system that increased materials discovery data rates roughly tenfold. The system could nail top candidates on its first try after training. The A-Lab at Lawrence Berkeley National Laboratory demonstrated autonomous inorganic powder synthesis with a 67% success rate in work published in Nature back in late 2023, though the findings still get cited frequently. Emerald Cloud Lab has been running a fully remote, automated lab using its open-source Symbolic Lab Language since 2023. Opentrons liquid handlers—basic models run around $5,000, according to vendor materials—have become a low-cost staple in many pipelines.
The economic logic is straightforward, even if the execution isn't. Compress discovery cycles, cut labor costs, and R&D margins improve. Market analysts seem convinced, at least in aggregate. Precedence Research pegs the AI-in-drug-discovery market at nearly $7 billion in 2025, growing to almost $18 billion by 2035. A different estimate from 360ResearchReports places the "Early Drug Discovery" segment at just over $1 billion in 2025, ballooning to more than $12 billion by 2034—a compound annual growth rate north of 30%. Lab automation more broadly sits around $6.2 billion in 2025, with a CAGR near 5.5% through 2035, per Expert Market Research. Numbers vary by methodology, naturally. But the trajectory is consistent.
Real Money, Real Bets

Lila Sciences, a company out of Flagship Pioneering unveiled in March 2025, announced a Series A and extension totaling roughly $350 million by October of that year. Third-party reports cite a valuation near $1.3 billion. The pitch? An "AI Science Factory" for discovery at scale. Chemify, which raised $43 million in mid-2023 and another $50 million in late 2025, is automating chemistry via what it calls "chemputation"—digital synthesis protocols executed by robots.
Autoscience, a startup billing itself as an "automated AI research lab," raised a $14 million seed led by General Catalyst in March 2026, according to Axios. The recursive angle here—AI agents building better AI models—is either elegantly meta or circular, depending on your level of skepticism.
Then there's the validation work Google published. The AML drug repurposing study showed tumor viability inhibition at clinically relevant concentrations. The liver fibrosis targets are reportedly being prepped for full publication by collaborators. The antimicrobial resistance hypothesis lined up, independently, with findings from human researchers working the same problem—a signal the system was doing something closer to reasoning than just pattern-matching. Perfect? No. But real.
Academic frameworks are proliferating at a pace that suggests either genuine innovation or a land grab. OmniScientist, a multi-agent end-to-end science system, appeared as a preprint in late 2025. Paper2Agent and AutoLabs, both from September 2025, focus on translating natural language into protocols for liquid handlers. TxAgent targets therapeutic discovery. AISAC, out of Argonne National Laboratory, launched in November 2025. The cadence hints at research groups racing to plant flags in different verticals before someone else does.
BixBench, a benchmark published around early 2025 by FutureHouse, tracks 53 scenarios and 296 research questions designed to test real analytical pipelines, not just Q&A. Vendor-reported scores are all over the map, which is perhaps inevitable when there's no centralized, independent leaderboard. K-Dense claims scores above 90% on a 50-item "verified" subset. Edison Analysis reported 46% on the full benchmark. BIOS claims the top spot. Synthetic Sciences' 92% on "BixBench Verified" comes from that Y Combinator social post, not a paper. Adoption decisions, then, rest on trust—or on running your own proof-of-concept trials.
The Gap Between Demo and Deployment

The distance between a splashy demo and a deployed system remains considerable, and experts are quick to point that out. A TechCrunch piece from March 2025 quoted researchers skeptical of "AI co-scientist" readiness, noting the grunt work real research demands versus pure reasoning exercises. A McKinsey report from late 2025 framed "scientific AI" as a once-in-a-century R&D productivity lever but stressed the need for responsible scaling. MIT Technology Review argued in December 2025 that AI materials discovery must escape the simulation and enter the real world—a necessary, if unglamorous, next step.
Policy is catching up unevenly. The EU AI Act, approved in mid-2024, grants research exemptions that reduce friction during R&D, but deploying systems beyond research triggers transparency and risk management obligations. The NIH banned generative AI in peer review back in mid-2023 and reiterated AI-use guidance for grant applications in mid-2025. A preprint on "Safe-SDL" from February 2026 proposed architectural safety mechanisms for AI-driven self-driving labs, highlighting model failures on a benchmark called LabSafety Bench.
Certain domains will likely see full automation before others. Materials and chemistry, with closed-loop experimentation and decent simulator availability, are easier to automate than complex in vivo biology. But infrastructure gaps remain stubborn: equipment interoperability, metadata standards, reproducibility frameworks. A Nature Communications paper from April 2025 cataloged these roadblocks in detail, and a Royal Society Open Science review from mid-2025 examined policy and workforce implications.
The companies worth watching are those that can tie agents to real experimental capacity. Lila Sciences is building vertically integrated "science factories." Chemify owns its chemputation platform. Emerald Cloud Lab runs the hardware. Synthetic Sciences, by contrast, positions itself as an orchestration layer—integrations across literature, code, compute, and writing, with optional biology workflows. Whether that approach scales hinges on whether third-party lab infrastructure becomes commoditized enough to orchestrate remotely. That's a bet, not a certainty.
Benchmark overfitting is an obvious risk. BixBench is already spawning "verified" subsets to counter prompt leakage, and AIRS-Bench, published in early 2026, attempts to set standards for "Frontier AI Research Science Agents." Expect more gold-standard tasks, human-in-the-loop validation, and third-party audits as vendors compete on state-of-the-art claims. The industry will need independent evaluation to separate signal from sales deck.
The Economics Are Shifting
For founders and R&D leaders, the calculus is moving beneath their feet. If an agent can compress a three-month literature review into three days, plan experiments while you're offline, and draft a manuscript overnight, the fixed costs of research start looking variable. The question isn't whether to experiment with these tools—most serious players already are. It's how quickly to weave them into core workflows, and how to avoid locking into platforms that may not survive the next funding cycle.
The teenage founders at Synthetic Sciences may or may not build the system that wins. But they're correct about one thing: the economics of R&D are being rewritten, one agent at a time. Whether those rewrites hold up under real-world pressure, or collapse into another wave of overhyped infrastructure that looked brilliant on a pitch deck, remains to be seen. The labs, after all, still need to deliver results—not just demos.
