Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

ai3Bio raises $48M to reset immune systems for remission

ai3Bio raises $48M to reset immune systems for remission
BiotechAutoimmune Disease+3
Healthtech & Biotech iconHealthtech & BiotechOctober 3, 2026

Halmos Labs automates biotech R&D with AI-driven design

Halmos Labs automates biotech R&D with AI-driven design
YcDrug Discovery+3
SaaS iconSaaSMarch 4, 2026

The AI Cost Crisis: How Context Compression APIs Aim to Cut Token Bills

The AI Cost Crisis: How Context Compression APIs Aim to Cut Token Bills
YcArtificial Intelligence+3
Climate / Social Tech iconClimate / Social TechMarch 4, 2026

AI Vision Systems Transform Fish Farming as Industry Races to Automate

AI Vision Systems Transform Fish Farming as Industry Races to Automate
Computer VisionAgtech+2

Founders Mentioned

Aayam Bansal

Synthetic Sciences

saas icon
SaaS

Ishaan Gangwani

Synthetic Sciences

saas icon
SaaS

Aayam Bansal

Synthetic Sciences

saas icon
SaaS

Ishaan Gangwani

Synthetic Sciences

saas icon
SaaS
Healthtech & Biotech iconHealthtech & Biotech
March 4, 2026
YcAi AgentsLab AutomationAi BenchmarkingAi Governance

The Race to Build AI Scientists That Can Run Research Alone

YC-backed Synthetic Sciences joins a growing field automating scientific research end-to-end. With benchmarks showing 17-92% accuracy ranges and EU AI rules looming, the stakes are high.

The Race to Build AI Scientists That Can Run Research Alone

The number landed like a challenge: 92% accuracy.

When Synthetic Sciences posted that figure to its Y Combinator profile this spring, claiming its biology mode had essentially solved BixBench Verified—a notoriously difficult bioinformatics benchmark—the claim stood out starkly against the field. The skepticism isn't personal. It's mathematical. Just a year earlier, when Future House and ScienceMachine first released BixBench in March 2025, the world's most advanced AI models were stumbling through real-world biology tasks with roughly 17% accuracy.

Now vendors are claiming they've quintupled that performance. Some of them, anyway.

The problem is that nobody can quite agree on what the numbers mean anymore, or even whether everyone's measuring the same thing. And that confusion matters more than it might seem, because pharmaceutical companies are making seven- and eight-figure infrastructure decisions right now, this quarter, based partly on which vendor demos look most impressive. Meanwhile, Europe's AI Act transparency requirements go live August 2—less than five months out—and American research institutions are still sorting out how NIH policies on AI-generated work should shape their grant strategies.

It's a mess, in other words. But it's a consequential mess.

When the Benchmark Became a Minefield

BixBench wasn't designed to be kind. The benchmark attempts to capture what computational biologists actually do: stitching together datasets from different sources, choosing appropriate analytical methods, interpreting results that don't come with answer keys. Messy, multi-step, cognitively demanding work.

Frontier models face-planted. 17% accuracy on open-answer tasks meant these systems were barely better than random guessing in some scenarios.

Specialized architectures started posting incremental gains. K-Dense Analyst, an autonomous bioinformatics system, reported 29.2% accuracy—a genuine improvement, though hardly revolutionary. More recently, Kepler AI claimed 33.4% on multiple-choice questions that included a "refuse to answer" option. Drylab listed 30% accuracy on its vendor page as of early 2026.

Then there's Synthetic Sciences at 92%. And BIOS by bio.xyz claiming scores above 64% in analysis mode.

Neither claim appears on an independent leaderboard. No third-party verification exists, at least none that's publicly documented. Which raises an uncomfortable question: Are these companies testing on different subsets of the benchmark? Using modified evaluation protocols? Or simply… optimistic?

The pattern suggests what one veteran AI researcher described to me, off the record, as "benchmark tourism"—vendors cherry-picking favorable conditions, then marketing the best number they can generate. It's not necessarily dishonest, but it's not science either.

A January 2026 paper titled "Why LLMs Aren't Scientists Yet" offers a more sobering picture. Researchers let autonomous systems attempt complete research pipelines—hypothesis through manuscript. Out of four attempts, only one produced a paper that got accepted to a workshop. Three failures, one success. The paper catalogs what went wrong: context degradation over long sessions, implementation drift, weak experimental design. The systems aren't incompetent. They're inconsistent, which might be worse.

The Infrastructure Gambit

Aayam Bansal and Ishaan Gangwani, both in their early twenties with competitive programming backgrounds, didn't set out to build just another AI wrapper. Their company, Synthetic Sciences, is making a specific bet about where the market is headed.

The product offers four modes. Research mode handles hypothesis framing through manuscript drafting. Biology mode specializes in wet-lab and computational biology—protein design, genomics, pathway analysis. Write mode produces publication-ready LaTeX documents.

But Flywheel mode reveals the real thesis.

In a LinkedIn post from early March 2026, the founders argued that research teams should "graduate" from renting frontier tokens to owning post-trained, task-specific models. The workflow they outline: start on GPT or Claude, collect production traces from real usage, fine-tune with supervised and reinforcement learning, then swap off the expensive frontier APIs for cheaper, faster, private alternatives.

It's a bet that today's foundation models are temporary scaffolding, not permanent infrastructure. And it's a pitch that resonates with pharmaceutical R&D directors watching their cloud bills. Running thousands of biology experiments through GPT-4 or Claude at current API pricing gets expensive, fast. Uncomfortably expensive.

The company raised $1.4 million in a late-2025 pre-seed round from Cory Levy at Z Fellows, Pioneer Fund, Amplo VC, and others, according to founder LinkedIn disclosures and December 2025 coverage, though major financial databases may show incomplete or variant totals. As of March 2026, pricing sits at $50 monthly for a Plus plan with 50 credits, $200 monthly for Pro with priority GPU access, and custom Enterprise deals offering unlimited credits and on-premise deployment.

The team? Still just the two founders, per the Y Combinator directory. That's either remarkably lean or a sign that scaling hasn't happened yet.

Adjacent companies are carving out different territory. Drylab prices its Lab plan at $149 per seat monthly. Emerald Cloud Lab operates a 105,000-square-foot remote-controlled lab facility in Austin—230-plus instrument types, entirely cloud-accessible. It's a hardware play for the automation wave. Benchling, once valued at $6.1 billion in November 2021, saw its valuation reportedly slashed to around $2.4 billion by September 2024, according to Contrary Research. The scientific software market, it turns out, can be brutal.

Then there's Isomorphic Labs, Alphabet's drug-discovery spinout, which has secured partnerships with Novartis and Eli Lilly worth nearly $3 billion combined—announced in January 2024 and expanded through early 2025. That's a completely different model: selling AI-powered insights directly to pharma giants rather than subscription software to individual labs.

The GPU Shadow Over Everything

Digital illustration for article section "The GPU Shadow Over Everything" in "The Race to Build AI Scientists That Can Run Research Alone" - A towering, abstract monolithic structure representing a next-generation GPU processor dominates the...

NVIDIA's hardware roadmap hangs over every conversation about AI research automation. The Rubin generation, announced in January 2026, promises up to 5× inference performance and 10× lower cost per token versus Blackwell. Deployments expected in the second half of 2026.

Those aren't marketing numbers. They're infrastructure decisions waiting to happen.

The U.S. Department of Energy and Argonne National Laboratory announced AI supercomputers with 100,000 NVIDIA Blackwell GPUs back in October 2025, explicitly positioned to accelerate scientific discovery. Recursion Pharmaceuticals' BioHive-2, with 504 H100 GPUs, ranked 35th on the Top500 supercomputer list as of May 2024—the largest pharma AI supercomputer at that time.

These aren't vanity projects. Running reinforcement learning environments for long-horizon agentic coding—the training approach Synthetic Sciences describes—requires serious GPU hours. The company's architecture integrations include Modal, Pinecone, and Weights & Biases, suggesting they're orchestrating compute across multiple clouds rather than owning hardware. Smart, perhaps necessary, but it introduces dependencies.

The economic calculation is straightforward enough. If Flywheel mode can genuinely produce task-specific models that match frontier performance at a fraction of the inference cost, the ROI could be substantial for labs running thousands of simulations monthly. If it can't, the entire pitch evaporates.

Regulation Arrives, Ready or Not

Digital illustration for article section "Regulation Arrives, Ready or Not" in "The Race to Build AI Scientists That Can Run Research Alone" - A modernist illustration representing the European Union's AI Act compliance timeline as a series of...

The European Union's AI Act entered force in August 2024, but compliance deadlines arrive in waves. General-purpose AI obligations went live August 2, 2025. The broader transparency rules—covering high-risk systems, disclosure requirements, energy reporting—take effect August 2, 2026.

Less than five months.

Any vendor marketing an "AI co-scientist" in Europe will need comprehensive documentation on model behavior, copyright compliance, and risk mitigation by summer. That's not guidance. That's law.

In the United States, the NIH issued policy effective for grant applications received September 25, 2025, and later: AI-substantially-developed applications are not considered original work. Reviewers cannot use AI to write critiques. The message is cautious but not prohibitive—disclose, document, proceed carefully.

Publishing norms are converging similarly. The International Committee of Medical Journal Editors and Nature both state that AI tools cannot be listed as authors and require disclosure of AI use. For a company promising to generate "publication-ready LaTeX drafts," that creates tension. The draft may be camera-ready, but journals expect human accountability for the claims inside.

A paper from May 2025 tested whether LLM agents could replicate findings from five published Alzheimer's disease studies. The agents achieved roughly 53% replication. Better than random, yes. Worse than a competent graduate student. For computational biology, where reproducibility crises have already damaged field-wide credibility, that number stings.

The Market Nobody Can Measure

Ask three market research firms to size "AI in drug discovery" and you'll get three wildly different answers.

Grand View Research, in a 2026 report, estimated the market at $2.35 billion in 2025, growing to $13.77 billion by 2033—a 24.8% compound annual growth rate. Precedence Research, publishing December 2025, said $6.93 billion in 2025, reaching $17.81 billion by 2035 at 9.9% CAGR. Vantage Market Research via PharmiWeb in March 2025 pegged the 2022 market at $1.3 billion, heading to $7.1 billion by 2030.

The divergence isn't sloppiness. It's definitional chaos.

Is "AI in drug discovery" just the software? Does it include consulting? What about pharmaceutical companies building in-house AI teams—are those dollars in the market or outside it? When analysts talk about "lab automation," are they counting robotic liquid handlers, or the software orchestrating them, or the AI agents writing protocols for them?

McKinsey's November 2025 State of AI report found that 62% of organizations are experimenting with AI agents. Yet only 39% report EBIT impact at the enterprise level. The gap between experimentation and value capture yawns wide. BCG's 2025 analysis, covering 1,250 companies, suggested roughly 5% are capturing meaningful AI value—with software, telecom, and fintech leading, while chemicals and fashion lag.

For AI research automation specifically, adoption will likely follow similar patterns. Early wins will come in computational domains where validation loops are tight and failure costs are low. Wet-lab integration, where a bad protocol wastes reagents and weeks of researcher time, will move considerably slower.

What the Academics Are Building

Researchers aren't waiting for startups to solve this.

The AI Scientist-v2 framework, detailed in an April 2025 paper, produced a fully AI-generated paper accepted to an ICLR workshop. The system introduced agentic tree-search and a manager agent to coordinate sub-tasks. Code is open-sourced. Anyone can replicate, extend, or critique it.

OmniScientist, proposed in a November 2025 paper, envisions co-evolving human and AI scientist ecosystems with collaboration protocols and evaluation via ScienceArena. Another framework, whimsically titled "freephdlabor" and published in October 2025, offers dynamic workflows, memory persistence, and non-blocking human intervention for continual science automation.

AutoLabs, presented September 2025, demonstrated near-expert procedural accuracy—F1 score above 0.89—in translating natural language instructions to automated liquid-handler protocols for complex syntheses. That's narrow, domain-specific, requires significant setup. But it shows the technical path forward for wet-lab automation exists.

These aren't products, though. They're proofs of concept. Research contributions advancing the state of the art by increments. The distance between a workshop paper and a commercial system is measured in engineering hours, reliability testing, customer support infrastructure, and unglamorous debugging.

It's work. Lots of it.

What Happens Next

Digital illustration for article section "What Happens Next" in "The Race to Build AI Scientists That Can Run Research Alone" - A conceptual, modernist illustration representing the critical validation phase for AI scientist sta...

The next twelve months will clarify whether the current wave of AI scientist startups can deliver on their benchmark claims in production environments. Synthetic Sciences, Drylab, Kepler, BIOS, and others will either validate their accuracy numbers with independent third-party testing or face erosion of credibility. The EU's August 2026 compliance deadline will force transparency in model documentation and energy reporting—potentially advantaging companies with simpler, more explainable systems over black-box architectures.

For R&D leaders making platform choices this quarter, the risk calculus cuts both ways. Adopt too early and you may lock in immature technology requiring costly rework later. Wait too long and competitors using AI-accelerated discovery may build insurmountable data advantages. The middle path—controlled pilots with clear success metrics and short evaluation cycles—is prudent. It's also unexciting, which is perhaps why fewer executives choose it than probably should.

The hardware trajectory, at least, is easier to predict. NVIDIA's Rubin generation will deliver better inference economics in the second half of 2026. Frontier model providers will continue their benchmark leaderboard races. Fine-tuning workflows will become more accessible, making the "graduate from frontier APIs" thesis more plausible for well-resourced labs.

What remains genuinely uncertain is whether autonomous AI scientists will reach the reliability threshold that scientific research demands. A 92% accuracy claim, if independently verified, would be transformative. A 33% success rate on bioinformatics tasks is interesting but not yet trustworthy for consequential decisions.

The gap between those numbers is where the real competition is happening—not in marketing copy or conference demos, but in training pipelines, evaluation rigor, and the unglamorous work of debugging failure modes that only surface in production, under time pressure, when the experiment matters.

The race isn't over.

It's possible—quite possible—that it's only now truly beginning.

More stories

  • ai3Bio raises $48M to reset immune systems for remission
  • Halmos Labs automates biotech R&D with AI-driven design
  • The AI Cost Crisis: How Context Compression APIs Aim to Cut Token Bills
  • AI Vision Systems Transform Fish Farming as Industry Races to Automate
  • Ditto Bio Mines Parasite Evolution to Design Autoimmune Therapies
  • YC-Backed Salus Launches Pre-Execution Guardrails for AI Agents
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.