The numbers are familiar to anyone who's spent time around pharmaceutical development, but they still sting. Nine out of ten experimental drugs fail before reaching patients. Not nine out of twenty, not half. Nine out of ten. And that's after surviving years of preclinical work and animal testing—tests that, according to industry veterans, "can't predict how a human will respond" no matter how promising they look in the lab.
The math gets worse the deeper you go. From Phase I to final approval, success rates hover around 7.9%, based on data tracking the 2011-2020 period. For every compound that clears the regulatory hurdle, dozens more collapse under the weight of ambiguous trial results and the tens of millions spent pursuing them. Phase 2 studies alone carry estimated median costs of $8.6 million; Phase 3 can run north of $21 million.
Now a clutch of startups and established players are pitching a different approach altogether. Using what they call "foundation models"—massive neural networks trained on millions of biological samples—these companies claim they can predict which patients will actually respond to a drug before enrollment begins. The promise: cut trial populations by 70% or more, potentially rescuing compounds that failed not because they didn't work, but because researchers tested them on the wrong people.
It's an audacious claim. Whether it holds up under scrutiny is another matter entirely.
When Biomarkers Don't Work
The problem these AI tools aim to solve is both straightforward and maddeningly complex. Traditional biomarkers, the molecular signals doctors use to match patients with treatments, often perform only marginally better than guesswork.
Consider PD-L1 testing, a standard checkpoint in cancer immunotherapy. Meta-analyses peg its predictive power—measured as area under the curve, or AUC—at roughly 0.6 to 0.67. For context, a coin flip scores 0.5. Tumor mutational burden, another widely used marker, fares no better. These tests became standard of care not because they excel at their job, but because the alternatives were worse.
That mediocrity comes at a steep cost. Trials enroll far more patients than would theoretically be necessary if researchers could reliably identify responders in advance. Pharmaceutical companies compensate for imprecise targeting by throwing enrollment numbers—and budgets—skyward. The inefficiency compounds with every failed readout.
Meanwhile, regulators have started paying closer attention to artificial intelligence in drug development. The FDA logged over 500 submissions with AI components between 2016 and 2023. In early 2026, the agency's drug evaluation arm published guiding principles for AI practice—though cynics note that principles don't equal approvals. The European Medicines Agency released its own reflection paper on AI in September 2024, kicking off a multi-year implementation process that runs through 2028.
Translation: the regulatory scaffolding is going up, but the framework remains under construction.
Enter the Foundation Models
The technical underpinning here mirrors what happened in language AI, just applied to biology instead of text. Foundation models—large neural networks pretrained on vast datasets—can be fine-tuned for specific tasks using relatively modest amounts of new data. It's transfer learning, the same concept that powers chatbots and translation tools, retooled for cellular prediction.
Atlas Discovery, a three-person startup out of Y Combinator's summer 2026 cohort, built what they describe as "foundation models of patient drug response." The San Francisco team demonstrated their system on UNIFI, a Phase 3 trial studying ustekinumab (marketed as Stelara) in ulcerative colitis patients. Working solely from baseline colon biopsy data, their model predicted week-8 clinical response with an AUROC of 0.76 across 358 patients.
That number—0.76—represents a meaningful improvement over existing biomarkers. More to the point, Atlas calculated that prospective deployment of such a predictor could have achieved equivalent statistical power using approximately 91 patients per trial arm instead of the 320 actually enrolled. A 3.5-fold reduction in sample size translates directly into millions of dollars saved and, potentially, months shaved off timelines.
The company has backing from Y Combinator, Pear VC, and Glasswing Ventures, though specific funding figures haven't been made public. Their technical blog post includes a formula linking biomarker AUROC to enrollment requirements—each 0.01 improvement in AUROC yields roughly 5% fewer patients needed. Tidy math, if it scales.
"Foundation-model latent space plus a small supervised head solves sample size limitations of trial-scale supervised modeling," the Atlas team wrote. Perhaps. Or perhaps this represents another case of retrospective analysis looking cleaner than prospective reality will allow.
The Broader Ecosystem

Atlas isn't working in isolation. In July 2026, researchers published results in Nature Medicine describing COMPASS, a cross-cancer immunotherapy response predictor trained on pretreatment tumor RNA sequencing from more than 1,100 patients spanning 16 cohorts. Validated against a held-out Phase 2 urothelial cancer trial, it outperformed both tumor mutational burden and PD-L1 testing. Patients flagged as likely responders showed meaningfully better outcomes than those the model predicted would fail treatment.
Tempus AI, which went public in 2024, presented research at the American Society of Clinical Oncology meeting in June 2026 demonstrating an RNA-based survival model for trastuzumab deruxtecan in HER2-expressing breast cancer. The company's pitch revolves around multimodal foundation models trained on aggregated data from thousands of trials and millions of patients—a scale advantage that smaller startups can't easily replicate.
Then there's the infrastructure layer, companies building the plumbing rather than the predictions themselves. Unlearn.AI secured a European Medicines Agency qualification opinion back in September 2022 for its digital-twin methodology, which uses prognostic covariate adjustment to boost statistical power and shrink enrollment in randomized trials. Medidata markets a Synthetic Control Arm product that draws on historical trial data and real-world evidence across what it claims are 38,000 trials and 12 million patients. ConcertAI offers external control arms positioned explicitly for regulatory use.
Even mechanistic approaches are gaining traction. VeriSIM Life announced a collaboration with the FDA's National Center for Toxicological Research in June 2026, applying its hybrid mechanistic-AI platform to challenges like predicting drug-induced liver injury. The proliferation of players and approaches suggests genuine momentum, though separating signal from promotional noise remains challenging.
Regulatory Limbo
Regulators, for their part, are moving deliberately—some might say glacially. The FDA's Model-Informed Drug Development program continues accepting applications focused on drug-trial-disease modeling and quantitative systems pharmacology. In June 2026, the International Council for Harmonisation published ICH M15, establishing harmonized principles for model-informed development across major markets.
Draft guidance released by the FDA in 2025 addresses AI in regulatory decision-making. The agency stood up a CDER AI Council in 2024 to coordinate strategy across therapeutic centers. Yet there remains a substantial gap between what these tools can demonstrate in retrospective analyses and what regulators will accept for pivotal trial design changes.
The European landscape adds layers of complexity. The EU AI Act entered force in August 2024, with phased implementation running through 2028 covering everything from prohibited AI applications to high-risk systems. The EMA's March 2026 consultation on virtual control groups—initially scoped to preclinical drug safety—signals tentative openness to model-based evidence. But regulatory qualification of any specific AI biomarker remains fundamentally a case-by-case evaluation.
In other words: positive signals, no clear path.
The Market Story

Industry analysts are, predictably, bullish. Grand View Research pegged the AI-in-drug-discovery market at $2.9 billion in 2026, projecting growth to $13.8 billion by 2033—a compound annual growth rate approaching 25%. IQVIA's March 2026 global R&D trends report noted "early evidence of stronger success rates for AI-enabled programs among emerging biopharma," though the sample size remains small enough to caveat heavily.
A Deloitte survey from mid-2026 found that 61% of life sciences executives are prioritizing partnerships specifically to access AI capabilities. McKinsey scenario modeling from 2024-2025 suggested potential cost reductions approaching 50% in certain trial operations, though those figures remain directional projections rather than demonstrated outcomes. Analysts love a growth story, and this one has all the right ingredients.
The underlying science continues advancing. Foundation models trained on 30 million to 100 million cells—scGPT, scFoundation, CellFM—now enable transfer learning for drug response and cellular perturbation prediction. A Nature Communications paper from June 2026 described CellFM, pretrained on 100 million human cells, adding another tool to the expanding repertoire of biological state embeddings.
All of which sounds impressive. The question is whether impressive translates to useful.
What Could Actually Change
The central question isn't whether AI can predict drug response in controlled analyses—the AUROC numbers demonstrate it can, at least retrospectively. The real question is how quickly these tools migrate from academic publications to registered trials, and whether they can ultimately reach routine clinical practice.
Several variables will determine the pace. First, regulators need confidence that AI-selected populations don't create biased efficacy estimates or inadvertently exclude patients who might benefit. The FDA's draft guidance on diversity action plans, issued in June 2024, may complicate how enrichment strategies interact with inclusion mandates. Narrowing trial populations to predicted responders could conflict with goals around demographic representation.
Second, the economics need to work in practice, not just on paper. Atlas Discovery's claim of 3.5-fold enrollment reduction would indeed save tens of millions per trial. But if developing and validating the biomarker itself costs $20 million and requires multiple years, the value proposition shifts considerably. Tempus and Medidata are betting that platform approaches with reusable infrastructure will solve this equation. Maybe. The business model remains unproven at meaningful scale.
Third, technical performance needs to hold up prospectively and generalize across populations. Most published results—including Atlas's UNIFI analysis—are retrospective validations. A model trained primarily on European trial participants might perform differently in Asian or African populations, a concern regulators will scrutinize intensely. Real-world deployment will surface edge cases that retrospective analysis misses.
Tempering the Narrative

For compounds that genuinely failed Phase 2 or Phase 3 due to poor patient selection rather than fundamental lack of efficacy, AI-driven stratification could prove transformative. Rescue programs might become viable where they're currently written off as lost causes. That nine-out-of-ten failure rate might tick downward, even if only marginally.
But expectations need tempering. AI won't fix a flawed mechanism of action or predict safety signals invisible in preclinical models. It won't solve recruitment challenges in rare diseases where every eligible patient already matters. And it certainly won't eliminate the need for rigorous Phase 3 trials—just potentially make them smaller and more precisely targeted.
That's not revolutionary, exactly. But in an industry where marginal improvements compound into billions of dollars and years of development time, marginal might be enough. The pharmaceutical industry has survived on smaller victories than that.
Whether this particular wave of AI enthusiasm delivers on its promises or joins the long history of overhyped healthcare technologies remains to be seen. The foundation models are real, the retrospective validations are encouraging, and the regulatory groundwork is being laid. What happens next depends less on the algorithms themselves and more on whether the humans making decisions—regulators, pharmaceutical executives, clinicians—trust them enough to act.
For now, it's still mostly promise. But the promise, at least, looks better than a coin flip.
