Pharmaceutical companies burn through $2.67 billion, on average, to shepherd a single drug from lab bench to pharmacy shelf. Nine out of ten still fail.
That brutal arithmetic hasn't changed much in decades. But something else has: regulators are starting to accept that artificial intelligence might offer a way out.
In March, the FDA did something remarkable. It allowed Certara, a drug development software company, to use its Simcyp physiologically-based pharmacokinetic model in place of ten human clinical studies for a leukemia therapy called asciminib. Instead of conducting those specific trials, the agency accepted results from Simcyp—a physiologically-based computer model that simulates how drugs move through the human body. No actual patients required for that particular set of tests.
Not quite the same as a foundation model trained on patient data, perhaps. But it signals a shift.
Consider Atlas Discovery, a three-person startup from Y Combinator's Summer 2026 batch. The company claims its AI model can retrospectively predict which ulcerative colitis patients would respond to treatment in a phase 3 trial—with enough accuracy to reduce enrollment by 458 people. That's not a marginal improvement. That's the difference between a $50 million trial and something considerably cheaper.
Or take Tempus, the oncology diagnostics company that went public and announced in May that its multimodal foundation model—trained on more than 45 million patient journeys—achieved a C-index of 0.802 for survival prediction in cancer patients. The company positions this as a tool for designing smarter trials and stratifying risk, though whether it will actually reduce trial costs at scale remains an open question.
These aren't academic exercises anymore. They're regulatory submissions. And the fact that agencies are even considering them marks an inflection point that founders, investors, and pharma executives should be watching closely.
The Economics Are Still Terrible
Clinical trial success rates remain dismally low—roughly 14.5 percent across recent cohorts, according to a November 2025 study in Nature Communications. That figure has barely moved despite billions in R&D spending. Deloitte reported in 2025 that industry returns on pharmaceutical innovation had crept up to 7.0 percent, but strip out GLP-1 obesity drugs and most pipelines still face punishing economics.
Then there's the patent cliff. Between now and 2030, somewhere between $230 billion and $300 billion in sales are at risk as blockbuster drug patents expire. Meanwhile, the clinical trials market itself keeps expanding—Grand View Research pegs it growing from $89 billion in 2025 to $158.4 billion by 2033, a CAGR of 7.7 percent. That growth, however, reflects volume and complexity, not efficiency. Pharma companies are running more trials, at higher costs, with failure rates that haven't budged.
AI adoption in clinical trials is accelerating, though market sizing estimates vary wildly—Fortune Business Insights projects the sector growing from $5.5 billion in 2026 to $77.3 billion by 2034 (a CAGR of 39.14 percent), while Mordor Intelligence forecasts something more conservative: $2.68 billion this year growing to $8.24 billion by 2031. The discrepancy reflects definitional chaos. What counts as "AI in clinical trials" spans everything from patient recruitment software to foundation models predicting individual drug response.
What's genuinely new is regulatory receptiveness. The FDA logged more than 500 drug and biologic submissions with AI components between 2016 and 2023. The agency proposed a framework for evaluating AI model credibility in January 2025 and released draft guidance on alternatives to animal testing this past March. Europe moved faster: the ICH M15 guideline on model-informed drug development reached finalization in February, harmonizing standards across the FDA, EMA, and Japan's PMDA.
Why Now?
Three forces are converging to make this possible.
First, regulatory frameworks have matured. ICH M15 establishes common principles for pharmacokinetic modeling, physiologically-based simulations, and quantitative systems pharmacology. It doesn't guarantee approval, but it gives sponsors and regulators a shared vocabulary. The FDA's Model-Informed Drug Development Paired Meeting Program, previously a pilot, is now ongoing under PDUFA VII. The EMA's qualification pathways for novel methodologies are issuing letters of support with increasing frequency.
The FDA Modernization Act 2.0, enacted back in December 2022, broadened acceptable nonclinical evidence to include in vitro, in silico, and microphysiological systems—organ-on-chip platforms and computational models can now replace some animal studies. The March guidance made explicit what had been implicit: validated computational models can support regulatory submissions.
Second, data scale has reached a tipping point. Tempus's foundation model was trained on more than 45 million de-identified patient journeys and over 500 petabytes of data, including deep multimodal clinical, genomic, and imaging records for 400,000 patients. That's not a research dataset. That's an industrial corpus. Xaira's X-Atlas perturbation sequencing atlas, released publicly over the past year, provides a shared resource for training models on how cells respond to genetic and pharmacological interventions.
Third, economic pressure is forcing sponsors to take risks. The patent cliff means companies need faster, cheaper paths to market. McKinsey estimated in 2024 that generative AI could unlock $60 billion to $110 billion in annual value across pharma, potentially cutting clinical costs by up to 50 percent and shaving a year or more off trial timelines. Those projections assumed mature deployment; we're just reaching the beginning of that curve now.
The Evidence Is Still Mostly Retrospective

Atlas Discovery represents the newest wave of foundation model companies targeting clinical response prediction. The startup's approach: large-scale pretraining across in vitro and clinical data, then fine-tuning on specific trial cohorts. In a June retrospective analysis of the UNIFI phase 3 trial for ustekinumab in ulcerative colitis, Atlas's model predicted baseline responders from patient biopsies with an AUROC of 0.76—performance the company says would enable equivalent statistical power with 182 patients instead of 640.
That's a backtest, not a prospective validation. But the consistency is notable: Atlas reports AUROC ranges between 0.7 and 0.9 across multiple trials. The company's technical work on discrete diffusion models for single-cell perturbations, presented at several machine learning conferences this year, claims a tenfold improvement in accuracy with 50 times less data. Backed by Y Combinator, Pear, and Glasswing, Atlas is positioning itself for partnerships with biopharma and hospital systems.
Tempus operates at a different scale. The company announced initial foundation model results on May 29, highlighting a C-index of 0.802 for overall survival prediction in oncology patients. Tempus frames its model as a tool for trial design and patient risk stratification, and its TIME trial network uses biomarker predictions derived from the model to accelerate enrollment. The company reported $261.1 million in diagnostics revenue for Q1 2026 and guided to roughly $1.6 billion for the full year. Collaborations with AstraZeneca and Pathos extend the model's reach into precision oncology.
A CURE AI presentation at AACR in July 2025 claimed a pan-cancer immunotherapy response model could reduce trial sizes by more than 50 percent and save over $100 million in average phase 3 costs through refined eligibility criteria. That work remains a conference abstract, not a peer-reviewed publication—but it reflects a pattern. Multiple teams are converging on similar approaches to patient stratification using large-scale clinicogenomic models.
Unlearn.AI has pursued a different path entirely, focusing on digital twin controls. The company received a draft qualification opinion from the EMA back in 2022 for its TwinRCT framework, which uses patient-level disease progression models to create synthetic control arms. The framework has advanced through EMA acceptance pathways, and Unlearn announced a collaboration with Acumen Pharmaceuticals on Alzheimer's trials in July. A JACC Basic to Translational Science overview in September 2024 outlined the EMA qualification status for certain continuous outcomes. Unlearn positions its approach as a way to shrink placebo arms or support external controls in rare diseases.
The most concrete regulatory win, though, came from Certara. Beyond the FDA acceptance for asciminib, the company released Simcyp v25 in March with expanded capabilities for regulatory submissions. Simulations Plus reported similar success with its DILIsym platform, which supported phase 3 dose selection for fezolinetant through quantitative systems toxicology modeling, published in April.
The Hard Part Comes Next

The 2026 regulatory environment is more permissive than at any prior point. ICH M15 provides a harmonized framework. FDA and EMA guidance documents are proliferating. But acceptance remains context-specific, not automatic. Models must be validated for their intended use, datasets must be representative, and transparency requirements are rigorous.
Atlas Discovery founder Shaamil Karim suggested in June that the evidentiary paradigm will shift "from the RCT to the predictor" as individual-level response models cross fidelity thresholds. Tempus leadership echoed that sentiment in May, arguing that "we should not be relearning biology for every trial; start from trained understanding."
Bold claims. Whether they hold up is another matter.
Prospective validation remains the bottleneck. Retrospective backtests are useful, but regulators want to see models perform on unseen future patients. Unlearn's EMA qualification pathway required years of dialogue. Certara's PBPK acceptance for asciminib built on decades of mechanistic modeling credibility. Foundation models are newer and far less interpretable, which raises the bar.
Generalization is an open question. Predictive performance on immunotherapy response varies significantly across cohorts and tumor types. A January JAMA Network Open study showed deep-learning CT response scores outperformed RECIST in some NSCLC cohorts but not others. A 2026 critique of organ-on-chip platforms emphasized the need for broader validation and cautioned against overgeneralization. The same concerns apply to foundation models: performance on one indication doesn't guarantee transferability.
Dataset bias and fairness present ethical and regulatory challenges. Foundation models trained predominantly on Western populations may perform poorly on underrepresented groups. The FDA's January 2025 AI framework proposal emphasized model credibility and transparency, implicitly acknowledging these risks. Sponsors will need to demonstrate that models work across diverse patient populations—not just the cohorts that dominate training data.
The concentration of R&D returns around GLP-1 programs creates a paradox, too. Deloitte noted in May that industry IRR had risen to 7.0 percent in 2025, but that improvement masks underlying fragility. If obesity drug margins compress or pipelines dry up, the industry faces renewed pressure to improve productivity across therapeutic areas. AI-driven trial prediction is one lever. It's not the only one, and it won't solve structural issues around market access or pricing.
Investment, meanwhile, keeps flowing. Owkin spun out Waiv in March with $33 million for federated learning and biomarker stratification. VeriSIM Life formalized a research collaboration with the FDA's National Center for Toxicological Research in June. Aitia, Quris-AI, and Recursion are all pursuing variations on causal modeling and digital biology. The Evaluate World Preview report in June forecasted pharmaceutical sales exceeding $2 trillion by 2032, driven by obesity and oncology, with China-origin licensing deals surging in Q1. That growth creates both opportunity and competition for AI-native companies.
What to Watch

Founders and sponsors should track regulatory guidance updates from FDA and EMA, particularly around model validation standards. Watch for prospective validation results from companies like Tempus and Unlearn. Look for peer-reviewed publications that move beyond backtests to real-world deployment.
The pilot phase is over. The question now is how fast validated models can displace or augment traditional trial designs—and whether that displacement creates genuine productivity gains or just shifts costs and risks around.
The pharmaceutical industry has spent decades trying to improve clinical trial efficiency through incremental changes. Foundation models and digital twins aren't incremental. They represent a fundamentally different approach to evidence generation.
Whether that approach delivers on its billion-dollar promises will depend less on model performance and more on the messy, slow work of regulatory acceptance, prospective validation, and industry adoption. That work is happening now, across boardrooms and FDA meetings and biopharma collaborations.
The race isn't to build the best model. It's to prove it works when it counts.
