Georgia Witchel pulls up a diagram on her laptop—layers of boxes and arrows mapping how data moves through a typical Phase 3 clinical trial. Or rather, doesn't move. Electronic data capture systems. Lab instruments. Imaging platforms. Genomics pipelines. Each box isolated, each arrow a manual export, a CSV file emailed between teams, months of reconciliation before anyone can ask whether the drug actually works.
"This," she says, "is why digital twins in healthcare are mostly theoretical."
Witchel would know. Her previous venture, Louiza Labs, built physics engines for surgical simulations—digital twins for operating rooms. She watched promising models stall not because the math was wrong, but because feeding them clean, standardized data was nearly impossible. Now, as founder of Mantis Biotechnology, a three-person startup fresh from Y Combinator's Winter 2026 batch, she's betting that before healthcare can fulfill the promise of virtual hearts and AI-powered patient replicas, it needs something far less sexy: plumbing.
The stakes are considerable, if you believe the projections. Mordor Intelligence forecasts the healthcare digital twin market growing from $2.81 billion in 2025 to $14.12 billion by 2031. Fortune Business Insights is more optimistic—or reckless, depending on your perspective—claiming a leap to $516.3 billion by 2034. Someone's math is off, but the directional signal is clear. Money is flowing toward computational models that can predict device failures, shrink trial timelines, and optimize hospitals in real time.
Except there's a catch, one that industry insiders acknowledge quietly and market reports tend to gloss over. The infrastructure to actually build, validate, and deploy these models at scale? It barely exists.
When Definitions Matter More Than You'd Think
Start with a basic question: what is a digital twin, exactly?
A 2025 scoping review in npj Digital Medicine examined studies claiming to use "human digital twins" and found only 12% met the National Academies of Sciences, Engineering, and Medicine's criteria. The rest were digital models, dashboards, or visualizations—useful, perhaps, but not twins. Not personalized, not dynamically updated, not capable of predictive decision support.
The NASEM published a consensus report in 2024 trying to impose order on the definitional chaos. It emphasized verification, validation, and uncertainty quantification—the unglamorous technical work that determines whether a model can be trusted. The report flagged persistent gaps: how digital twins integrate data across biological scales, how they maintain feedback loops between virtual and physical worlds, how they achieve reliability in clinical settings where patient safety hangs in the balance.
A separate paper in JMIR Medical Informatics went further in 2025, suggesting the industry abandon the notion of a single "patient twin" entirely. Instead, it proposed modular architectures—separate dashboards, virtual cohorts, prediction engines—that better align with regulatory realities and technical constraints.
What actually works today? Operational twins in hospitals have demonstrated tangible returns. Siemens Healthineers deployed a digital twin for Mater Private's radiology department in Ireland. GE HealthCare's Command Center uses virtual hospital models at Providence Swedish to optimize capacity planning and staffing. These aren't predicting individual patient outcomes—the scope is narrower, the data cleaner, the simulation more bounded.
On the drug development side, physiologically-based pharmacokinetic modeling has won regulatory acceptance. Certara's Simcyp platform became the first software to receive an EMA qualification opinion in 2025, with over 115 drugs now using it for label claims. Simulations Plus has multiple FDA grants validating its workflows. These tools model drug absorption and metabolism at scale. But they're organ-specific, system-specific. Not whole-body twins.
The gap between what exists and what's promised is where things get expensive.
The $15 Million Question
According to research Mantis Biotechnology cites, 80% of clinical trials face delays due to data inaccuracies. The average loss from data quality issues alone? Fifteen million dollars per trial.
When Tufts and Medidata estimate direct trial costs at $40,000 to $56,000 per day—and lost sales opportunity at $500,000 to $800,000 daily—those delays aren't merely inconvenient. They're existential. A three-month slip can mean competitors reach market first, or that patent clocks run down before launch, or that investor patience evaporates.
The friction sits upstream, in data collection and harmonization. Clinical trials pull information from a dozen sources: electronic data capture systems, clinical trial management platforms, lab instruments, imaging systems, genomics pipelines, electronic health records. Each has its own schema, its own terminology, its own quality control processes. Sometimes its own interpretation of what a "patient ID" means.
Harmonizing this mess into something a sophisticated model can ingest—that's where the $15 million disappears. Slowly, laboriously, into spreadsheets and manual checks and emails asking whether "glucose_level" in one dataset corresponds to "blood_glucose_mg_dl" in another.
Witchel frames it bluntly: the infrastructure layer for biomedical data doesn't exist the way it does in tech. Databricks provides unified data platforms for enterprises. Snowflake does the same for cloud analytics. Healthcare? EDC silos. HIPAA-compliant file transfers. FTP servers, if you can believe it.
Mantis positions itself as "Databricks for biomedical and clinical data"—domain-aware pipelines that encode biological meaning into reusable datasets. Anatomy. Physiology. Clinical context. Not just columns and rows, but data that knows what it represents.
The pitch resonates because everyone in the industry feels the pain. Oracle surveyed trial sponsors in 2024 and found data governance failures driving protocol amendments and enrollment delays across the board. The Cassava Sciences saga offers a cautionary tale: the company's Alzheimer's drug simufilam failed Phase 3 trials in November 2024 amid longstanding data integrity scrutiny and an SEC settlement. Traceability and provenance didn't just determine trial success—they determined corporate survival.
Standards exist, theoretically. HL7 FHIR released its R5 version in March 2023. The OMOP Common Data Model offers theoretical interoperability. The U.S. government's Trusted Exchange Framework and Common Agreement went operational in late 2023, expanding qualified health information networks throughout 2024.
Adoption, however, remains uneven. And pharma companies can't wait for public infrastructure to mature. They need something now.
Regulators Open the Door

What's changed the calculus—perhaps more than sponsors realize—is regulatory willingness to accept computational models as primary evidence.
The FDA Modernization Act 2.0, signed in December 2022, explicitly allows nonclinical tests, including in silico models, to replace animal studies where scientifically appropriate. A sentence in legislation, but one with profound implications for how drug development happens. The FDA's Model-Informed Drug Development Paired Meeting Program, formalized under PDUFA VII covering 2023 through 2027, now provides a pathway for sponsors to get feedback on modeling approaches before committing to pivotal trials.
On the device side, FDA guidance dating back to 2016 encourages computational modeling in premarket submissions. ASME's V&V40 standard, first published in 2018, provides a risk-based credibility framework—essentially, how to prove your simulation isn't garbage. A 2024 symposium co-hosted by the Medical Device Innovation Consortium found 65.7% of surveyed experts expect growing use of computational modeling in regulatory submissions. Not hope. Expect.
The European Medicines Agency has moved in parallel, sometimes faster. Its 2019 PBPK reporting guideline, adopted by Australia's TGA in 2020 and updated in 2025, standardized how physiologically-based models are documented for drug applications. In 2025, EMA made headlines by accepting AI-assisted pathology evidence—supervised by human experts—under its first-ever qualification for such methodology.
Unlearn, a California company building digital twins for clinical trial design, achieved EMA qualification for its PROCOVA covariate adjustment method and secured positive FDA feedback. CEO Charles Fisher has publicly stated that digital twins can reduce control-arm enrollment by 25% to 50% and shorten recruitment by four to five months. The company's collaboration with AbbVie in Alzheimer's trials suggests sponsors are willing to bet on the approach, provided regulators signal acceptance.
Dassault Systèmes' Living Heart Project, developed with FDA collaboration, entered a 2025 beta phase for its next-generation AI-powered virtual heart. Patient-specific and population-level configurations. Simulated device trials that reduce animal and human subject requirements. The company frames this as aligning with the 21st Century Cures Act's push for digital evidence.
Translation: the door is open. The question is whether the industry can walk through it.
Infrastructure's Moment
Mantis Biotechnology's YC launch crystallizes a thesis that's been percolating for months among investors and founders who watch healthcare from the outside. Before the industry can support sophisticated twins, it needs a foundational data layer. Not another simulation engine. Not more AI-powered replicas. The pipes. Integrations with EDC systems, lab instruments, omics platforms. Provenance tracking for every data point. HIPAA-compliant workflows that pass 21 CFR Part 11 scrutiny.
The company's Wellfound hiring page describes its positioning as "Palantir for biomedical"—a reference to how Palantir's Foundry platform became the backbone of the NHS Federated Data Platform in the UK. That's a £330 million, multi-year rollout that's faced public debate over governance but signals the scale of demand for unified clinical data. Mantis claims its platform enables "human-in-computer models" validated against real-world outcomes, with use cases spanning professional sports to medical device development.
It's a narrative that mirrors Databricks' trajectory in enterprise data, Snowflake's in cloud analytics. Those companies didn't invent data lakes or SQL. They made them usable at scale, with governance and tooling enterprises could trust. Mantis is betting healthcare digital twins need the same.
The incumbent landscape features established players, but with narrower scopes. Palantir's NHS work focuses on operational dashboards. Databricks has inked healthcare partnerships—Health Catalyst announced Delta Sharing integration; Novo Nordisk presented a multi-agent clinical data framework at the Data + AI Summit 2025—but isn't healthcare-native. Doesn't know the ontologies, the regulatory context, the workflow idiosyncrasies.
Unlearn, Medidata's Acorn AI, and Phesi target synthetic control arms. Adjacent, but not infrastructure. Certara and Simulations Plus own PBPK modeling but don't claim to unify trial data end-to-end.
The question facing Mantis and similar infrastructure startups? Whether "domain-aware" data platforms can command venture-scale returns or get absorbed into larger clinical trial software suites—Oracle, Medidata, Veeva—that already own sponsor relationships. The launch post's reference to the simufilam case suggests the sales pitch hinges on risk mitigation as much as speed. Data integrity as an existential issue, not just an operational one.
The Long Game

A systematic review in npj Digital Medicine from June 2024 found digital twins achieved 80% effectiveness across 45 measured outcomes in diverse conditions. Impressive, until you read the fine print: adoption remains early and fragmented. The National Academies' 2024 report envisions digital twins expanding into precision health with dynamic updates and multi-scale data integration, contingent on—there's always a contingency—standardized definitions, clinical validation, and regulatory pathways maturing in tandem.
Hospital operational twins will likely scale faster than patient-level replicas. Workforce shortages provide clear ROI. GE HealthCare and Siemens already have deployments showing reduced length of stay and improved access. Clinical trial applications—synthetic control arms, covariate adjustment, in-silico trial augmentation—are gaining traction as regulators formalize acceptance criteria and early adopters publish results.
The harder question is whether whole-body, continuously updated patient twins will move from academic prototypes to clinical practice within this decade. Twin Health has published outcomes in Type 2 diabetes: a 1.8% HbA1c reduction over one year, with medication deprescribing. Promising for metabolic conditions where continuous glucose monitors and wearables provide the data feed.
Cancer? Neurodegenerative disease? Rare conditions? They lack that real-time sensor infrastructure. Maybe always will.
For infrastructure startups, the window may be narrow. If Databricks, Snowflake, or Palantir decide healthcare data unification is strategic, they have capital and engineering depth to move fast. If clinical trial software incumbents integrate similar traceability and domain modeling into existing platforms, the "Databricks for biomedical data" positioning loses differentiation.
The advantage—perhaps the only sustainable one—lies in being healthcare-native. Understanding the ontologies, the regulatory context, the workflow idiosyncrasies that generic data platforms miss. Whether that's enough to build a venture-scale business remains an open question.
The broader trend, though, is undeniable. Computational models are shifting from supplements to primary regulatory evidence in defined contexts. PBPK. Device simulations. AI-augmented trial design. All crossing the credibility threshold, slowly but unmistakably.
What the industry lacks isn't vision or regulatory appetite. It's the unsexy, essential foundation—unified data that's traceable, reproducible, ready to feed models regulators will trust.
That's the race Mantis and others are running. And the $14 billion market projection? It assumes someone figures it out soon.
