Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

ai3Bio raises $48M to reset immune systems for remission

ai3Bio raises $48M to reset immune systems for remission
BiotechAutoimmune Disease+3
Healthtech & Biotech iconHealthtech & BiotechOctober 3, 2026

Halmos Labs automates biotech R&D with AI-driven design

Halmos Labs automates biotech R&D with AI-driven design
YcDrug Discovery+3
Healthtech & Biotech iconHealthtech & BiotechMarch 6, 2026

How AI Is Solving Biopharma's Hidden Bottleneck in Drug Development

How AI Is Solving Biopharma's Hidden Bottleneck in Drug Development
YcDrug Development+3
Climate / Social Tech iconClimate / Social TechMarch 6, 2026

Self-Charging Drones Land on Power Lines for Infinite Grid Monitoring

Self-Charging Drones Land on Power Lines for Infinite Grid Monitoring
YcPower Infrastructure+1

Founders Mentioned

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS
Healthtech & Biotech iconHealthtech & Biotech
March 6, 2026
YcAiDrug DiscoveryMachine Learning

How AI Foundation Models Are Solving Biology's Missing Data Problem

YC-backed Strand AI and competitors race to predict unmeasured biological data, promising to accelerate drug discovery in a market projected to reach $13.8B by 2033.

How AI Foundation Models Are Solving Biology's Missing Data Problem

Every clinical trial leaves something unmeasured. A researcher might sequence the genomes of hundreds of cancer patients but lack the budget to profile their protein expression. Another might have decades-old tissue slides sitting in a freezer, rich with visual detail but silent on the molecular activity that once coursed through those cells. It's not neglect—it's economics. Measuring everything costs too much.

Now a cluster of startups thinks artificial intelligence can predict what was never measured in the first place. Feed their models the data you have, they say, and they'll conjure the data you don't. Histology images in, protein levels out. Genotypes in, gene expression out. The promise is tantalizing: unlock insights from incomplete datasets without the expense of new lab work.

Strand AI, a company that recently passed through Y Combinator's accelerator program, is among the newest to make this pitch. The market they're chasing isn't small—Grand View Research pegged AI in drug discovery at $2.35 billion in 2025, with projections climbing toward $13.77 billion by 2033. But whether these tools work well enough to reshape how drugs get developed? That's still an open question, and one the industry is spending billions to answer.

Money Talks, But Does Biology Listen?

The pharmaceutical industry has never been shy about throwing capital at promising technologies. By early 2024, McKinsey was estimating that generative AI could create $60 billion to $110 billion in annual value for drug companies. Deloitte, perhaps more cautiously, put the near-term opportunity at $5 billion to $7 billion. Either way, executives started writing checks. McKinsey's surveys suggested the share of companies spending at least $5 million on generative AI would leap from 20 percent in 2024 to 32 percent the following year.

The deal-making followed predictably. Alphabet's Isomorphic Labs inked discovery agreements with Eli Lilly and Novartis worth up to $3 billion combined in January 2024, then pulled in $600 million in outside funding a year later. Recursion Pharmaceuticals absorbed Exscientia in late 2024, merging two high-profile players in AI-driven discovery. Lilly kept signing deals throughout 2025—research collaborations with Insilico Medicine, another with Creyon Bio.

But these platform companies, aiming to discover and optimize entire drug candidates, represent just one slice of the landscape. A different tier of startups has focused on something narrower and perhaps more tractable: predicting biological measurements that nobody actually took. Strand AI sits in this category. Founded by Yue Dai and Oded Falik, who both worked previously at Enable Medicine, the company builds what it calls "data generation and imputation models for biology." The mechanics are straightforward. You hand over whatever sparse data you collected—maybe histology slides and patient genotypes—and the model predicts the proteomics or transcriptomics you couldn't afford to measure.

It's worth noting that Strand isn't building this in a vacuum. Tempus AI, which went public in June 2024, has assembled a clinical and molecular database spanning more than 350 petabytes and north of 40 million patient records as of mid-2025. Enable Medicine, Strand's founders' former employer, partnered with Akoya Biosciences in April 2025 to launch what they described as the largest commercially available single-cell spatial proteomics atlas. These aren't just model-training exercises. They're infrastructure plays, assembling the datasets that make the models possible.

The Technical Bet

The underlying challenge is conceptually simple. Biological systems spew data across multiple dimensions—DNA sequences, RNA expression, protein levels, metabolite concentrations, tissue structure. Measuring all of it, for every patient, in every trial, costs more than most budgets allow. Clinical trials routinely grapple with missing data in the high teens overall, sometimes topping 20 percent for patient-reported outcomes, according to a 2024 review in BMC Psychiatry. Academic researchers often have histology slides archived for decades with no corresponding molecular profiles.

Foundation models trained on paired datasets—where multiple measurements were taken—offer a workaround. Show the model enough examples where both histology and gene expression exist, and it learns to predict one from the other. The academic literature has been inching this direction for years. Owkin published HE2RNA in Nature Communications back in 2020, showing that RNA expression could be inferred from tumor histology images. Google's Enformer model, which appeared in Nature Methods in 2021, predicted gene expression from genomic sequences. More recent papers have tackled spatial proteomics, including a January 2026 study in Nature Medicine demonstrating virtual proteomics profiles generated from routine slides.

Strand AI says its first model predicts spatial proteomics from histology, claiming to beat existing benchmarks. Early in 2026, the company released a dataset using something called VariantFormer to generate imputed DNA-to-RNA expression data for more than 500 individuals from the 1000 Genomes Project. On the infrastructure side—because these things matter when you're actually training models—Strand disclosed it's running multimodal biology foundation models on NVIDIA B200 GPUs. The team has even posted publicly about debugging performance snags, like a fivefold throughput regression traced to misconfigured Transparent Huge Pages. Not the stuff that makes headlines, but the kind of detail that determines whether a model trains in days or weeks.

The economic logic is clear enough. If you can reliably impute missing measurements, you unlock analyses that would otherwise require expensive new lab work or bigger patient cohorts. A trial built around genomic data could be retrospectively mined for protein-level insights. A biomarker hunt constrained by sample scarcity could be virtually expanded.

Regulatory momentum is building, too, though not without complications. The FDA released draft guidance in January 2025 outlining a credibility framework for AI models used in regulatory submissions, noting it had received more than 500 AI-related submissions since 2016. The EU AI Act, finalized in July 2024, reaches general applicability in August 2026, with stricter provisions for high-risk medical device applications following by mid-2027. The European Medicines Agency finalized a reflection paper on AI in drug development in September 2024, emphasizing transparency and validation.

From Benchmarks to Bedside

Digital illustration for article section "From Benchmarks to Bedside" in "How AI Foundation Models Are Solving Biology's Missing Data Problem" - A highly detailed, cinematic close-up of a singular, isolated biological strand or molecular structu...

Academic labs laid the groundwork, but turning research-grade models into production systems involves solving different problems. Strand AI's pitch is narrower than platform companies like Recursion or Isomorphic, which aim to shepherd drug candidates from concept to clinic. Strand focuses specifically on the imputation layer—the missing data problem. The founders frame it in terms drug developers recognize: sparse, fragmented patient data that limits what you can learn. Their stated approach emphasizes validation through downstream clinical utility—patient stratification, outcome prediction—rather than leaning solely on reconstruction accuracy scores.

That focus on validation matters more than it might seem. The academic literature from 2025 and 2026 has grown noticeably cautious. Nature Methods published a year-in-review piece in January 2026 flagging the need for robust validation, pointing out that some foundation models failed to outperform simpler baselines on perturbation prediction tasks. A critical assessment in Genome Biology from early 2025 noted that zero-shot performance of single-cell foundation models often disappointed and that benchmarking plus interpretability remained serious pain points. Microsoft Research and others have sounded similar warnings.

The imputation techniques themselves have advanced quickly. Papers from 2024 through 2026 describe methods like TransImpute (with uncertainty estimation), SpaIM (for spatial transcriptomics), and sCellST (using style-transfer frameworks). Large benchmarking efforts such as HESCAPE, presented at an ICCV workshop in 2025, are comparing cross-modal histology-to-gene-expression models. The Chan Zuckerberg Initiative Biohub released a VariantFormer model card in October 2025 for personalized expression prediction integrating whole-genome sequencing and RNA data. Earlier work on single-cell multi-omics, including totalVI and scVAEIT, established baselines for imputing protein measurements from RNA or the reverse.

But as Strand AI discovered with its throughput regression issue, moving from published methods to production systems means wrestling with engineering details that never make it into journal articles. The company is working on FP8 precision and compiler optimizations—unglamorous work that nonetheless determines whether your pipeline is practical or prohibitively slow.

Competitors are pursuing different angles. Owkin, which published the original HE2RNA work, secured a $180 million equity investment and collaboration with Sanofi in November 2021 and has kept building pathology-genomics tools. Enable Medicine, Strand's founders' former shop, announced the spatial proteomics atlas with Akoya in April 2025. Tempus, with its sprawling multimodal database, positions itself as both data provider and AI-enabled diagnostics company.

The consolidation trend is hard to miss. Recursion's acquisition of Exscientia in late 2024 created a combined entity with broader platform ambitions. Isomorphic's deals with Lilly and Novartis, followed by that $600 million funding round, signal that large-scale investment in AI discovery infrastructure is still flowing, perhaps more than some expected. These companies aren't focused exclusively on imputation, but they compete for overlapping enterprise budgets and face similar technical hurdles around data quality, model generalization, and validation.

What Comes Next

The next 18 to 24 months should clarify whether biological data imputation becomes standard practice in drug development or remains a research curiosity. Several dynamics seem likely to unfold.

Validation standards will almost certainly tighten. The academic literature has already pivoted in that direction, with multiple papers in 2025 and 2026 stressing the need for uncertainty quantification, external validation cohorts, and task-relevant endpoints rather than generic accuracy metrics. The FDA's draft guidance and the EU AI Act both push toward greater documentation and credibility assessment for AI models used in regulatory contexts. Companies that can show their imputed data actually improves patient stratification or predicts clinical outcomes will have a smoother path to adoption than those relying on reconstruction scores alone.

The gap between research-use tools and clinically validated systems will likely widen. Strand AI and similar companies initially target preclinical research and retrospective trial analysis, where regulatory burdens are lighter. But any imputation model feeding into a clinical decision support system or diagnostic will face much heavier scrutiny under the EU AI Act and potential FDA oversight. The boundary isn't always crisp, and navigating it will require care.

Data quality and provenance will matter as much as model architecture—maybe more. Training on clinical data implicates HIPAA compliance and patient consent frameworks. The EU AI Act and EMA guidelines raise documentation and audit expectations for any AI tied to regulatory submissions or medical devices. Companies with licensed, high-quality multimodal datasets and transparent data governance may have an advantage over those relying on publicly available but less controlled sources. Enable Medicine's partnership with Akoya to build a commercial spatial proteomics atlas exemplifies that strategy. So does Tempus's massive clinical database.

Expect hybrid approaches. The 2025 and 2026 literature on single-cell foundation models and spatial transcriptomics imputation consistently shows that zero-shot generalization across sites, assays, and populations remains challenging. Fine-tuning on cohort-specific data often beats off-the-shelf models. Some researchers are exploring reasoning-oriented biological foundation models, like BioReason, to support mechanistic hypothesis generation rather than pure pattern matching. But independent assessments warn that performance and interpretability gaps persist. The practical outcome will probably be pipelines combining foundation models with statistical methods, causal inference frameworks, and prospective wet-lab validation rather than AI alone.

The market will consolidate further, though niches will endure. Large platform companies pursuing end-to-end drug discovery have different needs than academic labs running retrospective biomarker studies or small biotechs trying to wring more insight from limited trial data. Strand AI's positioning as a focused imputation layer suggests a bet that not everyone wants—or needs—a full-stack AI drug discovery platform. Whether that bet pays off hinges on execution, validation, and whether imputation models deliver measurable value in actual development programs.

The Unforgiving Test

The broader opportunity is real enough. A market projected to reach $13.8 billion by 2033 represents substantial growth, and the underlying problem—sparse, incomplete biological data—isn't disappearing. But biology is unforgiving in ways that, say, language modeling is not. A protein prediction that's 80 percent accurate might be useless if the 20 percent error occurs in functionally critical regions. An imputed biomarker that looks promising in silico but flops in a prospective cohort wastes time and capital, two things drug development can't afford to squander.

Foundation models have proven their worth in other domains—text, images, code. Whether they can fill biology's missing data gaps at the scale and reliability drug development demands remains, for now, an empirical question. Strand AI, Owkin, Tempus, and the rest are racing to answer it. The companies that succeed won't just build better models. They'll build the validation infrastructure, the data partnerships, and the regulatory pathways that turn predictions into evidence.

And evidence, in the end, is what biology demands.

More stories

  • ai3Bio raises $48M to reset immune systems for remission
  • Halmos Labs automates biotech R&D with AI-driven design
  • How AI Is Solving Biopharma's Hidden Bottleneck in Drug Development
  • Self-Charging Drones Land on Power Lines for Infinite Grid Monitoring
  • DiligenceSquared Raises $5M Seed for AI-Powered PE Due Diligence
  • Digital Therapeutics 2026: Every FDA-Cleared Prescription App
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.