Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

Rhem Labs launches AI robot for aging-in-place monitoring

Rhem Labs launches AI robot for aging-in-place monitoring
YcSenior Care+3
Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

ai3Bio raises $48M to reset immune systems for remission

ai3Bio raises $48M to reset immune systems for remission
BiotechAutoimmune Disease+3
SaaS iconSaaSFebruary 14, 2026

MiniMax's M2.5 Hits 80% on SWE-bench, Matches Frontier AI at 1/10 Cost

MiniMax's M2.5 Hits 80% on SWE-bench, Matches Frontier AI at 1/10 Cost
Ai BenchmarkingOpen Source+2
Healthtech & Biotech iconHealthtech & BiotechFebruary 14, 2026

How AI Is Solving Drug Development's $18B Characterization Problem

How AI Is Solving Drug Development's $18B Characterization Problem
Drug DiscoveryArtificial Intelligence+3
Healthtech & Biotech iconHealthtech & Biotech
February 14, 2026
Drug DiscoveryArtificial IntelligenceBiotechMachine LearningPrecision Medicine

How AI Foundation Models Are Solving Biology's Missing Data Problem

New multimodal AI models can predict unmeasured biological data from patient datasets, transforming drug discovery and clinical trials. Industry reaches inflection point as startups and Big Pharma race to deploy the technology.

How AI Foundation Models Are Solving Biology's Missing Data Problem

The pathology lab at a mid-size biotech looks organized enough on the surface. Tissue samples from 50 patients line one freezer. Forty pathology slides sit filed in drawers. Genomic sequences from 30 individuals live on a server somewhere. Protein expression data? Maybe 15 samples, if you're lucky.

This isn't negligence. Multi-omic profiling costs a fortune. Samples degrade. Some assays simply weren't standard practice when the trial started five years ago. But those missing modalities—the measurements you wish you had but don't—can torpedo a study before it begins. They're the difference between discovering a viable biomarker and running an underpowered analysis that produces nothing. Between repurposing existing patient data and recruiting all over again.

Here's the thing, though: that sparse-data headache is about to become considerably less painful. A new generation of AI foundation models can now predict unmeasured biological information from whatever fragmented data you actually possess. Feed them a histopathology slide, and they'll generate what your gene expression profile probably looks like. Hand them a patient's DNA sequence, and they'll infer tissue-specific RNA levels. The technology has sprinted from academic proof-of-concept to commercial deployment faster than most anticipated. Maybe faster than anyone's fully ready for.

The implications ripple across drug discovery, clinical development, and precision medicine in ways that are only starting to become clear.

A Cambrian Explosion, With Caveats

The past two years have produced something approaching a Cambrian explosion in biological foundation models. Each one trained on a different slice of life's information architecture.

At the sequence level, models like EvolutionaryScale's ESM3 and various genomic transformers—Geneformer, scGPT—learn patterns from DNA, RNA, and protein sequences. DeepMind's AlphaFold 3, published in Nature last May, pushed structure prediction into multi-molecule complexes: proteins binding nucleic acids, ligands, ions. Antibodies engaging antigens.

Cell and tissue-scale models followed in quick succession. Single-cell foundation models now integrate expression data from over 100 million cells across species, as demonstrated by projects like GeneCompass. Spatial transcriptomics models can reconstruct gene expression maps from standard H&E stains—the kind pathologists have been using for a century. Recursion's phenomics foundation model, PHENOM-1 (available through NVIDIA's BioNeMo platform), generates embeddings from cellular microscopy images. It's trained on the company's 60-petabyte dataset, which sounds almost absurd until you consider the scale of biological information.

Then there's clinical multimodal integration, where things get interesting—and complicated. Companies like Owkin and its spinout Bioptimus raised $35 million in February 2024, then another $41 million this past January, to build what they're calling a "universal multimodal foundation model for biology." Initial focus: pathology. Noetik's OCTO "virtual cell" models combine spatial single-cell expression and protein patterns using masked-token multimodal pretraining. The company signed partnerships with Agenus last June and GSK in January.

The underlying technical approaches vary wildly. Transformers, diffusion models, contrastive learning, conditional transport. But the unifying insight remains constant: biology generates correlated data across modalities. Train a large enough model on paired observations, and it learns to predict one type of measurement from another.

Whether it learns well enough remains an open—and expensive—question.

From Morphology to Molecules

The academic precedent dates back to 2020, when researchers published HE2RNA in Nature Communications. The model inferred bulk RNA-seq profiles from whole-slide images across multiple tumor types. It demonstrated that morphology—what cells look like—encodes enough information to predict thousands of gene expression levels. More recent papers have extended the approach to spatial transcriptomics, using cross-modal mask reconstruction and contextual fusion to generate spatial gene expression maps from routine histology.

CZI Biohub released VariantFormer last November. A 1.2-billion-parameter model that predicts personalized gene-level RNA abundance across tissues from whole-genome sequences. The model trained on large paired WGS and RNA-seq datasets, learning the regulatory logic that translates genetic variation into expression. It's DNA→RNA imputation at scale. The kind of capability that could synthesize missing expression modalities in cohorts where only genomic data exists.

Which brings us to Strand AI—a two-person startup out of Y Combinator's Winter 2026 batch. Founded in 2025, based in San Francisco, the company describes its offering as "multimodal patient data you're missing." Models that generate or impute unmeasured biological modalities to fill gaps in sparse patient datasets.

The use cases read like a catalog of drug development's most persistent annoyances. Completing clinical trial datasets. Predicting proteomic or transcriptomic profiles from routine pathology slides. Augmenting rare-disease cohorts where comprehensive profiling costs more than most academic centers can afford. Enabling biomarker discovery in legacy datasets that lack omics measurements.

According to private tracking data from CB Insights, Strand AI raised $500,000 via convertible note roughly ten days before the database's most recent crawl. Y Combinator listed as an investor. The company's LinkedIn presence emphasizes "modality transformation models for biology" and includes a public demo using CZI Biohub's VariantFormer to generate RNA-seq on 1000 Genomes expansion samples. They claim a 37× faster inference pipeline compared to baseline implementations running on A100 GPUs.

It's early days. Team of two, half a million in the bank, no product announcements beyond demos. But the pitch maps to a problem that keeps biopharma executives up at night. Collecting complete multimodal data on patient cohorts is brutally expensive—a single spatial transcriptomics experiment can run thousands of dollars per sample. Proteomic profiling via mass spectrometry requires specialized sample prep and instrument time. If you can computationally "fill in" those missing modalities with reasonable accuracy, you unlock analyses that were previously impossible without re-recruiting patients and re-running assays.

If.

Big Pharma Isn't Waiting

Digital illustration for article section "Big Pharma Isn't Waiting" in "How AI Foundation Models Are Solving Biology's Missing Data Problem" - Create a conceptual visualization of accelerated pharmaceutical innovation where complex molecular s...

The major pharmaceutical players aren't sitting around waiting for startups to prove the concept works. Isomorphic Labs—the Alphabet spinout applying AlphaFold technology to drug discovery—kicked off 2024 with multi-target collaborations with Eli Lilly and Novartis. By March 2025, the company had reportedly raised $600 million led by Thrive Capital, with participation from both pharma partners. The Financial Times noted that Isomorphic expected to have candidate molecules in clinical trials by the end of 2025. It expanded its Novartis collaboration this past February.

Noetik's OCTO virtual cell models secured partnerships with Agenus last June for AI-enabled predictive biomarkers, and a separate deal with GSK in January. The technical approach involves masked multimodal pretraining on spatial single-cell data, allowing the model to simulate spatial expression and protein patterns. The company's pitch centers on biomarker enrichment and clinical response prediction. Essentially: using synthetic biology to identify patient subgroups before you enroll them.

Xaira Therapeutics launched in April 2024 with more than $1 billion in funding. The company assembled a leadership team and investor roster that reads like a who's who of computational biology and venture capital. Focus: end-to-end AI-enabled R&D, from target identification through clinical development. Insilico Medicine—another AI-native drug discovery company—reported positive Phase 2a topline results in November 2024 for ISM001-055, an AI-designed drug targeting idiopathic pulmonary fibrosis. The molecule progressed from discovery to human data in a timeframe that would have seemed implausible a decade ago.

The economics are starting to justify the hype. McKinsey estimated in January 2024 that generative AI could create $60 billion to $110 billion per year in value across pharmaceutical and medical products R&D, clinical operations, and commercial functions. A Bain report from February 2024 found that 40% of pharma executives budgeted expected savings from generative AI into their 2024 plans. Sixty percent set explicit cost or productivity targets.

Benchling's State of Tech in Biopharma 2024 survey—300 respondents, fielded August through September—showed a sharp adoption gap. Sixty-seven percent of large biopharma reported using AI/ML, compared to just 23% of small biotechs. The top investment priorities over the next three years? R&D data platforms and AI/ML capabilities.

Market-sizing estimates from vendors and analysts peg AI in drug discovery at $2 billion to $3 billion in 2025, growing at a compound annual growth rate of 25% to 31% through the early 2030s. Treat those numbers as directional rather than gospel, but the venture capital flows tell a similar story. Biotech AI funding in 2024 included Bioptimus, Isomorphic's $600 million round, and Chai Discovery's $70 million raise.

Infrastructure Catches Up (Mostly)

None of this would scale without cloud deployment infrastructure willing to handle the computational load. EvolutionaryScale released ESM3—its generative protein language model—with an open 1.4-billion-parameter variant available on AWS SageMaker and HealthOmics last June. The model family extends up to 98 billion parameters. Code and documentation live on GitHub. Generate:Biomedicines published its Chroma diffusion-based protein generator in Nature in 2023 and released code and weights for academic and nonprofit use. Profluent's OpenCRISPR-1, an AI-designed gene editor, appeared as a preprint in April 2024 and published in Nature last July. Free licensing for researchers.

NVIDIA's BioNeMo platform has become the de facto hub for biology foundation model deployment. The platform offers containerized microservices for protein folding (ESMFold), generative chemistry (MolMIM), and docking (DiffDock). 2025 integrations into electronic lab notebooks and LIMS systems from vendors like Sapio and Cadence Orion followed. Recursion's PHENOM-1 embeddings became available through BioNeMo in January 2024, providing external access to the company's phenomics foundation model trained on cellular microscopy images at scale.

The shift to cloud-hosted foundation models and turnkey APIs lowers the barrier for biopharma R&D teams to run proof-of-concept studies. You no longer need a dedicated AI team to fine-tune a protein structure predictor or generate spatial transcriptomics from histology.

That matters, particularly for smaller biotechs that lack the data science bench strength of a Genentech or Novartis. Though it also raises questions about what happens when everyone's using the same black-box tools without fully understanding what's happening under the hood.

The Validation Problem No One Wants to Talk About

Digital illustration for article section "The Validation Problem No One Wants to Talk About" in "How AI Foundation Models Are Solving Biology's Missing Data Problem" - A conceptual digital illustration visualizing the challenges of single-cell foundation models, depic...

The enthusiasm comes with caveats. Substantial ones.

A critical evaluation published in Genome Biology this past April by Microsoft Research found that zero-shot robustness in single-cell foundation models remains mixed at best. Models trained on one dataset or tissue type don't always generalize cleanly to new contexts without fine-tuning. The broader scientific community has also debated AlphaFold 3's closed-code release and limited server access for academics. Nature itself ran an editorial commentary on the implications for reproducibility and scientific openness.

Wet-lab and clinical validation remain the bottleneck. Despite pipeline progress, a Wired feature in 2025 noted the slow emergence of "AI drugs" on the market relative to the sheer number of companies claiming breakthroughs. Consultant estimates of 30% to 50% reductions in preclinical timelines under strong AI adoption sound impressive. They rest on uneven evidence. Insilico's positive Phase 2a data for ISM001-055 represents genuine progress, certainly. But one molecule doesn't constitute proof of systematic acceleration across an entire industry.

Regulatory frameworks are trying to catch up. The FDA released draft guidance on January 6, 2025, proposing a framework to advance the credibility of AI models used to support drug and biological product submissions. It's the first AI-specific guidance for drug development, emphasizing risk-based assessment, documentation, and verification/validation. The European Medicines Agency adopted its Reflection Paper on AI across the medicinal product lifecycle in 2024, aligning with the EU AI Act that entered into force last August. Both agencies encourage early regulatory engagement. Both stress that AI use in discovery carries lower risk than clinical decision tools. But data governance, bias assessment, and auditability remain critical.

For companies generating or imputing patient-level omics data, compliance gets thorny fast. The NIH updated its Genomic Data Sharing policy—effective January 25, 2025—to require NIST SP 800-171 security controls for controlled-access human genomic data. That applies to repositories, tools, and end-user environments. If you're hosting or serving generated paired datasets that include human genomic information, you're now in scope for those security requirements. Universities issued implementation notices throughout 2025 as researchers scrambled to understand the implications for cloud-based analysis pipelines.

Using AI-generated or imputed patient data in regulatory submissions will require clear documentation of data provenance, model training and validation, uncertainty quantification, and potential bias assessments. The FDA's draft guidance and EMA's reflection principles make that explicit. Exploratory biomarker work sits in lower-risk territory than pivotal endpoints. But sponsors will need transparency about which data points are measured and which are predicted.

That distinction matters more than most companies currently acknowledge.

What's Coming, Ready or Not

Digital illustration for article section "What's Coming, Ready or Not" in "How AI Foundation Models Are Solving Biology's Missing Data Problem" - A conceptual visualization of the future of biopharma R&D depicting the integration of domain-specif...

The next 18 months will likely see expanded cloud availability of domain-specific foundation models. Lower proof-of-concept barriers for R&D teams across biopharma. More ELN and LIMS integrations mean foundation models become embedded in routine workflows rather than standalone research projects. "Virtual cells" and "virtual tissues" will mature from academic curiosities to partnership-backed tools for cohort enrichment and patient stratification. The Noetik-GSK and Noetik-Agenus deals point in that direction.

DNA→RNA and pathology→omics imputation models will see real-world trials aimed at reducing re-assay costs and recovering statistical power in under-sized clinical cohorts. Expect guidance requests to the EMA and FDA on acceptable use in exploratory biomarkers. If the models prove robust enough—and that remains a meaningful if—they could change how sponsors think about trial design. Collect cheaper modalities up front. Impute expensive ones computationally. Validate predictions on a subset. Use the augmented dataset for hypothesis generation.

Protein design foundation models like ESM3 and Chroma, along with AI-generated gene editors like OpenCRISPR-1, will face more wet-lab benchmarking as the community stress-tests their predictions. Open releases will likely continue to spur ecosystem adoption, though commercial deployments may remain proprietary. AlphaFold 3 and future competitors will keep pushing the envelope on complex biomolecular interaction modeling. Access models and benchmarking protocols—CASP and beyond—will shape whether academic or industry applications dominate.

Regulatory clarity will improve incrementally. The FDA's January 2025 draft guidance is a starting point, not the final word. As more companies submit AI-enabled data in INDs and NDAs, case law will accumulate. Both sponsors and regulators will develop shared intuitions about what constitutes adequate validation. Data security requirements for genomic datasets will force infrastructure investments in secure compute environments and auditable pipelines.

The skepticism is warranted. Many of the performance and productivity claims come from consultants, vendors, or company press releases rather than peer-reviewed independent studies. The track record of transformative technology in drug development is littered with overpromises and underwhelming results.

But the trajectory feels different this time. Biological foundation models have moved from research prototypes to production systems faster than most observers expected. Whether they deliver on the promise of fundamentally cheaper, faster drug development depends on execution. By startups like Strand AI. By platform providers like NVIDIA and AWS. By pharma partners willing to bet billions on the technology. By regulators willing to accept computationally augmented datasets as valid scientific evidence.

The inflection point isn't coming. It's here—messier and more uncertain than the pitch decks suggest, but undeniably here. The question now is how quickly the industry figures out which problems these models can actually solve versus which ones still require old-fashioned experiments and time.

And whether anyone's willing to be honest about the difference.

More stories

  • Rhem Labs launches AI robot for aging-in-place monitoring
  • ai3Bio raises $48M to reset immune systems for remission
  • MiniMax's M2.5 Hits 80% on SWE-bench, Matches Frontier AI at 1/10 Cost
  • How AI Is Solving Drug Development's $18B Characterization Problem
  • Blissclub Taps Startup Founders, Not Athletes, for Men's Launch
  • Anterior Raises $3.2M Seed to Automate Healthcare Prior Authorization
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.