Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

Rhem Labs launches AI robot for aging-in-place monitoring

Rhem Labs launches AI robot for aging-in-place monitoring
YcSenior Care+3
Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

ai3Bio raises $48M to reset immune systems for remission

ai3Bio raises $48M to reset immune systems for remission
BiotechAutoimmune Disease+3
SaaS iconSaaSFebruary 27, 2026

UpGuard Raises $75M Series C for AI-Powered Cyber Risk Platform

UpGuard Raises $75M Series C for AI-Powered Cyber Risk Platform
Cyber SecurityArtificial Intelligence+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 27, 2026

Honest Health Lands $140M to Scale Value-Based Senior Care Platform

Honest Health Lands $140M to Scale Value-Based Senior Care Platform
HealthtechMedicare Advantage+2

Founders Mentioned

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS
Healthtech & Biotech iconHealthtech & Biotech
February 27, 2026
Drug DiscoveryArtificial IntelligenceBiotechMachine LearningPrecision Medicine

Foundation Models Are Learning to Fill Biology's Data Gaps

AI systems can now predict missing proteins, genes, and spatial data from routine lab tests—unlocking archived patient samples and reshaping drug discovery economics.

Foundation Models Are Learning to Fill Biology's Data Gaps

The tissue samples sit in freezers, carefully archived from clinical trials that wrapped years ago. Each one represents a patient, a moment in disease progression, perhaps a clue to why a promising cancer drug worked brilliantly for some but failed others. Researchers would love to run fresh assays—spatial proteomics, gene expression panels, the works. But there's never enough tissue. And retrospective analysis burns through what remains, often without answers.

This is the chronic condition of pharmaceutical research: every trial runs on partial information. A Phase III study might capture genomic data for 80% of enrolled patients, spatial proteomics for maybe 15%, gene expression for half if you're lucky. All three modalities for a single person? Almost never.

Now a new crop of computational tools is attempting something that sounds almost too convenient: filling those gaps without touching a tissue sample. Instead of physically staining slides or extracting RNA, these foundation models predict missing biological measurements from whatever data already exists. Feed them a routine H&E histology slide, and they'll generate virtual spatial proteomics. Hand over a patient's genome, and out comes predicted gene expression across tissues.

The implications ripple quickly through an industry built on expensive, time-consuming wet-lab work. In January, a team published results in Nature Medicine demonstrating they could predict 40 protein biomarkers from standard pathology slides alone—no immunohistochemistry staining required. The model, called HEX, was validated across 2,298 non-small cell lung cancer samples and matched or exceeded clinical risk scores while predicting immunotherapy response in advanced disease. One histology slide, digitized and analyzed, standing in for dozens of traditional assays.

If it sounds too good to be true, well, that's the question the field is wrestling with now.

Why Every Dataset Has Holes

Biological data collection is inherently messy. The UK Biobank holds whole-genome sequences for half a million people, but proteomics and metabolomics exist for only a fraction. The Cancer Genome Atlas cataloged genomics and transcriptomics for thousands of tumors, yet spatial context—which cells neighbor each other, how the immune microenvironment organizes itself—was captured for select samples only.

Clinical trials make this worse. Protocols change mid-study. New biomarker hypotheses surface after enrollment closes. A promising signal emerges in a subset too small to analyze with statistical confidence. Traditional solutions involve either going back to collect more data prospectively (expensive, slow, sometimes impossible if the trial has ended) or simply living with uncertainty.

Pathology labs occasionally re-cut archived FFPE blocks for additional stains, assuming enough tissue remains and the assay still works on aged samples. Genomic imputation has existed for years, inferring missing single-nucleotide polymorphisms from reference panels. But predicting entirely different data types—proteins from images, RNA from DNA—only became feasible at scale recently, maybe within the past 24 months.

The technical pieces came together quietly. Large paired datasets materialized from biobanks and research consortia. Self-supervised learning matured, letting models learn patterns from unlabeled biological data before fine-tuning on specific tasks. Compute got cheaper and more accessible—NVIDIA's BioNeMo platform, which launched large language model services for life sciences in 2022, has become something like standard infrastructure for pharma AI experiments.

The Convergence

Digital illustration for article section "The Convergence" in "Foundation Models Are Learning to Fill Biology's Data Gaps" - A surreal, highly textured macro composition visualizing the convergence of artificial intelligence ...

Foundation models trained on biological data share architectural lineage with the transformers that power ChatGPT, but the training corpus is different: millions of histopathology images, billions of genomic sequences, single-cell transcriptomes from 100 million cells or more.

Take VariantFormer, released by Chan Zuckerberg Biohub last November. The 1.2-billion-parameter model was trained on GTEx, ENCODE, and other datasets to predict gene expression from personal genomes across tissues. It outperforms older methods like Enformer in several benchmarks and can generate patient-specific expression profiles conditioned on individual genetic variants. Not perfect, but accurate enough to be useful.

The use case becomes clear quickly: you can augment sparse cohorts computationally. Strand AI, a startup founded by Yue Dai and Oded Falik (both previously at Enable Medicine, a spatial biology data company), recently optimized VariantFormer's inference speed—37 times faster on A100 GPUs—and ran it across the 1000 Genomes "expansion pack," generating RNA-seq predictions for over 500 individuals not in the original training set. They released the data publicly with an interactive visualizer.

Strand's pitch centers on turning "incomplete patient profiles into complete, multimodal datasets." Cross-modal prediction as a service layer for biomarker discovery and trial stratification. The company was accepted into Y Combinator's Winter 2026 batch and claims—though hasn't yet published—that its first foundation model predicts spatial proteomics from standard histology, "beating state of the art."

There's a whiff of hype in those claims, perhaps, but the academic literature has been moving fast enough to keep startup ambitions grounded in reality. A 2020 Nature Communications paper demonstrated gene-expression inference from whole-slide images (HE2RNA). By 2025, multiple groups had published spatial transcriptomics prediction from H&E using hypergraph learning and cross-modal mask reconstruction. The January 2026 HEX publication represents something more substantial: prognostic value across multiple institutions, predictive power for immunotherapy response, integration with existing foundation-model backbones. PictorLabs, another entrant, partnered with PathPresenter and Hamamatsu to deliver AI-powered virtual stains—labeled research-use-only for now.

Market forces are doing their part. The AI-in-drug-discovery market was sized at $3.6 billion in 2024 by Global Market Insights, with projections reaching $49.5 billion by 2034—a 30.1% compound annual growth rate. McKinsey estimated in 2024 that generative AI could unlock $60 to $110 billion in annual value for pharma and medtech. BCG's January 2026 analysis floated the idea that AI-first biopharma companies might see 5% to 15% revenue uplifts through faster timelines: discovery compressed to months, clinical trials shortened by a fifth, manufacturing yields up more than 20%.

These are consulting projections, not peer-reviewed science. But they capture the executive appetite for tools that de-risk pipelines and stretch budgets.

When Virtual Beats Physical

Digital illustration for article section "When Virtual Beats Physical" in "Foundation Models Are Learning to Fill Biology's Data Gaps" - Generate a realistic image of microscopic tissue slides under a modern, high-tech microscope to repr...

Cross-modal prediction has delivered results in places where traditional biomarkers stumble. Owkin, a French computational pathology company, showed that a multimodal model combining histology and clinical data outperformed PD-L1 expression—the current standard—for predicting progression-free survival and overall survival in NSCLC patients receiving immunotherapy. The concordance index improved by 15 points. PD-L1 testing isn't going anywhere, but computed biomarkers derived from routine slides and electronic health records might refine patient selection without consuming additional tissue.

HEX's validation remains the most extensive published to date. Trained on 755,000 co-registered tiles and validated across independent datasets, the model improved prognosis predictions beyond clinical risk scores and correctly stratified 148 advanced NSCLC patients by immunotherapy response. The authors noted that virtual spatial proteomics enables retrospective biomarker screening—you can interrogate archived cohorts for protein signatures without physically re-staining anything.

The appeal for pharma is obvious. Completed trials can be re-analyzed for secondary endpoints or exploratory hypotheses. Rare-disease studies with limited samples extract more information from the same tissue. Every archived block becomes computationally renewable, in a sense.

At the sequence level, models are tackling biology's central dogma head-on. GET, published in Nature in January 2025, builds a foundation model of transcription from chromatin accessibility and DNA sequence. EvolutionaryScale's ESM-3, a 98-billion-parameter model released mid-2024 and published in Science this January, reasons jointly over protein sequence, structure, and function—its API is in public beta and was used to design a novel fluorescent protein. Google DeepMind's AlphaFold 3, released in May 2024, expanded from protein-structure prediction to modeling interactions among proteins, nucleic acids, and small molecules. Isomorphic Labs, DeepMind's drug-discovery spinout, secured multi-billion-dollar partnerships with Eli Lilly and Novartis to co-develop therapeutics using these tools.

The single-cell field is consolidating around similar architectures. A 2024 Nature Methods benchmark evaluated 47 datasets and found that totalVI and scArches perform well for protein prediction from RNA, while newer models like CellFM (trained on 100 million cells) extend to perturbation prediction and integrative analysis. Early-stage biotech is adopting these rapidly: Blank Bio, another YC startup from Summer 2025, builds "foundation models that understand RNA" and counts Sanofi and GSK as partners for virtual cell modeling.

What Comes Next

Digital illustration for article section "What Comes Next" in "Foundation Models Are Learning to Fill Biology's Data Gaps" - A conceptual macro visualization of a regulatory timeline for the EU AI Act depicted as a series of ...

Regulators are catching up, though not quickly. The EU AI Act entered force in stages—prohibitions took effect in February 2025, general-purpose AI obligations in August, transparency requirements coming in August 2026, high-risk embedded systems in August 2027. Digital pathology and computed biomarker tools may qualify as high-risk depending on their role in clinical decisions. The European Medicines Agency finalized a reflection paper in September 2024 outlining principles for AI across the medicines lifecycle, encouraging early scientific advice.

In the U.S., the FDA issued comprehensive draft guidance in 2025 for AI-enabled medical devices, emphasizing total product lifecycle management and transparency, alongside 2023 final guidance for digital health technologies in clinical trials.

Computed biomarkers used for trial enrichment or as endpoints will need analytical and clinical validation—no shortcuts there. The FDA's Biomarker Qualification Program offers a pathway to qualify biomarkers, including computed ones, for specific contexts of use. That can streamline regulatory acceptance across programs. Real-world evidence is increasingly part of the conversation; an ISPOR analysis of 2021-2022 approvals found RWE frequently included in regulatory submissions, though controlling for bias remains critical.

Technical challenges linger. Distribution shift across sites and staining protocols remains a genuine risk, especially for models anchored to histology. HEX demonstrated strong external validation, but extensive clinical validation will be required before these tools graduate from research-use-only to diagnostic or companion diagnostic status. Evaluation metrics for virtual omics must correlate with biological and clinical endpoints, not just image fidelity. A 2025 arXiv paper critiqued common image-quality metrics like FID and SSIM, arguing for stain-accuracy measures and whole-slide assessments instead.

Commercial models are still forming. Options include dataset licensing, model-as-a-service APIs for virtual assays, and co-development partnerships for indication-specific biomarkers. NVIDIA's BioNeMo and platforms like EvolutionaryScale's Forge API are standardizing access, making enterprise pilots easier to launch. The buyer personas: translational oncology biomarker teams, clinical development groups exploring computed stratification, real-world-data analytics teams building synthetic control arms.

The long bet, then, is that patient cohorts will become computationally completable. A genome and a slide might be enough to predict spatial proteomics, transcriptomics, maybe metabolomics. Not perfectly—perfection isn't the point. But well enough to generate hypotheses, stratify trials, unlock archived samples.

Strand AI's founders frame it as filling gaps to enable better science. The broader industry positions it as infrastructure, the next layer in a data stack that's been incomplete from the start. Either way, the economics shift when historical samples can be interrogated without cutting another section, when every trial can be re-analyzed without destroying what's left in the freezer.

Whether the promise holds depends on validation that's only beginning. But the models are running, the datasets are growing, and the tissue blocks—for now—remain intact.

More stories

  • Rhem Labs launches AI robot for aging-in-place monitoring
  • ai3Bio raises $48M to reset immune systems for remission
  • UpGuard Raises $75M Series C for AI-Powered Cyber Risk Platform
  • Honest Health Lands $140M to Scale Value-Based Senior Care Platform
  • SHINE Technologies Proves Fusion's Commercial Value Beyond Energy
  • Ariel Investments Closes $250M Fund for Women's Sports
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.