Somewhere in a pharmaceutical company's archives sits a freezer full of tissue samples from a clinical trial that wrapped up five years ago. The samples have genomic data attached—sequencing was routine by then—but nobody ran the spatial proteomics assays that might have revealed why some patients responded to treatment and others didn't. Those assays cost thousands of dollars per sample. The trial is over. The budget is gone.
This is the kind of problem that keeps biotech executives up at night. Could you reconstruct the missing biology computationally, without thawing a single sample?
A handful of startups now argue the answer is yes—or at least, close enough to yes that it's worth trying. They're building what they call foundation models, borrowing the term from the language AI boom, and treating biological data as a translation problem. Feed the model a standard pathology slide, the kind pathologists have been staining since the 19th century, and it predicts the expensive molecular measurements you never ran.
The pitch resonates in an industry where retrospective analysis often hits a wall because someone years ago didn't collect the "right" data. And where a single spatial proteomics run can still cost enough that smaller biotech teams have to pick and choose which samples get profiled.
Whether these models actually work well enough to matter—well, that's the bet being tested in real time.
The Missing Data Problem, Quantified
Clinical trials in oncology routinely collect tissue samples. Comprehensive molecular profiling, though? That's selective. Budget constraints, sample degradation, shifting biomarker hypotheses—all conspire to leave researchers with datasets full of holes.
The numbers suggest the opportunity is substantial, if not precisely defined. McKinsey estimates from early 2024 projected that generative AI could add somewhere between $60 billion and $110 billion in annual value to pharma and medical products. Some portion of that, presumably, comes from squeezing more insight out of datasets that already exist.
Spatial omics represents one pressure point. Market research pegs the sector at roughly $498 million in 2025, climbing to about $550 million this year, with forecasts reaching $911 million by 2031. Companies like 10x Genomics, Akoya Biosciences, and Bruker keep releasing new platforms. Each assay adds capability. Each also adds cost and operational complexity.
If a model could predict spatial protein expression from a hematoxylin and eosin slide—the workhorse stain of pathology for more than a century—the economics shift. Maybe not entirely. But enough to change workflow decisions about which samples merit the expensive follow-up.
Academic Roots, Commercial Ambitions
The idea didn't emerge from startups first. Owkin published work in Nature Communications back in August 2020 showing that gene expression could be inferred from histology images. The paper, HE2RNA, demonstrated proof of concept. By 2025, the field had matured enough that a Nature Medicine paper—published online in January this year—showed virtual spatial proteomics from H&E slides improving prognostic models in lung cancer cohorts.
The mechanistic logic is straightforward: tissue morphology encodes molecular state. Cells arrange themselves in patterns. Those patterns correlate with gene and protein expression. Train a model on enough paired samples, and it learns the relationship.
That said, performance gaps remain. A benchmarking study in Nature Communications from 2025 found models predicting spatial gene expression from histology reached concordance indices around 0.52 for luminal breast cancer, compared to 0.58 for RNA sequencing. It works, in other words. But it's not a drop-in replacement.
What these models offer, perhaps more realistically, is triage. They can flag which samples warrant expensive assays and which don't. They can retrospectively enrich archived cohorts that can't be re-measured because the samples no longer exist or the money ran out years ago.
Strand AI: Two People, One Thesis
Strand AI entered Y Combinator's Winter 2026 batch as a two-person team with a straightforward value proposition: we'll be the missing data layer for biology foundation models.
Co-founders Oded Falik and Yue Dai bring complementary backgrounds. Falik led product at Enable Medicine. Dai worked on foundation models at Pathos AI—a Tempus initiative—and before that at Microsoft Research and Element AI. They launched the company in 2025. Their first product, POSTMAN, hit early access in March this year. It predicts spatial proteomics from routine pathology slides.
The use cases Strand AI describes sound practical rather than revolutionary: retrospective cohort profiling, triage before expensive assays, hunting for unmeasured protein biomarkers in archived samples. In January, the team released a dataset called VariantFormer—imputed RNA expression across 4,500 genes, 45 tissues, and 538 samples from the 1000 Genomes project.
One engineering detail stands out. They tuned their inference pipeline to run 37 times faster on A100 GPUs compared to H100s. That's prioritizing cost over raw speed, a choice that signals awareness of the economic realities facing smaller biotech teams who might be their customers.
Strand AI positions itself as a data augmentation layer. For clinical trial sponsors with incomplete molecular profiling across patient cohorts, the model offers computational gap-filling. The company is actively seeking design partners—pharma and biotech teams running trials or biomarker studies.
It's early days. Strand AI has two employees, operates out of San Francisco, and hasn't disclosed funding beyond the standard YC investment. But the timing aligns with broader industry momentum around multimodal biological AI.
A Crowded, Well-Funded Field

Strand AI has plenty of company, some with substantially deeper pockets.
Pathos AI announced a $365 million Series D in May last year at a post-money valuation around $1.6 billion. It's building an oncology foundation model in partnership with AstraZeneca and Tempus. That collaboration, announced in April 2025, included $200 million in data license fees to Tempus, which went public in June 2024.
In Europe, Bioptimus raised $35 million in seed funding in February 2024 and had reached $76 million in total funding by January this year. The French company recently released M-Optimus, a multimodal foundation model integrating histology, spatial and bulk RNA, and clinical data.
Paige, working with Microsoft Research, has developed PRISM and PRISM2—slide-level foundation models trained on datasets measured in the millions of slides. Enable Medicine and Akoya Biosciences launched what they described as the largest commercially available single-cell spatial proteomics atlas in April last year, creating a resource that both validates and trains cross-modal prediction models.
Even infrastructure providers are moving. NVIDIA announced in January that its BioNeMo platform had been adopted by major life sciences players and released new open models including RNAPro for RNA structure prediction.
The common thread: a belief that multimodal biological datasets—genomics, transcriptomics, proteomics, imaging—represent interconnected views of the same underlying biology. And that sufficiently large models trained on paired samples can learn to translate between them.
It's a thesis borrowed directly from the language model playbook. Whether biology is actually analogous to language in the ways that matter for this approach—well, that remains an open question.
The Regulatory Unknown

The promise of data imputation runs straight into some uncomfortable questions. How accurate must a prediction be to guide clinical decisions? What level of validation satisfies regulators if an imputed biomarker informs how patients get stratified in a trial?
The FDA had authorized more than 1,200 AI and machine learning-enabled medical devices as of mid-to-late 2025. But radiology dominates that list—roughly 75 to 80 percent of clearances. Pathology represents a smaller share. Predictive molecular models occupy an even narrower sliver.
In January 2025, the FDA held a webinar on final guidance for Predetermined Change Control Plans, a mechanism allowing AI-enabled devices to update post-deployment in a managed way. The EU's AI Act, which entered force in August 2024, will impose high-risk requirements on medical device-embedded AI starting in August 2027 for products falling under the Medical Device Regulation or In Vitro Diagnostic Regulation.
For companies building foundation models that might eventually power clinical decision support tools, regulatory strategy isn't an afterthought. It's a fundamental constraint.
There are data governance wrinkles, too. Executive Order 14117, issued in February 2024, and a Department of Justice final rule effective April 2025, restrict outbound transfer of "bulk sensitive personal data"—explicitly including human 'omic data—to countries of concern. That impacts data licensing deals, cloud compute arrangements, and cross-border collaborations. Larger players with compliant infrastructure can navigate these constraints more easily than seed-stage startups operating on YC funding.
Good Enough Versus Ground Truth
Performance benchmarks remain in flux, which is another way of saying the technology isn't fully baked yet.
The Nature Communications benchmarking study from last year showed histology-to-spatial gene expression models lagging behind direct measurements. The Nature Medicine paper from January this year showed that integrating virtual spatial proteomics with H&E improved outcomes prediction in lung cancer. The technology adds value. It doesn't replace ground truth entirely.
That creates an interesting commercial position. These models function as decision support tools rather than substitutes for wet-lab assays. That may simplify the regulatory path—they're augmenting human judgment, not replacing it—while simultaneously limiting how large the addressable market can grow.
Pharma executives surveyed by McKinsey in 2024 and 2025 reported universal experimentation with generative AI. About 32 percent said they were scaling efforts. Budgets increased in 2025. Whether that investment translates into models that pass regulatory muster and genuinely accelerate drug development remains to be seen.
The Hypothesis Being Tested

The industry is placing bets that biological foundation models will compress costs and expand the utility of molecular profiling. Strand AI's POSTMAN represents one specific bet: pick a high-value, high-cost assay—spatial proteomics—and make it computationally accessible.
If the model can triage samples accurately, flagging which ones merit expensive follow-up, it delivers immediate return on investment. If it can impute protein biomarkers in retrospective cohorts, it unlocks datasets that would otherwise remain stranded in freezers and archives. Neither scenario requires the model to be perfect. Just useful enough to change a workflow decision.
The broader ecosystem suggests this isn't a niche curiosity. When companies valued in the billions are building similar infrastructure, when NVIDIA is releasing platform tools, when Nature-tier journals are publishing validation studies—the technology has crossed from academic curiosity into serious commercial consideration.
Whether it becomes standard practice in clinical trials or remains a specialized tool for retrospective analysis depends on how the next round of validation studies shake out, how regulators respond, and whether the predictions hold up in prospective use.
For now, the proposition is simple, maybe even obvious: biology generates messy, incomplete datasets, and foundation models offer a way to fill in the blanks computationally. The companies building those models are wagering that pharma will pay for higher-resolution views of patient cohorts, even if—especially if—those views are synthesized rather than measured.
Strand AI, two people strong and barely out of Y Combinator, is testing that hypothesis. So are a dozen other companies, some with hundreds of millions in funding and partnerships with Big Pharma. The question isn't whether anyone will try. It's whether anyone will succeed in a way that changes how drug development actually works.
