Founderland Logofounderland
the ★ top ★ 100 ★ marketers ★
SavedSearch
FoundersFounders
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Product Launches
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Investment News
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
Research & Innovation
Industries
Fintech iconFintechClimate / Social Tech iconClimate / Social TechSaaS iconSaaSHealthtech & Biotech iconHealthtech & BiotecheCommerce iconeCommerceMedia & Entertainment iconMedia & Entertainment
FoundersFounders
Return

Recommended Articles

Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

Rhem Labs launches AI robot for aging-in-place monitoring

Rhem Labs launches AI robot for aging-in-place monitoring
YcSenior Care+3
Healthtech & Biotech iconHealthtech & BiotechOctober 4, 2026

ai3Bio raises $48M to reset immune systems for remission

ai3Bio raises $48M to reset immune systems for remission
BiotechAutoimmune Disease+3
Healthtech & Biotech iconHealthtech & BiotechFebruary 21, 2026

AI Co-Scientists Take Over Labs in Research Automation Race

AI Co-Scientists Take Over Labs in Research Automation Race
Lab AutomationAi Agents+3
Fintech iconFintechFebruary 21, 2026

Inscope Raises $14.5M Series A for AI Financial Reporting Platform

Inscope Raises $14.5M Series A for AI Financial Reporting Platform
YcAccounting Automation+3

Founders Mentioned

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS

Yue Dai

Strand AI

saas icon
SaaS

Oded Falik

Strand AI

saas icon
SaaS
Healthtech & Biotech iconHealthtech & Biotech
February 21, 2026
YcBiotechDrug DiscoveryClinical TrialsMachine Learning

Strand AI Launches Foundation Models to Fill Gaps in Biological Data

YC-backed startup enters race to build 'foundation models for biology,' promising to predict missing modalities in clinical trials and reduce costly lab assays.

Strand AI Launches Foundation Models to Fill Gaps in Biological Data

The pathology slides came back complete—nearly 300 of them, stacked neatly in a freezer in Boston. Whole genome sequences? Check, from about half the clinical trial participants. Proteomics? Maybe thirty samples, if the research coordinator was lucky and the assays didn't fail. Everything else? Gone.

Not missing because someone screwed up. Missing because running comprehensive molecular profiles on every patient in a Phase II trial costs north of $5,000 per person, sometimes double that. And certain assays—spatial proteomics, say, or tissue-specific RNA expression—simply don't scale when you're enrolling hundreds of people across multiple sites.

Which brings us to Strand AI.

The San Francisco startup, barely out of Y Combinator's Winter 2026 batch, is making what sounds like a preposterous pitch: it can generate the missing biological data. Not extrapolate it. Not estimate it statistically. Generate it—using foundation models trained to predict one type of molecular measurement from another. Co-founded by Yue Dai, who left the computational pathology startup Pathos to launch Strand, and CTO Oded Falik, the company raised a $500,000 convertible note and started shipping datasets almost immediately.

Last month, Falik took to LinkedIn to announce Strand's first public release: predicted RNA expression profiles spanning 45 tissues and 4,500 genes for more than 500 genomes in the 1000 Genomes Project expansion. The twist? None of those RNA measurements were ever actually collected. Strand generated them by running the Chan Zuckerberg Initiative's VariantFormer—a 1.2 billion-parameter transformer trained to predict gene expression from DNA sequences—and claims to have optimized the pipeline to run 37 to 65 times faster and cheaper than the academic baseline.

Whether this constitutes a genuine scientific advance or clever infrastructure repackaging remains an open question.

Biology's Data Problem

Multimodal biological datasets are, by nature, full of holes. Always have been. A 2024 review in Proteomics laid it out plainly: real-world biomedical data rarely captures every modality for every sample. Patients drop out. Assays fail halfway through. Budget constraints force researchers to triage—do we spend the remaining funds on proteomics for twenty people, or imaging for a hundred?

Statisticians have names for these patterns. "Missing completely at random" (MCAR) is the best-case scenario—pure chance determines what's absent. "Missing at random" (MAR) means the missingness is predictable from other observed variables. But "missing not at random" (MNAR)? That's when the absence itself carries information, and that's when things get dicey. If your sickest patients are the ones who can't complete the proteomics follow-up, any attempt to fill those gaps computationally risks baking in bias rather than correcting for it.

The scale of modern resources is staggering, yet coverage remains frustratingly patchy. The UK Biobank's March 2025 data release includes full-body imaging for 100,000 participants, brain MRI surface mapping for over 63,000, and polygenic risk scores for 485,000—but not every participant has every modality. HuBMAP now holds more than 5,000 datasets across 27 organs and 22 data types, yet spatial proteomics and spatial transcriptomics rarely overlap cleanly. The NIH's All of Us research program has sequenced 245,388 clinical-grade genomes, with 77% from underrepresented groups, but linking those genomes to corresponding tissue expression or imaging data requires additional layers of consent and collection that many participants never complete.

For pharma and biotech, this sparsity translates directly into cost. Running RNA-seq, proteomics, or spatial assays on every trial participant can easily add $5,000 to $20,000 per patient. Multiply that by a few hundred enrollees, and you're looking at millions in assay costs before the first efficacy readout.

Strand's pitch is elegantly simple: predict the expensive modality from the cheap, routinely collected data. Validate on a subset. Pocket the savings—or use them to expand your cohort.

What Strand Actually Built

Strand calls itself a "data augmentation company for biology." Its first public demonstration leveraged VariantFormer, the CZI Biohub model trained on GTEx, MAGE, ADNI, and ENCODE datasets to predict tissue-specific gene expression from whole genome sequences. According to Falik's post, Strand optimized the inference pipeline—comparing A100 versus H100 GPU costs, wringing out inefficiencies—and released the resulting synthetic paired WGS-RNA expression data for free, complete with an interactive visualizer.

Here's the catch: the underlying model isn't Strand's. VariantFormer is open-source, with a detailed model card published by the Chan Zuckerberg Initiative's Virtual Cell Models project. What Strand offers, in essence, is operationalization—taking academic foundation models, making them production-ready, and wrapping them in pipelines that pharma companies can actually deploy without hiring a team of ML engineers.

Whether running optimized inference on someone else's pretrained model and distributing the outputs constitutes defensible intellectual property is, well, debatable. Dai's stated plan includes recurring public dataset releases, which suggests Strand may be building credibility and distribution before unveiling proprietary models. Or perhaps the company is betting that execution and packaging matter more than model ownership—a not-unreasonable position in an ecosystem where many academic labs publish cutting-edge architectures but lack the infrastructure to commercialize them.

The company's website emphasizes use cases that sound plausible enough: filling gaps in clinical trial datasets, reducing costly assays by predicting proteomic or transcriptomic profiles from pathology slides, enabling rare-disease research where sample sizes are inherently tiny, and accelerating biomarker discovery via imputation. All reasonable. All also hard to validate rigorously before deployment.

A Crowded Field, Getting More So

Strand hardly enters virgin territory. The field has moved from academic curiosity to commercial deployment with surprising speed. Owkin's HE2RNA, published in Nature Communications back in 2020, demonstrated that deep learning could predict bulk RNA-seq expression from hematoxylin and eosin (H&E) stained whole-slide images—the kind of slides every pathology department already generates. The company followed with partnerships at Sanofi for target identification and drug positioning, embedding the technology in real discovery workflows.

Newer architectures keep pushing performance. SEQUOIA, published in Nature Communications in late 2024, uses linearized attention mechanisms and reportedly outperforms HE2RNA on some benchmarks. A 2025 Nature Communications paper benchmarked 11 different methods for predicting spatial gene expression from histology images. Virtual immunohistochemistry pipelines now use diffusion models to generate molecular stain proxies from standard H&E, improving pancreatic cancer subtyping and sidestepping expensive multiplexed antibody panels.

Single-cell foundation models have attracted even more capital and hype. scGPT, published in Nature Methods in 2024, was pretrained on over 33 million cells and claims to handle batch integration, multi-omic integration, and perturbation prediction. CellFM, released in Nature Communications in early 2025, scaled to 100 million human cells with an 800 million-parameter ERetNet backbone. The scvi-tools ecosystem—anchored by models like MultiVI and totalVI—offers pretrained probabilistic approaches for imputing missing modalities with calibrated uncertainty estimates, all distributed via the scvi-hub.

But critical evaluations have started to temper the enthusiasm. A Microsoft Research study and accompanying Genome Biology article found that zero-shot benchmarks expose serious limitations in single-cell foundation models. In many cases, scGPT and Geneformer failed to beat simpler baseline methods without task-specific fine-tuning. Batch effects—the technical noise introduced by different labs, instruments, or protocols—remained poorly handled.

The field is learning, perhaps a bit belatedly, that pretraining scale doesn't automatically guarantee generalization. Bigger isn't always better, at least not without careful validation.

Pharma's Appetite Is Real

Digital illustration for article section "Pharma's Appetite Is Real" in "Strand AI Launches Foundation Models to Fill Gaps in Biological Data" - A conceptual macro photography scene depicting the high-stakes intersection of pharmaceuticals and a...

Whatever the technical hurdles, pharma's appetite for multimodal AI is undeniable. Isomorphic Labs, DeepMind's drug discovery spinout, inked partnerships with Eli Lilly and Novartis worth nearly $3 billion in early 2024, then raised roughly $600 million in 2025 and expanded its Novartis collaboration again this February. Recursion Pharmaceuticals received a $50 million NVIDIA investment in 2023 to build biology and chemistry foundation models and has since announced the BioHive-2 supercomputer along with partnerships spanning Sanofi, Roche, and Genentech.

Noetik, which builds "world model" foundation models for spatial tumor data, signed a five-year licensing deal with GSK in January 2026—a shift from bespoke AI consulting services to reusable model infrastructure. Relation Therapeutics closed a partnership with GSK potentially worth up to $300 million. Tempus AI, which operates large multimodal real-world datasets, announced a collaboration with BioNTech and released Tempus One GenAI capabilities for analyzing unstructured multimodal clinical data.

PathAI received FDA 510(k) clearance in 2025 for its AISight Dx IMS platform for primary diagnosis and expanded label claims to Roche's DP200 and DP600 digital pathology scanners via the agency's Predetermined Change Control Plan pathway—a regulatory innovation designed to pre-approve certain types of algorithm updates. The company also launched a Precision Pathology Network to aggregate real-world AI-pathology data and formalized a companion diagnostic collaboration with Roche. These regulatory clearances lay critical infrastructure for deploying image-centric multimodal models in clinical workflows.

Market sizing varies widely depending on who's counting and how generously they define the category. Grand View Research pegs the AI drug discovery market at $2.35 billion in 2025, growing to $13.77 billion by 2033—a compound annual growth rate near 25%. Mordor Intelligence estimates $2.58 billion this year expanding to $10.29 billion by 2031. McKinsey, always the optimist, projects generative AI could unlock $60 billion to $110 billion in annual value across life sciences.

Yet a recent McKinsey survey found only 5% of organizations have realized competitive differentiation from these tools. That gap underscores how difficult scaling these technologies remains, even when the promise seems obvious on paper.

The Regulatory Thicket

Missing data imputation is statistically treacherous. When data are missing not at random—say, your sickest patients drop out before proteomics collection, or certain assays systematically fail in tumor subtypes with poor tissue quality—model-based imputation can amplify bias rather than correct for it. A 2024 review in the Journal of the American Society for Mass Spectrometry warned that metabolomics and omics imputation methods risk introducing systematic errors if the missingness mechanism is misunderstood.

It's not a hypothetical concern.

Regulatory expectations are tightening in response. The FDA's Center for Drug Evaluation and Research established an AI Council in 2024 and released draft guidance in 2025 on using AI to support regulatory decision-making, with "Guiding Principles of Good AI Practice in Drug Development" expected in 2026. The International Medical Device Regulators Forum published final "Good Machine Learning Practice" principles for AI/ML medical devices in January 2025, emphasizing transparency, risk management, and high-quality training datasets. The EU AI Act, which entered into force in August 2024, classifies high-risk healthcare AI and mandates conformity assessments, with staged deadlines extending into 2027.

The challenge for companies like Strand is proving that synthetic data—RNA expression never measured, only predicted—can pass muster in regulatory submissions or clinical validation studies. ICH E9(R1), the FDA's adopted addendum on estimands and sensitivity analysis, requires rigorous handling of intercurrent events and missing data in clinical trials. Imputation is permitted, sure. But sponsors must justify their assumptions and demonstrate robustness through sensitivity analyses. Hand-waving doesn't cut it.

Daphne Koller, CEO of Insitro, has argued in interviews that biology still lacks the foundational data infrastructure needed for an "AlphaFold-like" breakthrough in broader disease modeling. High-quality human data paired with experimental validation systems remain prerequisites, not luxuries. McKinsey's surveys suggest that while the estimated value from these tools is enormous, execution gaps around data governance, model validation, and integration into legacy workflows have left most organizations struggling to capture any of it.

What Success Looks Like

Digital illustration for article section "What Success Looks Like" in "Strand AI Launches Foundation Models to Fill Gaps in Biological Data" - A sophisticated macro photography composition depicting a miniature world of biological infrastructu...

Strand AI's early dataset releases position it as a potential infrastructure play—distribute free generated data, build community goodwill, and later monetize proprietary models or enterprise APIs. It's a familiar playbook from the open-source software world, though whether it translates to biology remains unclear.

The company's bet aligns with several converging trends: expanding multimodal public datasets (UK Biobank imaging updates, HuBMAP scale-ups), growing regulatory clarity for AI in drug development and digital pathology, and pharma's shift from paying for bespoke AI consulting to licensing reusable foundation models that can be deployed across multiple programs.

Virtual assays—predicting one modality from another—will likely see fastest adoption where regulatory rails already exist. Digital pathology for primary diagnosis is FDA-cleared. Virtual immunohistochemistry for triage before confirmatory wet-lab staining has clear cost savings and established validation frameworks. Imaging-to-omics transformations benefit from histology being routinely collected and relatively cheap compared to spatial proteomics or transcriptomics.

Single-cell foundation models, despite the breathless hype, face steeper hurdles. Zero-shot performance remains limited, and domain-specific fine-tuning is almost always required. The scvi-tools ecosystem's emphasis on pretrained probabilistic models with uncertainty quantification may prove more pragmatic than end-to-end transformers that struggle to generalize across batch effects and technology platforms.

For founders and R&D executives evaluating these tools, the calculus hinges on validation rigor, regulatory alignment, and transparency. The EU AI Act's high-risk requirements and FDA's emerging AI guidance favor vendors who ship model cards, bias assessments, and calibrated uncertainty estimates—features already standard in academic tools like MultiVI and VariantFormer. Strand's decision to release datasets publicly and build atop open models could prove shrewd if regulatory scrutiny rewards transparency over proprietary black boxes.

Then again, it could also prove naïve if competitors with proprietary architectures capture enterprise contracts while Strand gives away its lunch.

The Endgame

Digital illustration for article section "The Endgame" in "Strand AI Launches Foundation Models to Fill Gaps in Biological Data" - A conceptual macro photography scene depicting the high-stakes race to build foundation models for b...

The race to build foundation models for biology is accelerating, funding rounds are getting larger, and the rhetoric is getting bolder. But the winners won't necessarily be those with the largest parameter counts or the flashiest pretraining runs. They'll be the teams that understand how to make predictions trustworthy enough to stake a clinical trial on—and cheap enough that pharma will choose synthetic data over the real thing.

That's a higher bar than it sounds. Biology is messy in ways that language and images are not. Batch effects plague every dataset. Tissue heterogeneity confounds every model. And the consequences of getting it wrong aren't embarrassing chatbot responses—they're failed drug candidates and wasted years.

Strand AI is betting it can thread that needle by operationalizing academic models, building credibility through open releases, and moving fast while the regulatory landscape is still taking shape. Whether that bet pays off depends less on the elegance of the technology and more on the unglamorous work of validation, documentation, and earning trust from risk-averse pharma executives who've seen plenty of AI pitches before.

The synthetic biology play is underway. Now comes the hard part: proving the predictions are real enough to matter.

More stories

  • Rhem Labs launches AI robot for aging-in-place monitoring
  • ai3Bio raises $48M to reset immune systems for remission
  • AI Co-Scientists Take Over Labs in Research Automation Race
  • Inscope Raises $14.5M Series A for AI Financial Reporting Platform
  • The Visual AI Race: Who's Really Winning on Phones and Wearables
  • Tidy Launches No-Code AI Assistant That Learns Any App via iMessage
fintech icon
climate-social-tech icon
saas icon
healthtech-biotech icon
ecommerce icon
media-entertainment icon
Loading...

About

Dreamwell AIContact UsOur Story

Articles

Product LaunchesInvestment NewsResearch & Innovation

founderland

We Use Cookies

We baked up some cookies – the digital kind. They help Draper run like a well-oiled mid-century machine. Some are essential to the experience, others help us tailor things to your taste. We promise, no crumbs on your blazer. Take a moment to choose what works for you.