The slides were already in storage. Hundreds of them, tissue samples from a failed Phase II cancer trial, each one representing a patient who hadn't responded as hoped. Somewhere in those archived specimens might lie the answer to why—but pulling that answer out would require re-consenting participants (if they could even be found), shipping samples to specialized labs, and spending thousands of dollars per patient on spatial proteomics assays that didn't exist when the original study was designed.
For most drug developers, that math doesn't pencil out. The slides stay in storage. The insights stay locked away.
Unless, a small but growing group of companies now argues, you don't actually need to run those assays at all.
What's emerging is a new breed of artificial intelligence models that attempt to predict what expensive molecular tests would show—without ever running the tests. Feed them a standard pathology slide, the kind that's been stained with hematoxylin and eosin (H&E) for more than a century, and they'll computationally generate the spatial protein maps or gene expression profiles that previously required cutting-edge lab work. The promise: unlock decades of archived patient data, make clinical trials smarter about whom to enroll, and maybe—just maybe—inch up those brutal drug approval odds.
It's an audacious pitch. Whether it actually works at scale remains an open question.
The Data Problem Nobody Talks About
Here's what most people outside the pharmaceutical industry don't realize: clinical trials fail far more often because of incomplete biological information than because of bad science. A patient enrolled in an oncology study today might have tumor imaging, genomic sequencing, maybe even a liquid biopsy. But capturing the full picture—all 80-plus biological measurements that collectively describe how a cancer will behave—is prohibitively expensive. Logistically tangled. Often impossible when working with samples that were collected years ago under different protocols.
That incompleteness creates cascading problems. Biomarker discovery stalls because researchers can't see the molecular signatures they need in historical patient cohorts. "Trial enrichment"—the practice of selecting patients most likely to respond to a therapy—becomes guesswork when enrollment decisions rest on partial data. The consequences show up in the numbers: roughly 6.7% of drugs that enter Phase I testing eventually win approval, according to recent industry analyses. Incomplete patient stratification contributes to expensive failures in late-stage trials, where the stakes are highest.
Enter the imputation models. Rather than measuring everything from everyone, these AI systems use what's already been collected—often just those routine H&E slides—to predict the missing modalities. It's a bit like asking an experienced pathologist what she might expect to find if she could run every test imaginable, except the pathologist is a neural network trained on thousands of paired samples where both the slides and the molecular data actually exist.
The technical term is "cross-modal prediction." The practical translation: turning incomplete patient profiles into something closer to complete ones, at least computationally.
A Market Looking for Infrastructure
The spatial omics sector—technologies that map where proteins or RNA sit within a tissue sample—has been valued in the hundreds of millions, with some projections suggesting it could hit several billion dollars within the next decade. Yet adoption remains patchy. As of early 2025, Akoya Biosciences reported an installed base of 1,359 spatial phenotyping instruments worldwide, while 10x Genomics disclosed exactly 8,046 instruments sold through December 31, 2025. Respectable growth, perhaps, but a fraction of the laboratories running standard H&E pathology, the workhorse technique that's generated mountains of archived slides over decades.
Digital pathology—the infrastructure that turns glass slides into analyzable digital images—is expanding faster. Market estimates for 2025 hovered between $1.3 billion and $1.5 billion depending on whose research you trust, with projected annual growth rates ranging from around 9% to nearly 19%. Hospitals and research centers are digitizing tissue archives at scale, creating a curious asymmetry: tons of pixels, relatively little molecular context.
That gap has attracted attention. Owkin published early work in 2020 demonstrating RNA prediction from H&E slides and has continued developing spatial expression models. Lunit offers genotype prediction tools for research use. PathAI has presented findings linking tissue subtypes to gene expression signatures. Bioptimus, which announced cumulative funding in the neighborhood of $76 million as of early 2025, is building what it describes as a universal AI foundation model for biology. Recursion has partnerships with Roche and Bayer exploring related approaches.
The race, in other words, is on. Though perhaps "race" implies more clarity about the finish line than currently exists.
What's Driving This Now

Three forces are converging, each pushing this technology from lab curiosity toward something that might—maybe—reshape how clinical trials work.
First, the science is starting to deliver. A paper in Nature Medicine earlier this year described AI-derived "virtual spatial proteomics" from H&E that improved outcome prediction in non-small cell lung cancer cohorts. Multiple recent publications have demonstrated that standard histology slides contain predictive signals for molecular maps that used to require expensive specialized assays. The basic premise—that morphology and molecular state are related in ways machines can learn—keeps getting validated, even as questions about generalizability persist.
Second, regulators are waking up. The FDA released draft guidance on January 6, 2025, on using artificial intelligence to support regulatory decisions for drugs and biologics. (The agency noted it has fielded more than 500 submissions containing AI components since 2016, a tally that suggests the industry has been experimenting for a while.) The European Medicines Agency adopted a reflection paper on AI in medicinal product lifecycles last fall. The EU AI Act's general obligations kick in this August, with high-risk provisions phased through 2027—deadlines that will force vendors to formalize governance, document bias mitigation, and build transparency practices whether they like it or not.
Third, the economics are tightening. McKinsey has estimated that generative AI could add between 2.6% and 4.5% of annual revenues in pharma and medical products—somewhere in the range of $60 billion to $110 billion annually if you take industry benchmarks seriously. BCG reported late last year that about a quarter of biopharma respondents claimed at least 5% cost or revenue impacts from AI, though one suspects some of that is aspirational. Meanwhile, industry-sponsored clinical trial completions jumped 14.2% year-over-year in 2024, hitting just under 5,000 studies. More trials, same brutal failure rates—the pressure to improve efficiency is real.
If cross-modal imputation can retrospectively extract biomarkers from archived cohorts, or prospectively enrich trials without measuring every modality upfront, the ROI starts to look compelling. Assuming, of course, that the predictions hold up under scrutiny.
The Startup Approach
Strand AI exemplifies the current moment—young, technically ambitious, betting that the window is open now. The company emerged from Y Combinator's winter 2026 cohort, a two-person operation founded in 2025 by Yue Dai (who previously worked at Pathos AI, Tempus, Enable Medicine, and Microsoft Research) and Oded Falik (also ex-Enable Medicine). Their pitch: multimodal foundation models that transform incomplete patient profiles into complete datasets by computationally imputing what wasn't measured.
This March, Strand opened early access to something called POSTMAN, a model that predicts spatial proteomics from routine H&E slides. The company frames three use cases: retrospectively profiling archived slides without re-running expensive assays; triaging samples to identify which ones warrant the cost of actual spatial proteomics; and surfacing biomarkers that weren't directly measured in the original study design. Strand claims POSTMAN "beats state-of-the-art at a fraction of the cost," though independent benchmarks or peer-reviewed validation haven't appeared yet. They're recruiting design partners from pharma, biotech, and research institutions, dangling co-publication opportunities and roadmap influence as incentives.
In January, Strand also released a free dataset: imputed gene expression across 4,500 genes and 45 tissues for 538 samples from the 1000 Genomes Project, generated using a DNA-to-RNA model from CZI Biohub. The release demonstrated significantly faster inference on certain GPU configurations—a signal that the team is thinking about deployment economics alongside model performance. Dai's background suggests familiarity with both the data and compute challenges: he's claimed experience training models on one of the largest multimodal patient datasets in existence and leading ML infrastructure across more than 1,000 GPUs.
Whether a two-person startup can navigate the regulatory complexity and enterprise sales cycles that define pharma partnerships is another question entirely. But that's venture capital's bet to make.
The Broader Ecosystem

Strand isn't operating in isolation. BioMap published a 100-billion-parameter protein transformer in Nature Methods last year and claims roughly 210 billion parameters across its suite of multimodal life science models. EvolutionaryScale and the Arc Institute released genome modeling tools earlier this year. An analysis from Epoch AI noted that while biological AI models have proliferated since mid-2024, only a small proportion report formal risk assessments—a gap that regulatory frameworks will likely force closed.
Consolidation is already beginning. Quanterix announced plans to acquire Akoya Biosciences in January 2025, aiming to integrate liquid and tissue proteomic biomarkers across a combined installed base of around 2,300 instruments. Expect digital pathology platforms, spatial omics vendors, and computational biology startups to pursue similar tie-ups, vertically stacking data collection, imputation models, and clinical decision support. The wrinkle: whether these remain research-use tools or cross into regulated diagnostics, which triggers far stricter validation requirements and fundamentally different business models.
What Could Go Wrong
The technical critique is sharpening. A recent analysis comparing single-cell foundation models found that while they organize biological knowledge impressively, they exhibit what the authors called "limited regulatory logic" in their representations—meaning they recognize patterns but may not capture the causal mechanisms that actually govern cellular behavior. Cross-modal imputation faces identical scrutiny: does a model that predicts proteomics from histology genuinely understand biology, or is it exploiting dataset-specific correlations that won't generalize to different patient populations, tissue preparation protocols, or disease contexts?
That question matters enormously if you're a regulator being asked to trust these predictions for clinical trial decisions. If a model systematically mispredicts protein expression in certain demographic groups, or fails silently when presented with tissue from a cancer subtype it didn't see during training, the consequences aren't abstract. Trials get enriched with the wrong patients. Drugs fail for reasons that look like biology but are actually artifacts of bad imputation.
The FDA's draft guidance and the EU AI Act both emphasize transparency, reproducibility, and bias monitoring for exactly this reason. Companies treating cross-modal imputation as a black box will likely struggle. Those that build audit trails, characterize failure modes, and validate downstream utility on independent cohorts—preferably with prospective studies, not just retrospective analyses—will differentiate themselves. Documentation rigor, not just predictive accuracy, may determine who survives regulatory review.
Academic and commercial teams are starting to converge on hybrid approaches that layer in mechanistic constraints: pathway databases, causal graphs, physical modeling of molecular interactions. Pure pattern recognition, it turns out, has limits when the patterns matter clinically.
The Pragmatic Calculus

For biotech founders and clinical development executives, the decision tree is pragmatic, if not simple.
If your Phase II trial failed and you couldn't figure out why, and archived slides from that cohort are sitting in storage somewhere, cross-modal imputation offers a second look without the expense of re-consenting patients and shipping samples to specialty labs. If you're designing a precision oncology study and can't justify spending $2,000 per patient on spatial proteomics for 500 enrollees—a million-dollar line item before you've enrolled anyone—computationally triaging samples to identify high-value candidates for confirmatory testing might split the difference.
The technology won't replace wet-lab biology. Nobody serious thinks it will. But it might reshape which experiments happen when, and what can be learned from data already sitting on shelves.
That assumes, of course, that the models actually work as advertised when confronted with the messy reality of real-world clinical trials rather than carefully curated benchmark datasets. Assumes that regulators develop coherent frameworks for validating these tools without stifling innovation. Assumes that the companies building these systems can navigate the chasm between a clever demo and an enterprise product that pharma executives will stake their careers on.
Big assumptions. But then, drug development has always run on big bets with long odds. Adding computational ones to the mix might not fundamentally change the game—but it could, perhaps, nudge the percentages just enough to matter.
