Somewhere in a biopharma office, a clinical team is staring at a spreadsheet that doesn't add up. They have histology slides from 300 patients. Spatial proteomics—the expensive molecular imaging that might unlock which patients respond to treatment—from only 50. The trial design assumed complete data. The regulatory clock is ticking. The options are grim: abandon the cohort, wait months for costly assays that may exhaust tissue samples, or present a fragmented dataset that weakens any conclusion.
Nine out of ten drug programs that enter clinical trials never make it to approval. That brutal arithmetic has held steady for decades. Phase II remains the graveyard. And while the reasons vary—efficacy failures, safety signals, enrollment problems—a quieter culprit shows up again and again: incomplete patient data.
A startup out of Y Combinator's Winter 2026 batch thinks it has found a shortcut.
Strand AI, founded by former Enable Medicine colleagues Yue Dai and Oded Falik, emerged in March with POSTMAN, a multimodal foundation model that claims to predict spatial proteomics directly from standard H&E histology slides—the workhorse stain pathologists have used for more than a century. The pitch is disarmingly simple: turn incomplete, multimodal patient profiles into enriched datasets without running additional assays. No new tissue cuts. No waiting. Just inference.
Whether it works—really works, in the messy reality of regulatory submissions and pivotal trials—remains an open question. But the fact that serious research institutions and pharma giants are even entertaining the idea suggests the ground has shifted.
When Biology Meets Missing Values
Missing data has always been a methodological headache in clinical research, the kind of problem that shows up in regulatory guidance documents with words like "principled handling" and "sensitivity analyses." The EMA's missing data guideline and the ICH E9 framework both emphasize pre-specification. Don't improvise. Don't fill gaps with wishful thinking.
What's changed, perhaps faster than the guidance can keep pace, is the emergence of foundation models capable of cross-modal prediction with enough performance to warrant serious attention from pharma R&D leaders who have spent careers treating computational biology with polite skepticism.
A January 2026 study in Nature Medicine offers a proof point, or at least a credible attempt at one. Researchers developed HEX, an AI model that predicts spatial proteomics from H&E slides, and evaluated it across six non-small cell lung cancer cohorts (2,298 patients) and 12 other cancer types (5,019 patients). Combining H&E with AI-derived virtual proteomics improved prognosis and immune checkpoint blockade response prediction. Notably, the study framed performance not just in pixel-level metrics—where models can look impressive in abstract—but in downstream clinical utility, a distinction that matters when considering regulatory acceptance.
Similar efforts have appeared in rapid succession, almost a race. Nature Communications published ROSIE in August 2025, a model for robust in silico multiplex immunofluorescence from H&E. Nature Machine Intelligence followed with HistoPlexer, a generative model to infer 11-channel protein multiplexes. The academic literature suggests the field is no longer speculative. The question is whether these models can transition from research validation to industrial deployment, where the stakes involve not just publications but patient outcomes and billion-dollar programs.
Strand's POSTMAN enters this landscape with notable technical claims. The model was trained on 20,000 paired slides spanning more than 180 proteins and uses a 1-billion-parameter architecture. The company says it beats state-of-the-art performance at lower cost, though those benchmarks have not yet appeared in peer-reviewed journals—a caveat worth noting. Dai, who previously logged time at Tempus, Pathos AI, and Microsoft Research, and Falik bring petabyte-scale multimodal spatial biology experience from Enable Medicine. Now they're opening early access through a design partner program, explicitly seeking pharma and biotech partners with H&E archives to validate POSTMAN in real-world workflows.
Whether those partners materialize, and whether the model performs as claimed in independent hands, will determine the trajectory.
The Economics of Virtual Tissue
Cost dynamics add urgency, though the numbers are slippery. Strand estimates high-plex multiplex immunofluorescence can run $5,000 to $10,000 per slide—a company claim that's difficult to verify independently, as public pricing data for such workflows is scarce and context-dependent. Academic core facilities often list far lower figures for simpler panels, but industrial-grade validation with dozens of markers and optimization across cohorts can push costs substantially higher, though such figures are rarely published. If a sponsor has archived H&E slides from a completed trial but lacks spatial proteomics, the choice is either to re-cut and re-run assays (expensive, time-consuming, sometimes impossible if tissue is exhausted) or to accept data gaps.
Virtual proteomics offers a third path. Whether it's a credible one depends on validation rigor and regulatory acceptance, neither of which is guaranteed.
The digital pathology market provides context, if not precision. Analyst estimates peg the sector at roughly $1.5 billion in recent years, with compound annual growth rates ranging from 8 to 19 percent through the mid-2030s, depending on methodology. Spatial omics, the more specialized niche POSTMAN addresses, has been estimated in the range of several hundred million dollars, with projections suggesting continued growth. Meanwhile, AI in drug discovery more broadly is expected to expand significantly over the coming years, according to various market research firms.
(The specifics vary wildly depending on who's counting and what they include, a reminder that market sizing in emergent categories often tells you more about optimism than reality.)
Engineering Notes From the Frontier

Beyond POSTMAN, Strand AI has been running infrastructure experiments that reveal its technical seriousness—or at least its founders' comfort with granular model-ops work. In January, the team executed a community run of CZI Biohub's VariantFormer, a 1.2-billion-parameter DNA-to-RNA model, on hundreds of samples from the "1000 Genomes expansion" dataset. They released imputed RNA-seq data and reported 37 times faster inference on NVIDIA A100s versus H100s after pipeline tuning.
More recently, the founders shared lessons from deploying on NVIDIA B200 GPUs, detailing throughput issues, transparent huge pages configuration, NCCL fallbacks, and FP8 precision gains. These are the kind of granular engineering notes that don't make for flashy headlines but signal the difference between a demo and a production system. They also suggest a team that's thinking about scale before the revenue arrives, which could be wisdom or hubris depending on how the next twelve months unfold.
The competitive landscape, meanwhile, includes both startups and well-funded incumbents moving fast. Tempus announced a collaboration with AstraZeneca and Pathos in 2025 to build what they called the "largest multimodal oncology foundation model," leveraging Tempus's extensive data library. Paige holds early FDA authorization for AI in digital pathology (prostate) and has continued publishing validation studies. Ibex gained U.S. 510(k) clearance for prostate AI in early 2025. Owkin has been pushing multimodal patient data networks and tools like BRCAura, an AI-driven screening for gBRCA mutations from slides. Bioptimus, which has raised significant funding, launched H-Optimus-0 and later announced M-Optimus, described as a "world model" spanning histology, spatial and bulk transcriptomics, and clinical data.
Strand is smaller and earlier, with all the attendant risks. But the design partner approach suggests pragmatism over hype, at least for now.
The Regulatory Fog
Regulatory acceptance remains the hardest question, perhaps the only one that truly matters.
In January 2026, the FDA and EMA jointly published "Guiding Principles of Good AI Practice in Drug Development," a high-level, non-binding document that signals direction but offers little operational detail. The principles emphasize governance, transparency, and validation across the product lifecycle—sensible enough, but vague when it comes to answering whether a pharma sponsor can rely on AI-imputed spatial proteomics in a pivotal trial.
More concrete frameworks are in development. The FDA has been signaling credibility assessment standards for AI models in submissions, noting that the agency has received hundreds of submissions with AI components in recent years. But for cross-modal imputation like POSTMAN, the regulatory path is genuinely unclear. Model-aided enrichment and exploratory biomarker discovery are promising use cases. Confirmatory use in pivotal trials, though? That will demand rigorous, pre-specified statistical handling and sensitivity analyses, the kind that regulatory reviewers can scrutinize without mercy.
The EMA held a workshop on external controls in late 2025, suggesting rising openness under defined conditions, but formal reflection papers are still in development. FDA commentary earlier this year emphasized "one pivotal trial with confirmatory evidence" as the default—a policy signal that may influence how sponsors design studies and what role AI-derived data can play.
The EU AI Act adds another layer, though its impact on life sciences AI remains somewhat opaque. General-purpose AI obligations took effect in mid-2025, with broader transparency rules rolling out through 2026. Foundation model providers entering the EU face documentation and technical disclosure requirements. Deployers integrating models into medical products may become "providers" under the Act depending on modifications and branding, a compliance scope that could affect both Strand and its clients in ways that won't be fully clear until enforcement begins.
The Bet

Industry appetite, at least, is clear. Recent reports from Deloitte, McKinsey, and BCG all emphasize AI's potential to accelerate trial design, optimize eligibility criteria, and improve patient selection. McKinsey has been particularly vocal. BCG reported in 2025 that companies investing in AI foundations—data, process, operating model—realize outsized gains, a finding that suggests sustained capital allocation toward these capabilities even as broader tech funding contracts.
If virtual spatial proteomics proves reliable enough to rescue incomplete cohorts or shave months from timelines, the economics become compelling. The alternative—running high-plex assays retrospectively or abandoning patients from analysis—is neither cheap nor fast. Strand's POSTMAN is still in early access, with performance claims unverified by independent peer review. The founders are betting that pharma partners will take a chance on a small team with a bold technical claim and a design partner pitch deck.
It's the kind of bet that either looks prescient in hindsight or gets quietly shelved when the model hits real-world data and buckles. But the research literature, the regulatory trajectory, and the growing roster of pharma-AI collaborations all point in the same direction: missing biology data is too expensive to stay missing for long.
Someone is going to figure out how to fill those gaps reliably. Whether it's Strand, or Tempus, or a well-funded lab that publishes a breakthrough next month, the question is no longer if but when. And for the clinical teams staring at those incomplete spreadsheets, when can't come soon enough.
