The lab space is unremarkable, tucked into one of those generic San Francisco buildings where half the startups seem to land. But the claim coming out of it is anything but modest. Atlas Discovery—three people, fresh backing from Y Combinator—says its AI model can predict which ulcerative colitis patients will respond to a biologic drug before they ever take it. Using only baseline biopsy data from a past clinical trial, the team argues it could have shrunk enrollment from 640 patients down to 182 while keeping the math intact.
That's the sort of efficiency gain that could shave tens of millions off a Phase 3 trial budget, if it actually works.
It's also the kind of promise that's suddenly everywhere in biotechnology circles. "Virtual cells"—computational models designed to simulate how living cells react to drugs, genetic tweaks, or disease—are racing from academic curiosities into the pitch decks of startups and the strategic plans of pharmaceutical giants. The vision is seductive: predict what a compound will do before you synthesize it. Identify which patients will benefit before you enroll them. Skip some of the expensive, time-consuming guesswork that has defined drug development for decades.
The reality, at least for now, is more incremental. But the stakes are high enough that billions of dollars are flowing into the effort anyway.
What Changed, and When
The idea of simulating cellular biology on a computer isn't new. Mechanistic models of metabolic pathways and signaling networks have been around for years, grinding away in academic labs with mixed results. What's different now is the collision of three things: single-cell genomics that can profile millions of individual cells at once, high-throughput CRISPR screens that systematically knock down thousands of genes, and the AI foundation models that can digest all that data and spot patterns humans never would.
By the middle of this decade, researchers could routinely generate datasets of staggering size—millions of cellular snapshots capturing which genes are switched on, how cells respond to perturbations, even where they sit in tissue. The bottleneck shifted from data collection to making sense of it all.
The field took shape publicly when Arc Institute launched its Virtual Cell Challenge on June 26, 2025, inviting teams worldwide to build models capable of predicting how cells would respond to genetic changes they'd never seen before. NVIDIA, 10x Genomics, and Ultima Genomics put money behind it. The results, published late that year, were instructive. Models could capture broad trends—yes, this gene drives proliferation, that pathway lights up under stress—but they couldn't yet replace actual experiments. "The virtual cell instead of experiment is not yet realized," Arc's postmortem noted, calling for better benchmarks.
That assessment hasn't slowed anyone down.
On January 13 this year, Illumina unveiled what it called the Billion Cell Atlas, a dataset profiling genetic perturbations across roughly 200 cell lines. Two months later, Xaira Therapeutics announced X-Cell, a 4.9 billion-parameter model trained on more than 25 million perturbed single-cell transcriptomes. The same day—March 17—PerturbAI surfaced from stealth with a 7.7 million-cell dataset built entirely in mouse brains, mapping CRISPR knockdowns in living tissue rather than plastic dishes. The company worked with NVIDIA, the Allen Institute, and 10x Genomics to pull it off.
It felt, suddenly, like a land rush.
Why Now? (And Why So Much Money?)

Three forces converged to make this moment possible, or at least fundable.
Data. Sequencing costs collapsed, and technologies like single-cell RNA sequencing matured from research tools into production platforms. Spatial transcriptomics—capturing not just which genes are active but where in tissue—added another dimension. Grand View Research estimated the spatial omics market at $711 million in 2024, heading toward $1.7 billion by 2030.
That's generating raw material at unprecedented scale. In December, Tahoe Therapeutics contracted with Parse Biosciences to produce 300 million single-cell profiles through a platform called GigaLab. Illumina's atlas, Xaira's expanding datasets, and the Human Cell Atlas (which released more than 45 papers in a November 2024 collection) are all feeding what people in the field like to call the "data flywheel." More data trains better models, which attract more funding, which generates more data.
AI architecture. Foundation models pretrained on genomic-scale datasets—scGPT, trained on 33 million cells; scFoundation; Geneformer—showed they could learn representations of cellular biology that transferred across different tasks. But evaluations published in Nature Methods in 2024 and Genome Biology the following year highlighted shaky zero-shot performance and a need for larger, cleaner training sets. The 2026 wave of models, including Xaira's X-Cell and Atlas Discovery's ExpressionVAE (which uses discrete tokenization, not continuous embeddings), represents a second generation trying to close those gaps.
Money. Pharma is paying attention because the economics of drug development are brutal. Bringing a new medicine to market costs upward of $2 billion and often takes a decade. The AI-in-drug-discovery market, depending on which forecast you believe, is projected to grow from somewhere around $3 billion to $4 billion this year to anywhere between $14 billion and $44 billion by the mid-2030s. Those wide ranges reflect both optimism and uncertainty, but the trend is clear.
On May 20, 2026, Incyte paid Genesis Therapeutics $80 million upfront to expand a partnership built around AI-driven discovery. In January 2024, Isomorphic Labs—Demis Hassabis's drug discovery spinout from DeepMind—signed deals with Eli Lilly and Novartis worth nearly $3 billion in potential milestone payments. Those numbers get pharma executives' attention.
The Players

Xaira Therapeutics launched on April 24, 2024, with more than $1 billion from ARCH Venture Partners and Foresite Capital. That's a staggering sum for a startup without a product, but the pitch was ambitious: build a virtual cell capable of understanding biology at a causal level, not just correlation. "What does it take to bring a disease-state cell back to healthy?" was how the founders framed it.
On June 17, 2025, Xaira open-sourced X-Atlas/Orion, then the largest public genome-wide perturb-seq dataset at around 8 million cells. Nine months later came X-Atlas/Pisces and the X-Cell model—4.9 billion parameters trained to predict cellular responses to genetic knockdowns, chemical compounds, even different doses. Xaira positions this as the largest "causal perturbation model" built to date, though the team hasn't yet published peer-reviewed benchmarks. The field is moving fast enough that open-sourcing datasets and models has become a competitive strategy in itself.
Atlas Discovery is playing a different angle. Founded by Sanjukta Bhattacharya, Christian Gensbigler, and Shaamil Karim, the three-person team came through Y Combinator's Summer 2026 cohort with a focus on patient drug response. Their ExpressionVAE model uses discrete latent representations—think tokens, like in language models, rather than continuous vectors—and they claim improvements of three to twentyfold over continuous-latent baselines when tested on public datasets.
More provocatively, Atlas ran a retrospective analysis on the UNIFI trial, which tested ustekinumab for ulcerative colitis. Using baseline biopsy gene expression data, they report an AUROC of 0.76 in predicting which patients would respond. They modeled a scenario where enriching for likely responders could have reduced enrollment from 640 patients to 182 while maintaining statistical power. That's an internal claim, not independently validated. If it holds up and survives regulatory scrutiny, though, it's the kind of trial redesign that could reshape Phase 3 economics. A big "if."
Chan Zuckerberg Biohub announced a $500 million Virtual Biology Initiative in May 2026. The stated aim: build datasets and predictive cellular models comprehensive enough that AI can simulate biological processes in silico. It's a moonshot, and it comes with the baggage Mark Zuckerberg always brings—genetic data as the new frontier, and a track record with user data that invites skepticism. Still, half a billion dollars signals how seriously billionaire-backed philanthropy is taking this bet.
PerturbAI went public on March 17, 2026, with something the field had been missing: in vivo data. Most virtual cell work uses cultured cell lines, convenient but artificial. PerturbAI profiled 7.7 million cells from mouse brains, mapping how CRISPR-mediated gene knockdowns affect cells in their native tissue environment. The collaboration with NVIDIA, the Allen Institute, and 10x Genomics underscores the computational and experimental firepower required. They put the data on bioRxiv and Hugging Face, part of a broader trend toward open datasets that fuel competitive model-building.
Other names in the mix include Cellular Intelligence, which rebranded from Somite in January and touts a "universal virtual cell-signaling model," and Stratica Bio, which frames its work as "biological intelligence infrastructure" for virtual tissue models. Recursion Pharmaceuticals, which partnered with NVIDIA back in 2023, continues training foundation models on multi-petabyte phenomics data plus oncology real-world data from Tempus.
What Happens Next

The immediate challenge is validation, and it's not a small one. Arc Institute's 2025 challenge made clear that predicting perturbation responses in held-out contexts remains genuinely difficult. Models can capture trends—this gene influences proliferation, that pathway responds to inflammation—but quantitative accuracy and generalization to disease-relevant biology lag behind the hype. The 2026 iteration of the challenge, expected to open this summer, will test whether newer architectures and larger datasets actually close the gap.
Regulation looms larger than many founders probably want to think about. If virtual cell outputs start informing clinical trial design or patient selection, they'll need validation pathways—likely through the FDA's Biomarker Qualification Program or something similar. The FDA issued draft guidance on AI for regulatory decision-making in drug development in early 2025 and has been coordinating with the European Medicines Agency on "good AI practice" principles.
The EU AI Act, finalized in June, sets transparency requirements starting this December, with high-risk AI systems (which could include those used for medical decisions) facing compliance deadlines ranging from late next year to mid-2028 depending on implementation. For founders, that means planning for documentation, lifecycle controls, and early agency engagement if they want their models to support regulatory submissions. It's not the sexy part of the pitch, but it matters.
The market opportunity is real, if conditional. Pharma companies are expanding AI partnerships—Incyte's $80 million to Genesis is one signal among many—but they're also watching for clinical proof points. The first virtual cell-guided drug program to enter trials with a validated patient enrichment strategy will set the template, and everyone knows it. Until then, the field remains a peculiar mix of ambitious science and speculative economics.
Data will continue to be both the bottleneck and the enabler. Multiple efforts to produce datasets an order of magnitude larger—Illumina's Billion Cell Atlas, Xaira's expanding collections, Tahoe's 300 million-cell project—point to the next year or two as a period of rapid "data flywheel" build-out. If the hypothesis holds that scale unlocks emergent capabilities, as it has in language models, then the models trained in 2027 might finally deliver on some of the current promises.
Perhaps the most revealing indicator is where the money flows. When a Y Combinator-backed three-person startup can credibly pitch foundation models of drug response, and when pharmaceutical companies pay eight figures to access AI platforms still in development, the industry believes something is about to shift. Whether virtual cells actually replace the messy, expensive, failure-prone process of traditional drug discovery remains very much an open question.
But the race is on. And judging by the dollars being committed, the bets are already placed.
