The news came quietly, as these things often do. Pfizer discontinued its hemophilia B gene therapy earlier this year, not because the science failed—the treatment had cleared FDA approval, after all—but because doctors and patients weren't using it. A one-time genetic fix for a chronic bleeding disorder, shelved. Industry watchers took note, though perhaps not for the reasons Pfizer might have hoped.
The real story wasn't about market dynamics or pricing strategy, though those mattered. It was about a question that's haunted gene therapy since its earliest days: what happens when the genetic intervention you've engineered works too well, or in places you didn't intend?
Hepatotoxicity. Cytokine storms. Immune attacks that can overwhelm a patient's system. The field's safety ledger reads like a cautionary tale, and as gene therapies inch beyond ultra-rare diseases toward larger patient populations, the room for error narrows considerably.
Which brings us to a problem that sounds almost mundane compared to the headline-grabbing work of CRISPR editing or viral capsid design: expression control. How do you ensure a therapeutic gene turns on at the right intensity, in the right cells, and nowhere else?
It's less glamorous than the genetic manipulation itself. It might also be more important.
The Mechanics of Getting It Wrong
Here's the logic, stripped down. AAV gene therapies—those using adeno-associated viral vectors to ferry genetic cargo into cells—need to deliver their payload without triggering an immune mutiny. When that payload expresses at supraphysiologic levels, meaning far above what the body naturally produces, it presents what immunologists call an antigenic burden. The immune system notices. Sometimes it reacts with devastating force.
CAR-T therapies face a parallel challenge. The promoters driving CAR expression can fire too aggressively, leading to what's clinically termed "tonic signaling"—T-cells that exhaust themselves or, worse, trigger cytokine release syndrome and immune effector cell-associated neurotoxicity syndrome. Both conditions can be life-threatening.
A 2024 editorial in EMBO Molecular Medicine captured the field's emerging consensus in five words: "more is not always better." Developers have understood for years that dialing down expression or confining it to target tissues could reduce adverse events. The trick has been finding the right regulatory DNA—the promoters and enhancers that function as dials and switches for gene activity—without running thousands of trial-and-error experiments.
Enter machine learning, stage left.
Or more accurately, a cohort of AI models trained to decode the regulatory instructions embedded in our genomes.
Prediction at Scale
The October 2024 publication in Nature felt like a turning point, at least to those paying attention. Researchers from the Broad Institute, Yale, and Jackson Laboratory demonstrated that synthetic cis-regulatory elements—the DNA sequences controlling when and where genes activate—could be machine-designed and validated in living cells. They called their predictor "Malinois," and it achieved Pearson correlations between 0.79 and 0.91 across multiple cell types. More importantly, the sequences it helped design actually worked in vivo, steering gene expression to specific cell types with measurable precision.
Then, six months later, DeepMind unveiled AlphaGenome. The multimodal model, which received a major update in January following its Nature publication, could predict regulatory variant effects across a million base pairs of context. It drew on datasets from ENCODE, GTEx, and FANTOM5—the field's major reference libraries—to interpret how non-coding DNA influences gene activity.
Impressive work. Prediction at genuine scale.
But here's the thing: prediction alone doesn't build therapies. You need generation—sequences designed to do something specific. Turn on in liver hepatocytes but stay silent in cardiac muscle. Ramp up when a tumor microenvironment signal appears. Shut down after a therapeutic window closes.
That's a design problem, and it's where a new wave of startups is placing bets.
Four People and a Genomic Code-Breaking Ambition

Origin emerged from Y Combinator's Winter 2026 batch with a team of four and an outsize ambition. Founded by Yash Rathod and Malhar Bhide—both computer science graduates from the University of Illinois at Urbana-Champaign with machine learning research backgrounds—the San Francisco startup announced its "Axis" platform in October. Axis, they claim, is the first AI model that both generates regulatory DNA and predicts its function.
Whether it's truly the first is debatable—competitive claims in this space move quickly—but the dual capability matters. Prediction tells you what a sequence might do. Generation lets you create sequences that do what you need.
The company reports a 6.7% average improvement over AlphaGenome on regulatory element-activity prediction benchmarks. They claim up to 9-fold enrichment of transcription factor binding motifs when prompted for high-affinity sequences. Those numbers come from internal evaluations, cross-checked using Malinois as an independent validator. So: provisional. Benchmarks in an emerging field where wet-lab validation remains the ultimate judge.
But the roster of advisors suggests they're serious. Dr. Manolis Kellis from MIT and the Broad Institute. Dr. Nicole Paulk, an AAV specialist affiliated with Siren Bio. Dr. Rashid Bashir, Dean of Engineering at UIUC. It's the kind of advisory board that signals credibility in both computational biology and translational gene therapy—if you can keep them engaged.
The Data Moat
Origin's pitch hinges on something less sexy than algorithmic innovation: a dataset play. They're building what they describe as the largest proprietary library of experimentally validated regulatory elements—millions of sequences tested across multiple cell and tissue types.
If that materializes, it could become a genuine competitive moat. Machine learning models are only as good as their training data, and regulatory DNA is notoriously context-dependent. A promoter that drives robust expression in neurons might barely flicker in hepatocytes. Sometimes the empirical results defy conventional wisdom entirely. A 2025 study in Molecular Therapy found that GFAP, a promoter classically associated with astrocytes, actually outperformed the liver-standard LP1 promoter in hepatocytes—in both mouse and humanized models.
Findings like that argue for data-driven, context-specific design rather than leaning on canonical assumptions about what promoters "should" do. They also underscore the value of scale. A proprietary dataset spanning dozens of cell types and thousands of regulatory variants could provide a significant edge over competitors relying solely on public repositories.
Origin isn't alone in this pursuit. MeiraGTx, a clinical-stage gene therapy company, developed CLARA, a convolutional neural network for in silico promoter prediction, integrated with riboswitch controls for multi-dimensional regulation. Asimov launched its AAV Edge suite in September 2024, combining AI models with genetic tools—including tissue-specific promoter libraries—for end-to-end AAV design and manufacturing. Senti Biosciences is engineering synthetic gene circuits with logic gates (AND, NOT) to spare healthy tissue in CAR-NK and CAR-T therapies.
Academic efforts are accelerating too. An HPI–MIT collaboration is reportedly using deep learning to design cancer-specific regulatory DNA. Preprints like "Ctrl-DNA," which appeared in mid-2025, describe constrained reinforcement learning methods for generating cell-type-specific promoters with improved specificity.
It's crowded territory, in other words. And crowded usually means someone's onto something valuable.
Big Deals and Renewed Optimism

The broader gene therapy sector is experiencing something of a recalibration. After years of hype punctuated by high-profile failures, dealmaking activity picked up notably last year. ARM CEO Tim Hunt noted in a January interview that three acquisitions exceeding $1 billion closed in 2025: AbbVie's $1.5 billion purchase of Capstan Therapeutics (in vivo CAR-T via targeted lipid nanoparticles), Bristol Myers Squibb's acquisition of Orbital Therapeutics (RNA platform with AI-driven design), and Eli Lilly's roughly $1 billion deal for Verve Therapeutics (gene-editing cardiovascular therapies).
Around $4 billion flowed into public markets, with approximately 15 private financings topping $70 million each. Not stratospheric numbers by biotech standards, but directionally encouraging after a challenging period.
Regulatory infrastructure is evolving in tandem. The FDA has convened town halls on gene therapy manufacturing and chemistry, manufacturing, and controls readiness. It issued draft guidance on individualized ultra-rare therapies recently and published frameworks for innovative trial designs in small populations. The CMS Cell and Gene Therapy Access Model, which launched in early 2025 with a focus on sickle cell disease treatments, had expanded to 33 states plus D.C. and Puerto Rico by mid-year, covering the substantial majority of Medicaid beneficiaries with the condition.
Yet uptake remains frustratingly uneven. Pfizer's BEQVEZ discontinuation wasn't an isolated incident—it reflected persistent challenges in payer negotiations, center enablement, and the chronic shortage of long-term outcome data. The American Society of Gene & Cell Therapy's quarterly landscape reports tracked over 3,200 global trials as of Q3 2025, but converting clinical promise to commercial reality requires solving for manufacturing scale, safety, and cost.
Market projections vary widely. Depending on which analyst you consult, the sector could reach anywhere from $117 billion to over $200 billion by the mid-2030s. The directional trend, however, is consistent: double-digit compound annual growth rates across therapeutics, tools, and manufacturing segments. Capital is being deployed. Infrastructure is expanding.
Whether that translates to medicines patients can actually access remains the open question.
Programming Biology, One Promoter at a Time

The case for AI-designed regulatory DNA rests on a straightforward premise: if you can program where, when, and how much a therapeutic gene expresses, you can separate efficacy from toxicity. That's not purely theoretical. Muscle-biased promoters like MHCK7 have concentrated micro-dystrophin expression in Duchenne muscular dystrophy programs, potentially reducing off-tissue antigen burden. CAR-T developers are exploring promoter modulation to dampen tonic signaling without sacrificing tumor control.
But—and there's always a but—the tools remain early. Origin's Axis platform has benchmarks, not clinical data. Malinois and AlphaGenome have demonstrated predictive power in controlled settings, but generative design at therapeutic scale hasn't been proven in patients. The open question is whether startups like Origin can build wet-lab datasets large and diverse enough to train models that reliably produce regulatory elements that work not just in silico, but in the messy, unpredictable environment of human biology.
The broader industry shift suggests something fundamental is changing. In vivo CAR-T. RNA-based delivery. AI-enabled design. The architecture of gene therapy is being reconstructed, piece by piece. Regulatory DNA is one component. Capsid engineering, payload design, and manufacturing process optimization are others. All of them need to converge before the field can credibly move beyond rare diseases into the chronic, high-prevalence conditions that define commercial viability.
The Unsexy Work Ahead
For now, the early movers are placing bets on data, models, and a proposition that would have seemed fanciful a decade ago: that the genome's regulatory code—long dismissed as "junk DNA"—holds keys to making powerful therapies safer and more controllable.
Whether those bets pay off will depend less on algorithmic cleverness than on wet-lab execution, clinical validation, and the unglamorous work of turning computational predictions into medicines. Manufacturing at scale. Navigating regulatory pathways. Convincing payers that outcomes justify price tags.
It's the kind of work that doesn't generate splashy headlines. But then, neither does the careful calibration of gene expression levels. Until someone gets it right, and a therapy that might have triggered immune catastrophe instead works quietly, safely, in exactly the cells where it's needed.
That's when the field will know these tools have moved from promising to proven. Until then, it's data collection, model training, and validation experiments. One promoter sequence at a time.
