The gene therapy industry is confronting an uncomfortable truth, though not everyone wants to say it out loud. Even as the field celebrates milestone approvals and billion-dollar market projections, safety signals keep mounting. Last November, the FDA slapped a boxed warning on Elevidys, Sarepta's Duchenne muscular dystrophy therapy, after investigating deaths from acute liver failure in non-ambulatory patients. Liver toxicity. Dorsal root ganglion damage. Off-target expression. These aren't edge cases anymore—they're friction slowing a therapy class that was supposed to revolutionize medicine.
Now a cluster of startups is applying artificial intelligence to one of gene therapy's most overlooked components: the regulatory DNA that controls when and where therapeutic genes actually turn on. It's unglamorous work. But as the industry learned with Elevidys, getting expression wrong can be fatal.
The Optimism (and the Cracks Beneath It)
Numbers tell part of the story. Fortune Business Insights projects the gene therapy market will swell from $3.43 billion in 2024 to $13.83 billion by 2032. MarketsandMarkets goes further—$36.55 billion by the same year. The optimism isn't manufactured. December 2023 delivered the first CRISPR therapy approvals for sickle cell disease. In January 2026, Encoded Therapeutics' CNS gene therapy ETX101 received FDA Breakthrough Therapy Designation. The FDA's Office of Therapeutic Products, established in 2023 specifically to handle the surge in cell and gene therapy submissions, is scaling review capacity to meet demand.
Yet beneath that momentum, safety concerns simmer. Beyond Elevidys, preclinical studies published in 2024 and 2025 documented dorsal root ganglion toxicity across species and cargo types—including RNAi, which theoretically should have been cleaner. The problem isn't limited to AAV vectors or specific diseases. When a therapeutic gene is delivered systemically, it often expresses in tissues where it shouldn't. Liver. DRG. Other off-target sites. The result: inflammation, or worse.
Traditional viral promoters, which drive gene expression broadly across cell types, amplify the issue. And regulators are tightening scrutiny. The FDA's long-term follow-up guidance, in force since 2020, already emphasized extended safety monitoring. The newer ICH S12 guidance on biodistribution studies, effective since 2023, demands more rigorous nonclinical data. The screws are turning.
The AI Pivot: Designing DNA That Knows Where to Stop
Regulatory DNA—promoters, enhancers, the non-coding "switches and dials" that govern gene expression—has historically been selected from a limited palette of viral or native human sequences. Over the past three years, that's changed. Artificial intelligence has moved from predicting how regulatory DNA behaves to designing it from scratch.
DeepMind's AlphaGenome, announced in June 2025 and widely covered by January 2026, can process 1 megabase-pair contexts and predict regulatory activity across multiple modalities. It outperformed Enformer, which itself represented a 2021 breakthrough for using transformers to predict gene expression over roughly 100-kilobase windows. Academic groups have published diffusion models and reinforcement learning frameworks for generating cell-type-specific promoters and enhancers. One 2025 RL system, Ctrl-DNA, demonstrated controllable tissue specificity while capturing biologically plausible transcription factor motifs.
But prediction is one thing. Clinical application? Another entirely. The challenge isn't just generating sequences that look active in silico—it's generating sequences that work in vivo, at the right levels, in the right cells, without triggering toxicity in the wrong ones.
The Startups (and the Old Guard) Making Bets

Origin Bio, a Y Combinator Winter 2026 company founded by Yash Rathod and Malhar Bhide, is building a platform around this problem. The startup is assembling what it describes as a large proprietary dataset—millions of experimentally validated regulatory elements across tissues and cell states. In October 2025, Origin released Axis, a multifunctional DNA model that both predicts regulatory function and generates new sequences. According to the company, Axis achieves 6.7% better transcription factor binding activity prediction than AlphaGenome on average, with 9× enrichment of TF motifs under high-affinity prompts. Axis is "promptable," meaning users can specify desired expression patterns and tissue targets.
Origin's scientific advisory board includes MIT's Manolis Kellis, Stanford gene therapy researcher Nicole Paulk, and the University of Illinois's Rashid Bashir—lending academic credibility to a young company. The pitch centers on safety: designing regulatory DNA that localizes gene expression to disease-relevant cells while minimizing off-tissue activity.
Encoded Therapeutics offers a more mature case study. The CNS-focused biotech has been engineering promoters, enhancers, and UTRs with machine learning assistance since its founding. Encoded's lead program, ETX101, uses a cell-selective regulatory element to control an engineered transcription factor that upregulates SCN1A in inhibitory neurons for Dravet syndrome. Preclinical data Encoded presented at the American Society of Gene and Cell Therapy conference in 2022 showed its regulatory elements could decrease liver expression while maintaining CNS activity—directly addressing common off-target liabilities. ETX101's Breakthrough Therapy Designation in January 2026 suggests the regulatory approach is passing clinical scrutiny. Encoded has also licensed its regulatory elements to Eli Lilly's Prevail Therapeutics, a sign that pharma sees commercial potential in the technology.
The space isn't crowded, but it's not empty either. AskBio acquired Synpromics in August 2019, integrating the UK company's synthetic promoter platform into its AAV gene therapy work. Chromatin Bioscience offers cell-selective synthetic promoters to cell and gene therapy developers. Meanwhile, adjacent players are tackling related problems: Dyno Therapeutics uses AI to design AAV capsids with improved tissue tropism, partnering with Roche, Sarepta, and Novartis. In May 2025, Dyno unveiled new capsids for CNS, eye, and muscle at ASGCT. The convergence is clear. Capsid design and regulatory element optimization are increasingly viewed as complementary routes to specificity.
A different set of companies is pursuing epigenetic editing, which modulates gene expression without cutting DNA. Chroma Medicine is advancing programs in HBV and PCSK9 silencing. Omega Therapeutics is developing "epigenomic controllers" targeting MYC and other oncogenes. Senti Bio's logic-gated CAR-NK platform uses gene circuits with Boolean logic—OR/NOT gates—to control cytokine expression in cell therapies. These approaches all grapple with the same underlying question: How do you turn genes on or off with surgical precision?
The Data Problem (and Why It Might Be a Moat)

The quality of AI models depends on the quality of training data. With regulatory DNA, that's not a given. Massively parallel reporter assays (MPRAs) and STARR-seq have become standard for generating sequence-to-function maps, but a 2025 meta-analysis published in Genome Biology highlighted troubling cross-assay variability and inconsistent enhancer calls across labs. Harmonization remains a problem.
Origin Bio's emphasis on building a proprietary, experimentally validated dataset may be strategic. If public data is noisy, private datasets with rigorous quality control become a competitive moat. The company hasn't disclosed the exact scale or tissue coverage of its dataset, though the "millions" figure and references to multiple cell states suggest ambition—perhaps more than one startup can realistically deliver alone.
Academic benchmarking efforts like DART-Eval from Stanford's Kundaje lab found mixed performance from DNA language models versus task-specific models, exposing evaluation gaps. The implication: raw model architecture matters less than training data quality and task-specific tuning. In other words, flashy AI won't save you if your dataset is garbage.
Regulatory Winds (Blowing in Favor of Precision)
The regulatory environment is evolving to match the science. The FDA's Office of Therapeutic Products is explicitly focused on improving cell and gene therapy review timelines and consistency. The European Medicines Agency's new guideline on investigational advanced therapy medicinal products, effective July 1, 2025, clarifies early-phase requirements and supports platform technologies. ICH S12, now adopted globally, standardizes biodistribution study expectations.
These frameworks create pressure to demonstrate tissue-specific expression and minimize off-target activity early in development. Synthetic regulatory elements, if they deliver on specificity claims, align with that pressure. Encoded's CEO has emphasized in public statements that regulatory element engineering mitigates toxicity by restricting expression away from liver and other vulnerable tissues—exactly the message regulators want to hear.
On the access side, the Centers for Medicare & Medicaid Services launched the Cell and Gene Therapy Access Model in January 2025, initially covering outcomes-based agreements for sickle cell gene therapies in 33 states plus DC and Puerto Rico as of July 2025. The model ties payment to outcomes, incentivizing safety and efficacy. Affordability remains gene therapy's existential risk. A July 2024 Nature paper argued that ten-fold cost reductions are possible through manufacturing innovation and alternative business models, but current pricing makes any safety signal potentially catastrophic for market access.
What Happens Next (and What Could Go Wrong)

The next 18 months will clarify whether AI-designed regulatory DNA moves from proof-of-concept to clinical standard. Encoded's ETX101 readouts will matter. If the trial succeeds, it validates regulatory element engineering in humans. Independent benchmarking of Origin's Axis against AlphaGenome on public datasets would help, too—so far, the performance claims come from the company's blog. Preprints or third-party evaluations are pending.
Several questions remain open, and they're not trivial. Will AI-designed synthetic elements outperform the best hand-screened promoters in non-human primates? How will regulators handle de novo synthetic regulatory DNA in Biologics License Applications and Investigational New Drug applications? What comparability data will they require? What off-target transcription thresholds will they tolerate? The FDA's Office of Therapeutic Products has signaled openness to innovation, but novel DNA sequences introduce variables that don't exist with natural promoters. Regulators will be cautious.
Dataset scale matters, but so does relevance. Most MPRA data comes from episomal contexts; integrated viral vectors behave differently. Tissue coverage is uneven. Cross-species translation remains tricky. Origin Bio and others will need to demonstrate that their models generalize to in vivo human contexts, not just cell lines grown under perfect lab conditions.
Perhaps the most interesting dynamic is the convergence of delivery and expression control. Dyno's AI-designed capsids improve tissue targeting. Origin's AI-designed promoters improve cell-type selectivity. Together, they represent a layered approach to precision: get the vector to the right organ, then turn the gene on only in the right cells. If both technologies mature in parallel, the combinatorial possibilities multiply. That's the optimistic scenario.
The Bottleneck Shifts
The gene therapy field spent years solving the delivery problem. It's now discovering that delivery alone isn't enough. Expression control—where, when, how much—may be the next bottleneck. AI is offering a way through, or at least a better tool than random trial-and-error screening.
Whether that path leads to safer therapies or just more elegant failures will depend on how well these models translate from silicon to tissue. The data so far is promising, but it's also early. And early data in biotech has a way of disappointing later. Still, the industry doesn't have many other options. Safety signals aren't going away, and neither is regulatory scrutiny. If AI-designed regulatory DNA can reduce off-target expression and toxicity, it won't just be a nice-to-have feature. It'll be table stakes for getting gene therapies through the clinic—and keeping them on the market once they're there.
