Zolgensma works. So does Casgevy. Ask the families of children with spinal muscular atrophy or patients freed from sickle cell disease, and the answer is unequivocal: gene therapy has arrived. But ask the scientists who designed these treatments, and you'll hear a different story—one about brute force and acceptable compromise.
The problem isn't efficacy. It's precision. Or rather, the lack of it.
Current gene therapies often require high doses to reach their target cells, with the therapeutic gene expressing itself indiscriminately along the way—in liver tissue that doesn't need it, in heart cells where it might cause trouble, in immune cells that might mount an attack. The FDA approves these treatments because they work well enough to change lives. But "well enough" leaves considerable room for better.
Enter an unlikely solution: the same generative AI technology that writes sonnets and conjures images is now learning to write DNA. Not just any DNA—the regulatory sequences that function as genetic dimmer switches, controlling when genes turn on, where they express, and how loudly they speak.
Whether this computational approach can solve gene therapy's control problem remains an open question. But the early evidence suggests we're past the point of asking "if" and moving rapidly toward "how soon."
When More Expression Isn't Better
The toxicity reports tell a story that approval headlines don't. AAV vectors—the delivery vehicles undergirding many approved gene therapies—can trigger hepatotoxicity at higher systemic doses. Myocarditis. Thrombotic microangiopathy. Often these reactions trace back to immune responses against proteins expressed in the wrong places or at the wrong levels.
Consider the challenge facing CNS therapies. You need transgene expression in specific neuron subtypes—inhibitory interneurons, perhaps—but emphatically not in excitatory neurons. And definitely not in liver or cardiac tissue. CAR-T and CAR-NK cancer therapies face a similar dilemma: kill the tumor without obliterating healthy tissues that happen to share the same surface antigen.
The traditional approach has been part curation, part educated guess. Borrow a promoter from nature, tweak an enhancer, run experiments, repeat. Encoded Therapeutics built its ETX101 program for Dravet syndrome around custom regulatory elements designed to upregulate the SCN1A gene selectively in inhibitory neurons. The FDA granted it both Regenerative Medicine Advanced Therapy and Breakthrough Therapy designations—the company is now running a global Phase 1/2 trial. AskBio acquired Synpromics in 2019 and launched its PromPT platform for tissue-selective promoters. Chromatin Bioscience markets machine learning tools to accelerate the hunt for better promoters.
Yet until very recently, these were fundamentally engineering exercises. Smart engineering, certainly, but still iterative refinement rather than true design from first principles.
The Models That Learned to Compose
Something shifted in 2025. Not with fanfare exactly, but with the kind of peer-reviewed demonstration that makes scientists lean forward in their chairs.
A team published the DNA-Diffusion framework in Nature Genetics, showing they could generate 200-base-pair synthetic regulatory elements with cell-type-specific activity. They didn't just predict which sequences might work—they designed them computationally, then validated 5,850 AI-generated elements using STARR-seq assays. The elements actually functioned. They modulated endogenous genes like AXIN2 in living cells. A Nature Genetics News & Views commentary called it "a step toward precision control of gene expression"—measured language from a field that's learned to temper enthusiasm, but notable nonetheless.
DNA-Diffusion wasn't operating in isolation. Reviews published over the past two years catalog a small ecosystem of approaches: reinforcement-learning frameworks like Ctrl-DNA and Regulatory DNA RL, flow-matching methods for discrete biological sequences, foundation models trained on genomic datasets that would have seemed absurdly large a decade ago. The architecture is converging. Diffusion for generative design. Foundation models for long-range sequence context. Reinforcement learning for constraint-aware optimization.
The training data evolved too, in ways that matter. Early models ingested curated transcription factor binding motifs—structured data, clean but limited. Now they consume functional readouts from MPRA and STARR-seq libraries that profile tens of thousands of sequences, scMPRA assays resolving promoter activity across individual cell types, chromatin accessibility maps. A 2025 cross-assay study in Genome Biology highlighted persistent technical variability between MPRA and STARR-seq, which pushed researchers toward standardized pipelines. That standardization isn't academic hair-splitting. Industrial-grade model training demands reliable, reproducible functional data.
Follow the Money

Wall Street and Basel are paying attention. Twenty of the 30 largest biopharmaceutical companies now have active cell and gene therapy programs. Market projections—directional at best—suggest the CGT space could hit $117 billion to $128 billion by 2034, expanding at roughly 18–19% annually. Manufacturing capacity is projected to grow faster than 25% per year, which tells you where the bottlenecks currently sit.
The deal flow reflects a certain urgency around delivery and expression control. Voyager Therapeutics licensed its TRACER capsids to Novartis for Huntington's, spinal muscular atrophy, and additional CNS indications. Capsida Biotherapeutics expanded its collaboration with AbbVie after demonstrating IV-delivered capsids that achieve neuronal expression while largely avoiding liver and dorsal root ganglia in primates. Dyno Therapeutics partnered with Roche and Spark around its CapsidMap platform for CNS and liver-directed therapies.
These deals pair AI-designed delivery systems with the emerging promise of AI-designed expression control. Two pieces of the same puzzle.
On the regulatory element front, Senti Bio is advancing logic-gated CAR-NK cells—SENTI-202—that use OR/NOT genetic circuits to improve tumor selectivity while sparing hematopoietic stem cells. It's expression control layered atop cell engineering, with promoters and sensor circuits working together to define therapeutic activation. Early updates suggest the logic-gating concept has some traction, though definitive data remains forthcoming.
Then there's Origin Bio, a California startup registered in January 2025 that's making expansive claims. The company, which says it's backed by Y Combinator and advised by researchers including Manolis Kellis, Nicole Paulk, and Rashid Bashir, launched "Axis"—an AI model it claims can both generate regulatory DNA and predict its function. Origin positions Axis as outperforming "AlphaGenome" on binding-activity benchmarks and hints at building the largest proprietary dataset for regulatory DNA design.
Independent peer-reviewed validation isn't yet available, which means the appropriate response is watchful skepticism. But the company's positioning is clear: it intends to program gene expression for safer cell therapies. Whether Axis actually delivers remains very much to be seen.
The Data Headache
The gap separating peer-reviewed models like DNA-Diffusion from clinical readiness is partly—perhaps primarily—a data problem.
MPRA and STARR-seq assays measure regulatory activity, but cross-lab variability remains stubbornly high. The 2025 Genome Biology harmonization study found that standardized computational pipelines improved enhancer call concordance, yet wet-lab protocol differences still introduce considerable noise. scMPRA expands the dataset by profiling at single-cell resolution, which is valuable. It also multiplies technical complexity, which is... less valuable.
For companies planning to submit INDs—investigational new drug applications—that variability poses real liability. The FDA finalized guidance on human gene therapy products incorporating genome editing in January 2024, with emphasis on product design, manufacturing, testing, and nonclinical safety. The EMA adopted an investigational ATMP guideline effective July 1, 2025. Regulators want reproducible, well-characterized products, which is perfectly reasonable. If your AI model's training data comes from assays with high lab-to-lab variance, how confident should you be that the designed sequence will perform consistently in a clinical vector?
Probably not very.
The field is gravitating toward what some researchers call "programming stacks"—multiple control mechanisms layered together. Cell-type-specific promoters and enhancers, yes. MiRNA de-targeting that suppresses expression in unwanted tissues (inserting miR-122 binding sites to block liver expression, for instance). Sensor circuits responsive to disease biomarkers. Each layer adds specificity. Each layer also adds complexity and potential failure modes. The challenge becomes designing these stacks computationally with sufficient confidence to bet a clinical program on them.
There's also a nascent dual-use concern. A 2025 preprint warned that DNA foundation models might be vulnerable to adversarial attacks that generate harmful sequences, calling for safety alignment and traceability. It's an early worry, perhaps premature. But given how quickly AI capabilities are advancing, worth tracking.
What Happens Next

Short term? Expect more peer-reviewed demonstrations of AI-designed regulatory elements with experimentally validated cell-type specificity. Expansion of scMPRA and similar assays will feed improved training data, assuming standardization efforts gain traction. Clinical readouts from precision gene-regulation programs—Encoded's ETX101 data in Dravet syndrome, for example—will begin showing whether computationally designed regulatory control actually improves therapeutic index in humans.
That last point matters more than all the computational benchmarks combined.
Mid term, the bottleneck likely shifts to reimbursement. The U.S. CMS launched the Cell and Gene Therapy Access Model to negotiate outcomes-based agreements for sickle cell disease gene therapies. As of July 15, 2025, 33 states plus DC and Puerto Rico had joined. If the model works—if it creates a viable framework for reimbursing $2 million-plus treatments—it could fundamentally reshape uptake dynamics and create real-world volume signals for developers.
Gene therapy approvals are accelerating. Luxturna in 2017, Zolgensma in 2019, Hemgenix in 2022, Casgevy and Lyfgenia in 2023, Lenmeldy this past March. But commercial success depends on payer willingness to cover these treatments at scale, and that negotiation is still very much in progress.
Competitive dynamics are shifting too. Pharma partnerships around AI-designed capsids suggest major players will pay for differentiated delivery technology. If regulatory DNA design proves out clinically, expect similar deal structures: platform companies licensing expression-control stacks to developers who need them for specific indications. The open question is whether any single platform—Origin's Axis, AskBio's PromPT, Chromatin's ML-aided approach—can build enough proprietary data and validation to become a de facto standard. Or whether the field fragments into indication-specific custom solutions.
Longer term, the most intriguing possibility is convergence. AI-designed regulatory elements combined with AI-designed capsids and gene circuits. A fully programmable gene delivery system where the capsid determines which cells get transduced, the promoter and enhancers control expression magnitude and timing, miRNA sites suppress unwanted tissues, and sensor circuits respond to disease state.
It's a lot of moving parts. Also the logical endpoint of precision medicine applied to gene therapy.
A Moment Between Questions

For now, the field sits at an inflection point. The fundamental question has evolved from "Can AI design functional regulatory DNA?" to "How do we industrialize and validate it for clinical use?" Early commercial players are placing their bets. Data infrastructure is maturing, if unevenly. Regulators are writing the rules in real time.
Gene therapy has wrestled with the control problem for decades. AI probably won't solve it entirely—biology has a way of resisting complete solutions. But it's beginning to offer something better than trial and error, something closer to intentional design.
Whether that's good enough to transform gene therapy from an effective but crude instrument into a precision tool remains to be written. The next few years should provide an answer. Perhaps more than one.
