On April 18, 2024, the FDA did something unusual. It plastered a boxed warning—the agency's starkest alarm—across every CAR-T therapy targeting CD19 and BCMA on the U.S. market. The reason: T-cell malignancies cropping up in patients who'd received what was supposed to be a one-time cure.
A year later, in 2025, a boy with Duchenne muscular dystrophy died from acute liver failure after receiving Sarepta's Elevidys, a gene therapy that had offered hope where almost none existed. These weren't isolated events. They were symptoms of a design flaw that's been hiding in plain sight.
Gene therapy has a math problem.
More than 3,200 clinical trials are running worldwide. Industry analysts project the market will hit somewhere between $117 billion and $232 billion by 2034—a range so wide it almost betrays the uncertainty lurking beneath the hype. The science works, at least in principle. You can deliver functional genes to correct devastating genetic diseases. But here's what the early trials keep revealing: the body's reaction to high-dose viral vectors and uncontrolled gene expression remains unpredictable enough to kill people.
The therapeutic gene—the payload everyone focuses on—is only half the equation. The regulatory DNA that controls when, where, and how much that gene gets expressed? That matters just as much, maybe more. Get it wrong and you flood healthy tissue with toxic protein levels. You trigger immune storms. You inadvertently wake up oncogenes that should have stayed dormant.
The field has tried hand-engineering synthetic promoters, embedding microRNA target sites, building kill switches. These approaches work, to a point. But they're slow, constrained by human intuition, and they often miss the staggering complexity of how cells actually decide which genes to turn on.
When More Therapy Means More Danger
The toxicity data is piling up. AAV vectors—adeno-associated viruses, the workhorse of modern gene therapy—trigger hepatotoxicity, immune activation, and thrombotic microangiopathy when doses climb above 1×10^14 viral genomes per kilogram. That's not a fringe observation; it's a class effect.
A 2025 meta-analysis in iScience found that intravitreal AAV delivery carries higher uveitis rates than subretinal administration, forcing clinicians into prophylactic immunosuppression protocols that carry their own risks. The FDA's April 2024 boxed warning now mandates lifelong monitoring for anyone treated with autologous CAR-T products targeting CD19 or BCMA. Lifelong.
These aren't edge cases. As of Q1 2025, the Alliance for Regenerative Medicine and Citeline tallied 2,154 gene therapies in development, with 1,070 targeting rare diseases—heavily weighted toward oncology. Each one carries this expression-control risk baked into its design.
Overshoot in the liver, and you get acute failure. Miss the target cell entirely, and efficacy vanishes. The therapeutic window is narrower than anyone wants to admit publicly.
The 98% That Matters
The solution, if there is one, lies in the 98% of the genome that doesn't code for proteins. Promoters and enhancers—molecular switches and dials that respond to transcription factor networks—differ dramatically between cell types and disease states. Design the right regulatory element and you can theoretically program a therapy to activate only in cancerous B cells while staying silent in healthy liver tissue. Or express at therapeutic levels in muscle without triggering the immune response that comes with systemic overdose.
The challenge? This regulatory code is staggeringly complex.
A single promoter might integrate signals from dozens of transcription factors across kilobases of sequence. Context matters: chromatin state, DNA methylation, three-dimensional genome architecture. Traditional approaches involve cloning natural promoters or manually shuffling motifs, then testing candidates one by one in cell culture. It's slow. It's expensive. And it's constrained by the imagination of the designer.
Companies like AskBio's Synpromics division and Chromatin Bioscience offer libraries of pre-validated synthetic promoters for common targets—liver, muscle, natural killer cells. SynGenSys takes a knowledge-driven computational approach backed by in vivo data. These methods work, to a point. But you can't easily ask for a promoter that activates specifically in hypoxic tumor cells expressing PD-L1 while remaining off in adjacent stroma.
That's where machine learning enters the picture. Or at least, that's the pitch.
A Startup's Bold Claim

Origin, a San Francisco startup that emerged from Y Combinator's Winter 2026 batch, announced its Axis model in October 2025 with a claim that raised eyebrows: it outperforms DeepMind's AlphaGenome by 6.7% on average across regulatory element activity benchmarks.
Founded by Yash Rathod and Malhar Bhide—both UIUC graduates with backgrounds spanning computer vision and machine learning research at Wadhwani AI—the company is building what it calls the largest proprietary dataset of experimentally validated regulatory sequences. Bhide published a paper in Nature Scientific Reports on disease modeling while still in high school, which either speaks to unusual precociousness or the changing pace of computational biology. Probably both.
The pitch is straightforward: Axis generates novel regulatory DNA sequences and predicts their function, offering what Origin describes as "DNA switches and dials" to control gene expression with cell-state specificity. The founders frame this as addressing the core safety problem. Program therapeutic genes to activate only in the intended disease context, and you can theoretically lower the effective dose, reduce off-target toxicity, improve the therapeutic window.
Origin has assembled a technical advisory board that spans regulatory genomics (Manolis Kellis at MIT and the Broad Institute), AAV gene therapy development (Nicole Paulk, CEO of Siren Bio), and bioengineering (Rashid Bashir, Dean of Engineering at UIUC). The company, incorporated January 31, 2025, lists a team of around four people and is actively seeking partnerships with pharma and biotech running cell and gene therapy programs.
No public financing beyond Y Combinator has been announced as of mid-February 2026, which is either unusually quiet for a company making technical claims this aggressive, or a sign they're lining up a Series A that hasn't leaked yet.
The AI Race Nobody's Watching
Origin's claimed 6.7% advantage over AlphaGenome matters because DeepMind set a high bar when it launched AlphaGenome on June 25, 2025. Published in Nature in January 2026, the model predicts regulatory activity across 11 epigenomic modalities with 1 megabase sequence context, trained on 5,930 human genomic tracks. DeepMind offers noncommercial API access, positioning the tool as a research platform rather than a direct therapeutic design engine.
Scientific American and IEEE Spectrum coverage emphasized the community's enthusiasm for long-context noncoding DNA modeling. But questions about generalization remain—especially when you move from prediction to generation, from "what does this DNA do?" to "design DNA that does this."
Independent models like ChromBPNet and Malinois serve as baselines. ChromBPNet provides base-resolution chromatin accessibility prediction with bias correction. Malinois, developed at the Broad Institute, validates in vitro MPRA predictions for cis-regulatory element activity across K562, HepG2, and SK-N-SH cell lines. It's widely used as an independent scorer in academic benchmarks, which gives it credibility Origin will need if they want pharma partners to trust their output.
The DART-Eval framework from the Kundaje lab found that current DNA language models perform inconsistently compared to task-specific models on regulatory prediction tasks. Translation: better data curation and evaluation standards are needed before this field moves from interesting to reliable.
Generative design is moving from concept to wet lab, though. Researchers at the Broad Institute and Mass General published a Nature Genetics paper in 2026 demonstrating DNA-Diffusion, a generative model that produced functional, cell-type-specific cis-regulatory elements validated in lab assays. The team showed these AI-designed sequences could reactivate a protective gene in leukemia cell lines. That's a meaningful step toward therapeutic use, even if it's still preclinical.
Reinforcement learning approaches like Ctrl-DNA have demonstrated controllable cell-type specificity in silico. Whether that specificity holds up when you actually synthesize the DNA and test it in animals is the question that separates science fiction from FDA submissions.
How the Industry Solves This Today (Imperfectly)

While AI-designed regulatory elements remain experimental, the field has deployed several expression-control strategies clinically, with varying success.
Incorporating microRNA target sites into AAV vectors is one proven method. Add miR-122 binding sites to the 3' untranslated region and you reduce liver expression 20- to 100-fold in mice—a technique used to detarget hepatotropic vectors. Similar approaches use miR-124 to reduce neuronal expression, miR-204 for retinal detargeting. These methods work by recruiting endogenous RNA-induced silencing machinery in cells expressing high levels of the target microRNA. Elegant, but limited to tissues with known microRNA signatures.
Inducible kill switches offer another layer of control, particularly for cell therapies. The CaspaCIDe iC9 system—originally developed by Bellicum Pharmaceuticals and acquired by MD Anderson Cancer Center in February 2024—allows clinicians to rapidly ablate engineered T cells by administering rimiducid, a small molecule that activates an inducible Caspase-9. Case reports document rapid toxicity mitigation, though MD Anderson later halted Bellicum's separate GoCAR-T program over unrelated safety signals. The field has a habit of taking one step forward, two steps sideways.
Tune Therapeutics raised $175 million in January 2025 for its TEMPO epigenomic control platform, which dials gene expression up or down without making permanent DNA edits. The company plans first trials in hepatitis B in New Zealand and Hong Kong. Regel Therapeutics is pursuing targeted epi-editing with dCas9-based epimodulators paired with proprietary promoters to confine intervention to disease cells.
Academic reviews and vendor white papers increasingly describe integrative strategies that layer capsid detargeting, promoter engineering, and miRNA control to maintain efficacy at lower doses. The goal: decouple therapeutic benefit from toxicity. It's a recurring theme in every AAV safety conference, and for good reason.
When Two AI Problems Meet
Perhaps the more interesting question—and the one that could define the next phase of gene therapy—is what happens when AI-designed capsids meet AI-designed regulatory elements.
Dyno Therapeutics has inked partnerships worth more than $1 billion in milestones with Roche, Sarepta, and Astellas to engineer AAV vectors for neurological diseases using its CapsidMap and LEAP platforms. Novartis acquired Kate Therapeutics in January 2025 for up to $1.1 billion, bringing in machine learning-engineered liver-detargeted, muscle-tropic AAV capsids. These deals signal that AI-guided vector design has moved from academic curiosity to strategic pharma investment, the kind of validation that matters in this industry.
The natural next step? Pairing those capsids with programmable regulatory elements.
An AI-designed capsid that homes to diseased muscle cells becomes more powerful if the therapeutic gene inside activates only when those cells are under metabolic stress. A liver-avoiding vector becomes safer still if the cargo carries both capsid mutations and expression control elements that doubly ensure the transgene stays silent in hepatocytes. Defense in depth, borrowed from cybersecurity.
No company has publicly announced a clinical program explicitly combining AI-designed capsids with AI-designed promoters, but the logic is compelling enough that such combinations seem inevitable. Nicole Paulk, who advises Origin and runs Siren Bio, has repeatedly emphasized platform strategies to broaden AAV use in oncology—an indication where precise expression control matters even more than in monogenic rare diseases, where you can sometimes afford to be less surgical.
Watch for strategic partnerships or acquisitions that bundle delivery and expression technologies. The companies that figure out how to integrate both will have a significant advantage.
The Regulatory Backdrop (Less Hostile Than You'd Think)
The FDA's January 11, 2026 communication on flexible chemistry, manufacturing, and controls oversight for cell and gene therapies signals regulatory willingness to accommodate innovation—assuming sponsors can demonstrate comparability and safety. The agency established the Office of Therapeutic Products in February 2023 specifically to expand review capacity for CGTs, a recognition that traditional biologics frameworks don't quite fit.
International harmonization efforts like the ICH Q5A(R2) guideline, effective June 2024 in the EU, encourage next-generation quality control methods including NGS-based viral safety testing. Translation: regulators are preparing for a wave of novel manufacturing and design approaches, even if they're not explicitly endorsing AI-designed regulatory elements yet.
On the payer side, CMS's Cell & Gene Therapy Access Model launched in January 2025 for Medicaid coverage of sickle cell disease gene therapies Casgevy and Lyfgenia. The model now includes 33 states plus DC and Puerto Rico. The outcomes-based agreements tie state payments to real-world effectiveness—a model that could extend to other indications if proven administratively feasible. That's a big if.
What Matters Next

Independent, peer-reviewed benchmarking will matter. Origin's claim of a 6.7% edge over AlphaGenome on regulatory activity prediction is company-asserted. Third-party replication or shared benchmark datasets will be critical for payer and regulator confidence—not to mention pharma partners who've been burned before by computational predictions that didn't survive contact with biology.
DART-Eval and similar frameworks offer a path toward standardized evaluation, but the field remains fragmented. Academic progress on DNA-Diffusion and similar generative models is still preclinical. The first IND-enabling packages that reference AI-designed cis-regulatory elements measured via massively parallel reporter assays or perturb-seq will set important precedents, the kind that determine whether this approach becomes standard practice or remains a niche tool.
The gene therapy market recorded its highest number of approvals in 2024. Pfizer's Beqvez for hemophilia B priced at $3.5 million. Telethon's Waskyra for Wiskott-Aldrich syndrome. Novartis's intrathecal Zolgensma formulation extending to older SMA patients. The Alliance for Regenerative Medicine forecasts up to 10 CGT blockbusters by 2031, with 20 of the top 30 biopharma companies now invested in the space.
Yet the boxed warnings and safety incidents keep coming—a reminder that scaling this technology requires more than manufacturing capacity and regulatory pathways. It requires solving the fundamental control problem that Origin, DeepMind, and a small cohort of academic labs are racing to crack.
Whether AI-designed regulatory DNA becomes the safety unlock the field needs, or just another incremental tool in a still-uncertain toolkit, will become clear in the next 24 months as first-in-human data begins to surface. The patients who died from acute liver failure and T-cell malignancies deserve better than iterative improvements. Whether the industry can deliver on that expectation is the question that haunts every clinical trial now enrolling.
