The trouble started quietly. Elevidys, the first agency-approved gene therapy for Duchenne muscular dystrophy, carried a black-box warning and far narrower prescribing guidelines. Somewhere between the laboratory and the bedside, the dream of one-time curative treatments had run into a problem the field hadn't fully anticipated.
Delivering therapeutic genes turned out to be only half the equation. Controlling where those genes switch on, how loudly they express, and for how long—that's the harder part, and perhaps the part that matters more.
Enter a four-person startup out of Y Combinator with an unusual gambit. On March 13, 2026, Origin Bio released something most biotechs guard jealously: 10,000 AI-designed DNA regulatory sequences, complete with prediction scores, binding annotations, and a public web explorer anyone could poke around in. The move was equal parts technical flex and philosophical bet—that machine learning might crack what decades of trial-and-error haven't quite solved.
The Invisible Architecture of Expression
The weakness in gene therapy isn't usually the gene itself. It's the regulatory scaffolding around it—the promoters and enhancers that act like dimmer switches and volume knobs, determining whether a therapeutic gene whispers or bellows, whether it fires up in liver cells or indiscriminately everywhere. Get the tuning wrong and CAR-T cells exhaust themselves through tonic signaling. Push micro-dystrophin expression too hard in the wrong tissue and you invite immune-mediated toxicity, or worse, dose-limiting liver damage.
This has been known, if not always heeded. Papers published in 2024 and 2025 documented that adeno-associated virus (AAV) doses above 1×10¹⁴ viral genomes per kilogram tend to wreak havoc on endothelial and hepatic tissue in nonhuman primates. More recent primate studies tied dorsal root ganglion damage to high intrathecal AAV doses. The pattern held: when expression overshoots or lands off-target, the therapeutic window narrows—sometimes dangerously.
For years, most developers reached for the easy answer: strong, ubiquitous viral promoters like CMV or CBA. They worked reliably in early research, so they stuck around. The shift toward tissue-specific, compact, de-immunized regulatory elements is relatively recent, spurred by AAV's strict payload limits (about 4.7 kilobases) and mounting evidence that broad expression invites trouble. At the 2024 ASGCT meeting, Tenaya Therapeutics showcased chimeric cardiac promoters designed to squeeze within AAV's packaging constraints while keeping expression confined to heart muscle. AskBio markets its Synpromics-derived PromPT platform on a similar premise: specificity reduces both off-tissue antigen burden and inflammatory side effects.
The catch? Designing these elements by hand—testing combinations of transcription factor binding sites, tweaking spacing and orientation—remains laborious. A 2025 preclinical study comparing MND, MSCV, EF1α, and CMV promoters in CAR-T constructs found that MND delivered functional tumor killing without excessive cytokine spikes, but only after systematic wet-lab iteration. The field's open question is whether artificial intelligence can compress that cycle from years to months.
A Sector Racing Past Its Own Infrastructure
As of late 2025, the FDA had approved 46 cell and gene therapy products, according to a WCG clinical research trends report published in January 2026; the agency itself cited "close to 50" approvals over the preceding decade. The sector raised roughly $11.1 billion across 216 financings in 2025—about 18% of total biotech therapeutic deal value, per the Alliance for Regenerative Medicine's State of the Industry briefing at the JPMorgan healthcare conference in January 2026.
Market forecasts swing wildly. Fortune Business Insights pegs the CGT market at $13.17 billion in 2025, climbing toward $200.54 billion by 2034 at a 35.6% compound annual growth rate. Other analysts put 2034–2035 figures anywhere from $90 billion to $232 billion. The spread reflects uncertainty—about reimbursement, durability, manufacturing scale—but the directionality is unmistakable.
Regulatory winds, meanwhile, are shifting, if tentatively. On January 11, 2026, the FDA pledged "increased flexibility" in chemistry, manufacturing, and controls requirements for gene and cell therapies, adopting a more lifecycle-based oversight model to accommodate rapid iteration. Ten days earlier, the FDA and European Medicines Agency jointly released 10 guiding principles for "good AI practice" across the medicines lifecycle—non-binding, but a signal that AI-generated data might enter regulatory submissions with clearer guardrails. And in January 2025, the Centers for Medicare & Medicaid Services launched a Cell and Gene Therapy Access Model targeting sickle cell disease therapies; by July, 33 states plus D.C. and Puerto Rico had signed on, negotiating outcomes-based agreements to manage list prices hovering around $2.2 million (Casgevy) and $3.1 million (Lyfgenia).
Yet the safety shadow lingers. A 2024 report documented dorsal root ganglion toxicity after intrathecal AAV-RNAi delivery in animal models. Broader reviews stress that dose escalation—the traditional route to efficacy—may be bumping against biological ceilings. If regulators and payers start demanding proof of dose-sparing, tissue-confined expression, the value proposition for better regulatory DNA design sharpens considerably.
Three Forces Converging

The push is coming from three directions at once: biological necessity, computational maturity, and—more tentatively—regulatory receptivity.
The biological urgency is most acute in AAV programs. Duchenne developers using the muscle-biased MHCK7 promoter to drive micro-dystrophin face a persistent trade-off: boost expression enough to matter, but not so much that systemic exposure triggers toxicity. Retinal gene therapy teams are gravitating toward "mini-promoters"—compact human sequences that dial down pro-inflammatory responses compared to CMV or chicken beta-actin backbones, according to a 2024 study. Cell therapy faces similar pressures. A 2024–2025 review noted that tuning CAR promoter strength can modulate tonic signaling and cytokine release, but the design space remains mostly unexplored.
On the computational front, deep learning for regulatory DNA has matured quickly. Massively parallel reporter assays (MPRAs) generate training data at scale: tens of thousands of synthetic sequences tested for activity in living cells. Models like Malinois—a convolutional neural network from the Broad Institute, Jackson Laboratory, and Yale, published in 2024—can predict and guide cell-type-specific enhancer design with MPRA-validated accuracy across K562, HepG2, and neuronal lines. A February 2026 Nature paper (based on 2025 research) described PARM, an MPRA-driven deep learning framework that models promoter "grammar" and ports predictions to patient-derived organoids. Earlier still, a 2023 Nature study demonstrated cell-type-directed enhancer optimization using Enformer and ChromBPNet, with in vivo validation in mice.
Foundation-scale genomic language models represent the next wave. A November 2025 review in Bioinformatics Advances charted the shift toward multimodal, transformer-based architectures—exemplified by Google DeepMind's AlphaGenome—that handle regulatory prediction and variant effect modeling. But these models train on public datasets like ENCODE v4, which skew toward well-characterized cell lines. Whether they generalize to therapeutic contexts—integration versus episomal, diseased tissue versus healthy—remains an open question. Independent robustness studies in late 2024 and early 2025 flagged domain-shift sensitivity as an active research problem.
Regulatory receptivity is the most tentative of the three. The FDA/EMA AI principles emphasize transparency, validation, and lifecycle monitoring, but offer no binding approval pathway for AI-designed sequences. An FDA draft guidance from January 2025 outlined "credibility" frameworks for AI models supporting regulatory decisions, yet practical precedents are sparse. What seems to be changing is the willingness to engage—provided sponsors bring orthogonal validation: wet-lab data, independent predictors, reproducibility across platforms.
A Small Team With Big Ambitions
Origin Bio sits squarely at the intersection of these trends. Founded in 2025 and part of Y Combinator's Winter 2026 batch, the San Francisco-based team bills itself as "AI for safer cell and gene therapies." CEO Yash Rathod—computer science at the University of Illinois Urbana-Champaign, first prize in the 2022 OpenCV AI Research Competition—and CTO Malhar Bhide—ML research at Wadhwani AI, previously at YC alum Automorphic—framed the company around a dataset strategy: build the largest proprietary library of experimentally validated regulatory elements, then train generative models to design new ones.
The first public milestone landed October 8, 2025, with Axis—a multifunctional transformer that both predicts regulatory element activity and generates sequences. Origin claims Axis delivers a 6.7% average improvement over AlphaGenome on ENCODE-derived binding and activity benchmarks, and can enrich transcription factor motifs up to 9× under "high affinity" prompts. (This is company-reported performance; independent third-party replication hasn't surfaced publicly.) In March 2026, the team released Switch: 10,000 AI-designed proximal enhancer-like sequences for K562 (myeloid leukemia), HepG2 (liver), and SK-N-SH (neuroblastoma) cell lines. Each sequence includes Malinois-predicted activity, transcription factor binding site annotations, quality control metrics, and predicted 3D structure—a deliberate move toward transparency and external validation.
The advisor roster suggests ambition and a certain savvy about credibility. Manolis Kellis, a professor at MIT and member of the Broad Institute, is a fixture in computational genomics and the ENCODE consortium. Nicole Paulk, CEO of Siren Biotechnology and an AAV gene therapy veteran who advises Dyno Therapeutics, Astellas Gene Therapies, and Metagenomi, brings deep domain expertise. Rashid Bashir, dean of the Grainger College of Engineering at UIUC, rounds out the scientific firepower.
Yet as of mid-March 2026, no institutional funding round beyond Y Combinator had been disclosed. The YC launch page emphasizes a go-to-market strategy centered on pharma and biotech partnerships: offering "DNA switches and dials" to confine therapeutic gene activity to target cell states, thereby lowering tonic signaling in CAR-T or improving AAV specificity. The pitch resonates with known pain points. Whether Origin can convert research-stage designs into clinical-grade components—and secure the capital to do so—remains to be seen.
An Ecosystem Tackling Adjacent Problems

Elsewhere, companies are attacking related layers of the problem. Dyno Therapeutics uses machine learning to engineer AAV capsids; in 2025, Roche exercised an option for a CNS-targeting capsid, and Dyno unveiled new muscle, eye, and neurological variants at ASGCT. Capsida Biotherapeutics secured IND clearance in May 2025 for CAP-002, a CNS-tropic AAV for STXBP1-related epileptic encephalopathy, and in January 2026 announced a $40 million AbbVie opt-in based on primate data. Both companies focus on the delivery vehicle. Origin is betting on the cargo control system—the regulatory DNA that programs expression once the vector reaches its destination.
AskBio's Synpromics platform, acquired in 2019, offers proprietary synthetic promoters via the PromPT bioinformatics engine. Chromatin Bioscience markets synthetic promoters for CGT manufacturing and in-vivo selectivity. Senti Biosciences explores logic-gated CAR-NK and CAR-T constructs, with partnerships spanning Spark (next-gen AAV) and BlueRock (iPSC). Each represents a thesis that gene therapy's next bottleneck is control, not just delivery or editing.
Academic consortia are also in motion. The Broad/JAX/Yale collaboration behind Malinois and the CODA framework released genome-wide CRE scanning and design tools validated across multiple cell lines. The 2023 Nature synthetic enhancer work—using Enformer and ChromBPNet to optimize enhancers in vivo—remains a methodological benchmark. PARM's promoter grammar modeling, published in early 2026, extends MPRA-derived insights into patient organoids, hinting at clinical translation. These efforts yield public datasets and independent predictors that startups like Origin can (and evidently do) leverage for benchmarking.
The Translation Chasm
The immediate question is whether AI-designed regulatory elements can clear the translational chasm. Many models train on MPRA data—episomal, reporter-based assays that may not fully predict behavior in genomic integration or disease-relevant chromatin contexts. The 2026 Nature paper on PARM took a step by validating in patient-derived organoids, and CRISPRa screens offer another path to interrogate endogenous loci. But the field still lacks large-scale clinical datasets linking synthetic regulatory sequences to therapeutic outcomes.
Regulatory validation remains hazy. The FDA/EMA AI principles are a start, but they don't specify how sponsors should document training data provenance, model versioning, or failure modes for generative DNA models. The January 2025 FDA draft guidance on AI credibility is conceptual. Practical precedents—INDs or biologics license applications with AI-designed promoters or enhancers as part of the construct—aren't yet public knowledge. Origin and peers will likely need to furnish extensive orthogonal wet-lab validation, third-party MPRA or CRISPRa data, and in-vivo proof-of-concept before regulators treat these sequences as more than computational hypotheses.
The safety bar, meanwhile, is rising. Broader scrutiny of gene therapy safety signals that regulators may demand tighter demonstration of dose-sparing, tissue-specific expression—precisely the value proposition Origin and others claim. But that also means longer, costlier development timelines and potentially more scrutiny of every component in a therapeutic construct, from capsid to transgene to regulatory cassette.
Reimbursement experiments like the CMS Cell and Gene Therapy Access Model (33 states plus D.C. and Puerto Rico as of July 2025) could ease payer resistance to multimillion-dollar therapies, provided outcomes-based agreements prove durable. If payers see evidence that better-controlled expression reduces adverse events and improves efficacy, they may be more willing to cover upfront costs. Conversely, high-profile safety failures could tighten purse strings and slow uptake.
Competitive dynamics will hinge on data moats and partnership velocity. Origin's strategy—building a proprietary assay dataset while releasing public benchmarks like Switch—is a dual play: demonstrate technical credibility to attract pharma collaborators, while accumulating private data that compounds model performance. Dyno and Capsida pursued similar strategies in capsid engineering—publish enough to establish legitimacy, retain enough to maintain an edge. The winner in regulatory DNA design may be whoever accumulates the most diverse, clinically relevant validation data fastest.
What Founders Are Weighing

Executives in gene and cell therapy face a strategic choice. Do they wait for off-the-shelf synthetic promoters and enhancers to mature, or invest now in partnerships with platforms like Origin, AskBio, or academic consortia? The cost of inaction is programs stuck with suboptimal, one-size-fits-all regulatory elements—potentially hitting dose-limiting toxicity or failing efficacy endpoints. The cost of early adoption is validation risk: AI-generated sequences that shine in silico may underperform in the clinic, and regulatory agencies may demand evidence standards that don't yet exist.
The broader arc, though, feels clear. As the field shifts from first-generation, broadly expressed therapies toward precision medicine—tissue-specific AAVs, tuned CAR-T constructs, logic-gated cell therapies—the ability to program gene expression with single-digit percentage precision becomes a competitive and clinical imperative. Machine learning can accelerate that. But only if the models generalize beyond public datasets, the regulatory path becomes navigable, and developers treat regulatory DNA as seriously as they treat the therapeutic transgene itself.
For now, 10,000 AI-designed sequences sit in a public repository, waiting to be tested. Whether they represent the vanguard of a new design paradigm or another incremental tool in a crowded toolbox will depend on what happens in the wet lab, the clinic, and the FDA review room over the next few years. The data, as always, will tell the story. And in gene therapy, perhaps more than anywhere else, the story has a way of surprising even those writing it.
