AI Decodes the DNA 'On Switch' Behind 60% of Genes
UC San Diego trained AI on 500,000 DNA variants to decode the initiator, a gene 'on switch' with new power to predict disease mutations.
A switch scientists could see but not read
Every one of your genes needs a precise spot where the cellular machinery lands to start reading it, and for decades, researchers knew that spot existed without being able to describe it with any real accuracy. That gap just closed. Researchers at the University of California San Diego have used machine learning to decode the exact DNA sequence pattern of the initiator, a molecular element sitting at the very first base of a gene's transcription start site, present in roughly 60% of focused human gene promoters. The findings, published July 31 in Genes & Development, give scientists their first quantitative tool for predicting whether a specific DNA mutation at that spot will disrupt gene activity enough to cause disease.
The work was led by graduate researcher Torrey Rhyne-Carrigg in the laboratory of Professor James T. Kadonaga in UCSD's Department of Molecular Biology, according to the university's August announcement. It's a modest-sounding accomplishment on the surface, decoding one small stretch of DNA, but the initiator turns out to govern exactly the kind of genes where getting activation wrong carries the highest stakes: the genes that separate a liver cell from a neuron, and healthy tissue from a tumor.
Why nobody had cracked this before
The initiator element was first formally characterized back in 1989 by Smale and Baltimore at MIT, so this isn't a newly discovered piece of biology. What was missing was precision. Earlier researchers had only managed to produce "consensus sequences," essentially a shorthand list of which DNA letter most commonly shows up at each position along the initiator. That approach could describe an average case, but it couldn't capture how the full combination of letters, working together, determines whether the switch actually functions. A single mutation at one position might matter enormously or barely at all, entirely depending on what sits three positions over, and consensus sequences have no way to represent that kind of interaction.
Scale was the fix. Rhyne-Carrigg's team generated roughly 500,000 different DNA sequence variants of the initiator region, measured how active each one was using high-throughput sequencing, then trained a support vector regression model, a machine learning method built to capture exactly the kind of position-by-position interaction effects that had stumped earlier approaches. Once trained, the model computed predicted activity for all 1,048,576 possible ten-letter combinations, a complete activity map that no lab experiment could have generated directly. That let the team pin down the optimal sequence with a level of precision nobody had managed before, and to measure exactly how far a real-world sequence could drift from that ideal while still working. Kadonaga said the AI models delivered, for the first time, strong predictions of whether the initiator is present or absent in a given human gene.
600 million years of proof
One detail in the study does more to validate the model than any statistic about prediction accuracy could. When the researchers compared their newly decoded human initiator pattern against the equivalent sequence in Drosophila melanogaster, the common fruit fly used in countless genetics experiments, the patterns came back essentially identical. Humans and fruit flies last shared a common ancestor roughly 600 million years ago. A sequence surviving that unchanged across such a vast evolutionary distance is not an accident; it means any deviation carried a real fitness cost severe enough that natural selection weeded it out, generation after generation, across the entire history of complex animal life.
That cross-species match matters for a second reason too. If the machine learning model had simply overfit to statistical noise in its training data rather than capturing genuine biology, there would be no reason for a fly's DNA, on an entirely separate evolutionary branch, to land on the same answer. The fact that it did is independent confirmation the model found something real.
What this means for reading cancer mutations
The immediate practical value sits in interpreting disease-linked mutations, and it's a bigger deal than it might sound. A large share of DNA variants tied to disease don't sit in the protein-coding parts of genes at all; they sit in the regulatory sequences that control when and how strongly a gene turns on, and many of those variants cluster right around the transcription start site, exactly where the initiator lives. Before this study, researchers had no way to say, quantitatively, whether a single-letter change at that spot would knock gene activity down by 5% or by 50%, or whether that drop would be enough to matter clinically. The Kadonaga lab's model now provides that answer, and the team specifically flagged cancer as the disease area where the predictive power will be most immediately useful, since transcription initiation-level regulation has already been documented as a defining feature across dozens of human cancers.
That kind of precision timing at the genetic level echoes a pattern researchers are increasingly finding throughout the body's molecular systems, not unlike the strict daily schedule UT Health San Antonio discovered governing which proteins the liver releases: biology, it turns out, runs on far more precise internal rulebooks than researchers assumed even a decade ago.
A tool for building genes, not just reading them
The same model that predicts how mutations weaken an initiator can run in reverse, and that's where the applications get more ambitious. A quantitative description of what makes an initiator strong or weak is effectively a design blueprint for building synthetic promoters, engineered DNA sequences that could switch a therapeutic gene on at a specific strength in a specific cell type. That kind of precision matters enormously in gene therapy, where a treatment gene that's too weak does nothing and one that's too strong can cause its own problems. It's a complementary approach to work happening elsewhere in synthetic genomics, including a generative AI model called DNA-Diffusion published by researchers at the Broad Institute and Mass General Brigham in January 2026, which builds new regulatory sequences from scratch rather than decoding the rules behind existing ones. Together, the two lines of research represent opposite ends of the same problem: understanding nature's code, and then rewriting pieces of it on purpose, a capability that could eventually feed directly into cancer treatments built around restoring gene activity that tumors have shut down, the same broad goal driving tumor-focused immunotherapy research elsewhere in the field.
The bigger map still has holes
Kadonaga was careful not to oversell the finding as a complete solution. The initiator is one piece of core promoter architecture, sitting alongside other elements like the TATA box, and even a complete map of every core promoter switch would only be a fraction of the full regulatory system, which also includes enhancers, transcription factor binding sites, and chromatin structure. Roughly 30% of focused human promoters still have no identified core element explaining how they function at all, a gap the same experimental-plus-machine-learning approach is now aimed at closing. As Kadonaga put it, within the six billion DNA bases inside each human cell sits a gene expression code specifying when, where, and how strongly every gene should switch on, and this initiator model is a small but real piece of finally reading it.
Written by
Dr. Anand Sharma
Doctor and science communicator.