Team uses AlphaFold AI to redesign gene-editing proteins to make them safer
Google’s AlphaFold can help ID what parts of a gene editing protein enable mistakes.
A couple of decades after the discovery of systems that could selectively target DNA, we’re starting to see the first therapies based on gene editing. One challenge these developments have faced is safety. While we can make them pretty specific to the gene we want edited, the human genome is very large, and even rare DNA sequences can appear a couple of times by chance.
As a result, all the original gene-editing systems had known rates of what are called off-target effects, in which they simply edit the wrong sequence. This may be a low-probability event, but edit enough cells—and therapies generally have to edit many—and errors become inevitable.
A lot of effort has gone into finding ways to minimize or eliminate off-target edits. In a recent issue of Nature, researchers described modifying the AI protein-folding software AlphaFold to help identify key areas of gene-editing proteins responsible for off-target effects. Those areas were then modified to reduce the problems.
Gene editing and off-target effects
Gene-editing systems have three key components. The first is guide RNA, which can base-pair with the targeted genome sequence. It’s possible to design many guide RNAs that all target the same gene, so one approach to making the system safer is to pick sequences that don’t share much similarity with any other genomic locations. This is now widely used as part of the basic design process for a gene-editing system.
The second component is a Cas protein named after the CRISPR system’s Cas9 protein. It interacts with both the guide RNA and genomic RNA and helps enforce the specificity of the interactions.
If there are too many mismatches in the base pairing between the RNA and DNA, Cas9 (or other Cas family members) won’t stick there. Various approaches have led to improved Cas family members that have reduced the tendency to enable off-target edits.
The final key component of the system is the protein that interacts with Cas9 when it’s bound to DNA and modifies the DNA. In the original CRISPR system, this cut both strands of the double helix, producing damage that’s difficult to control. Researchers have since modified other proteins to interact with Cas9 but catalyze more subtle changes to DNA, such as lopping off a single base or making chemical modifications that alter how it base-pairs.
Overall, the length of the base pairing between the guide RNAs and the genome is on the order of 18 bases long. That should show up in random DNA sequences only about once in 70 billion bases, and our genome is only about 3 billion bases. By that measure, we should be good. But it turns out that Cas9 can tolerate a small number of mispaired bases without losing its ability to stick to DNA. The exact number and location of the bases where variations are tolerated can vary somewhat, making it difficult to identify in advance which guide RNAs might pose a danger.
One route to improving safety is to better understand how these off-target interactions occur.
Making contact
The team behind the new work, based at a variety of institutions in China, reasoned out their approach in advance. A perfectly matched DNA-RNA hybrid will have one structure, while one with one or more mispaired bases will have a slightly different structure. Evolution has optimized the structure of the Cas9 to stick to the former. But it apparently hasn’t prevented Cas9 from adopting slightly different conformations that can interact with mispaired structures.
If we can identify the portions of Cas9 that mediate these problematic interactions, we can modify and potentially block them.
The team’s first step was to build a large library of off-target editing sites. They did this by using a modified CRISPR system that converts the DNA base adenine to a related chemical, inosine, and then isolating any DNA fragments that contain it. They repeated this process with 10 different guide RNAs and analyzed a large number of modified DNA fragments from each to get a broad picture of the types of off-target sequences present.
The next step was to examine how the CRISPR complex interacted with them, using the AlphaFold AI-based protein-folding software. Updated versions were designed to handle interactions between proteins and nucleic acids, as well as complexes of multiple proteins. So the team fed AlphaFold versions of a target DNA sequence, along with a guide RNA, the Cas9 sequence, and an enzyme that chemically modifies bases and can stick to Cas9.
Unfortunately, it choked, placing one of the proteins in what was clearly the wrong location.
Undeterred, the team simplified things and fed AlphaFold only the DNA, RNA, and Cas9 protein, since the latter is the primary factor determining its sequence specificity. This worked much better, producing a structure that agreed with ones determined by experiments with actual nucleic acids and proteins.
By comparing the structures AlphaFold generated when fed different on- and off-target sites, the researchers found a general pattern. Many (about two-thirds) of the off-target sites caused the Cas9 protein to adopt a slightly different structure. But nearly all (over 95 percent) of them altered which amino acids contacted the RNA. So there are clearly some cases where Cas9 maintains its normal structure but amino acids within it flex around in ways that accommodate the mispaired bases of off-target sites.
Conveniently, AlphaFold was already set up to identify what is termed the “contact probability,” namely, the chance that any two items, such as amino acids or nucleotides, are within a very small distance (eight Angstroms). The researchers could take the output of the contact probability analysis for on- and off-target sites and compare them, identifying exactly which amino acids in Cas9 have altered contacts when there’s a mismatch between the guide RNA and the DNA. They termed this computerized analysis setup “ContactSeek.”
Better targeting
On its own, ContactSeek tended to produce a large list of amino acids that shift around when bound to an off-target site. So the researchers focused on regions of the Cas9 protein where these amino acids clustered, viewing this as a sign that these areas were adapting to the differences caused by mismatched bases. They then began to test versions of Cas9 with different amino acids at these sites.
In all, the researchers made 23 different swaps, putting a different amino acid into one of 10 key positions identified by their work with AlphaFold. This allowed them to find a variant with activity similar to the normal Cas9 at sites that matched the target sequence, while its off-target activity dropped from 28 percent to 5 percent. Similar results were possible when the researchers used different guide RNAs. They also showed that the approach worked for a similar system that used a different Cas protein (Cas12) to recognize the DNA/guide RNA combination.
As mentioned above, other research teams have used different approaches, such as directed evolution, to develop Cas9 variants that are less prone to off-target editing. When tested against these variants, the newly designed versions tended to produce similar or even slightly better activity and specificity. The big difference here is that the changes identified by this approach may be somewhat more specific to a given guide RNA/mismatch combination rather than more generally effective.
Of course, some of the individual changes identified here could potentially be combined with those in the previously developed Cas9 versions for even greater improvements, although that wasn’t tested here.
In any case, the work is potentially useful because it describes a general method for producing gene-editing systems tailored to prevent known off-target events. If that ends up being a bottleneck in developing a therapy, this work could be a significant breakthrough. But the researchers also suggest that the approach would be useful more generally for fine-tuning protein-DNA interactions, which could have applications far beyond gene editing.
Nature, 2026. DOI: 10.1038/s41586-026-10794-z (About DOIs).
John is Ars Technica's science editor. He has a Bachelor of Arts in Biochemistry from Columbia University, and a Ph.D. in Molecular and Cell Biology from the University of California, Berkeley. When physically separated from his keyboard, he tends to seek out a bicycle, or a scenic location for communing with his hiking boots.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0

Comments (0)