Gene editing has promised a way to treat some diseases, including inherited genetic disorders, blood conditions and cancer, at their genetic roots. But even as tools such as CRISPR, which allows scientists to alter DNA sequences and modify gene function with high precision, have become increasingly powerful, a practical challenge remains: Many genome editors are too large to fit easily into the delivery vehicles used to carry them into cells.
A gene-editing system must get inside a cell before it can make the desired genetic change, but cell membranes prevent large molecules from simply entering. To overcome this obstacle, researchers can use lipid nanoparticles to deliver editor proteins or messenger RNA, or leverage engineered viruses to deliver the DNA encoding the editor. The larger the editor or its genetic instructions, the more difficult this delivery process can be.
One possible solution to widen the use of gene editors is to make them smaller.
Adeno-associated viruses, or AAVs, are commonly used to deliver gene therapies, but they can carry only a limited amount of genetic material. Compact gene-editing tools could help overcome this constraint by leaving more room for the other components needed to direct and control editing.
Now, Penn Engineers have developed a more powerful version of Fanzor2, a small, RNA-guided DNA-cutting protein that is being explored as a potential alternative to CRISPR-Cas systems, using artificial intelligence (AI) to help guide the process.
In a study published in Nature Biotechnology, a team led by Xue Sherry Gao, Presidential Penn Compact Associate Professor in Chemical and Biomolecular Engineering and in Bioengineering, postdoctoral fellow Shijie Wan, and others have developed FanzMAX, an engineered Fanzor2-based editor that achieved an editing efficiency as high as 97% at the best-performing genomic target. The development of this editor was supported by a machine-learning approach created by the team called EvoMax, which prioritizes candidate mutations using small amounts of experimental data.
“Fanzor2 is unusually small compared with many gene-editing proteins, and that matters because a smaller editor leaves more space in a delivery vehicle for the other components needed to make the system work, such as the RNA that guides the editor to the right genomic site,” says Gao.
The compact size of Fanzor2 allowed the team to design single-AAV editor systems. But small size alone is not enough; the editor also needs to be active, accurate and safe.
The lead faculty member of this work was Xue Sherry Gao, Presidential Penn Compact Associate Professor in Chemical and Biomolecular Engineering and in Bioengineering.
A Small Editor with a Big Engineering Problem
CRISPR-associated proteins such as Cas9 originated in bacteria, where they evolved as part of microbial immune systems. Fanzors are different. They are RNA-guided DNA-cutting proteins also found in eukaryotic organisms, the same broad category of organisms that includes humans.
They are also remarkably compact. Fanzor2 proteins in particular are fewer than 500 amino acids long, making them considerably smaller than Cas9, which can be well over 1,000 amino acids long, and potentially well suited to delivery by AAVs.
However, Fanzor2 proteins studied previously have generally shown only modest DNA-editing activity when tested in mammalian cells.
One option to tackle this protein engineering problem is to test mutations experimentally. But a protein roughly 500 amino acids long has an enormous number of possible sequences, making it prohibitively burdensome to test them all.
That’s where artificial intelligence can come in to narrow the search. But AI models also need data, usually lots of it to make accurate predictions. And, because Fanzor2 is a relatively new protein family with little experimental information available, even an AI approach would need to be innovative.
“We had only about 200 measured single-mutation activities,” says Wan. “That is a tiny sample compared with the full mutational space of a protein that is roughly 500 amino acids long.”
Shijie Wan (left) and Pranay Vure, a Ph.D. student in the Gao lab (right).
Teaching AI to Learn From a Small Dataset
The researchers developed EvoMax to tackle this sparse-data problem.
EvoMax combines three types of information. One part learns from mutations researchers have already tested in the lab. A second part looks at what evolution has already explored in natural protein sequences. And a third part asks whether a proposed mutation seems compatible with the protein’s three-dimensional shape.
“In our previous work, computational tools have helped us compare protein sequences, predict protein structures and generate ideas for experiments,” says Gao. “In this project, we used computational and machine-learning tools more directly to help decide which mutations to test. The computer does not give the final answer. Instead, it helps us choose a smaller and better set of candidates to test experimentally.”
Instead of testing thousands of possibilities, EvoMax generated candidates, and the researchers selected a small number (around 10 to 20) to test experimentally. The best-performing variant from each round became the starting point for the next, while the researchers adjusted how they scored the different traits they were looking for.
After three rounds of optimization, FanzMAX v3 paired with the optimized RNA scaffold produced more than 13-fold higher reporter activation than the naturally occurring Fanzor2 enzyme paired with the original RNA scaffold.
“AI is just a tool,” says Wan. “The predictions are far from perfect, so every candidate still has to be tested in the lab. AI helps you move the project forward faster, but it doesn’t replace the human scientist. More often than not, we’re working with AI tools like collaborators, as partners in experimental design to help us decide what to test next, but we, as humans, conduct the experiments that tell us whether those predictions are actually right.”
Building a Better System
Critically, Fanzor2 works with an RNA molecule that helps guide the protein to its DNA target. The team therefore needed to engineer the RNA component as well, developing a shorter scaffold that improved the system’s performance. They also added a human RNA-binding protein called “La” to help stabilize the RNA.
The resulting system, FanzMAX v3-hLa, combined improvements to the protein, RNA and supporting components.
At one genomic target, the original Fanzor2 system edited about 12% of DNA molecules. The optimized protein increased that to 66%, while adding La brought editing efficiency to 97%.
Across 19 genomic targets, FanzMAX v3-hLa also substantially outperformed the original Fanzor2 system.
The researchers further showed that EvoMax could be applied to other Fanzor2 proteins, including related versions of the protein that showed no detectable DNA-editing activity in the initial test.
“This study shows that the approach can work across different members of the Fanzor2 family,” says Gao. “We have not yet shown that the same version of EvoMax will work for unrelated protein families. In principle, a similar strategy could be adapted to other proteins if researchers have a way to measure activity, some initial experimental data and useful sequence or structural information. But the model would likely need to be adjusted and tested again for each new problem. Extending EvoMax beyond Fanzor2 is an important future direction, but it still needs direct experimental validation.”
When More Power Isn’t Necessarily Better
The team then tested three single-AAV FanzMAX configurations in mice, targeting the PCSK9 gene, which plays a role in regulating cholesterol.
One version produced about 25% editing in the liver and reduced circulating PCSK9 levels by 34% without detectable toxicity during the three-week study.
Two later configurations, AAV 2.0 and AAV 3.0, showed higher activity in cultured cells, but produced different results in mice. The mice developed severe toxicity, and the researchers observed larger DNA deletions. However, because these configurations also differed in other design features, the experiments did not establish which individual component caused the toxicity.
“You hope your effector can be as strong as possible,” Wan says. “But these results show that editing activity must be considered together with the kinds of DNA changes it produces and how well it is tolerated in the body.”
This need to balance efficacy and safety is common across work in this field.
“The goal is controlled editing, not maximum editing,” says Gao. “Future work needs to identify why some versions were tolerated and others were not, including the dose, expression level, guide RNA design, duration of editor activity and unwanted DNA changes.”
More Than a Better Gene Editor
Proteins can have an almost unfathomable number of possible sequences, while researchers rarely have enough experimental data to test more than a tiny fraction of them. In this study, EvoMax helped bridge that gap by combining a small amount of high-quality experimental data with information encoded in evolutionary history and protein structure to prioritize experiments.
“AI can’t do everything,” says Wan. “At least at the current stage, we still rely primarily on human intelligence instead of artificial intelligence.”
The result was not simply a more active Fanzor2, but a demonstration of how combining different approaches can turn a promising but limited biological system into a useful gene-editing tool.
“This final editor did not come from one single breakthrough,” says Gao. “The project involved finding natural Fanzor2 proteins, improving the RNA that guides the editor, using computation to prioritize mutations, testing those mutations in cells and then evaluating delivery in animals. It was a collaboration between computational modeling, molecular biology, genome editing and animal studies. This reflects our lab’s larger goal, which is to combine biological discovery, engineering and computation to turn naturally compact but initially limited systems into useful tools.”
Learn more about the work being done in the Gao lab here.
This work was supported by National Institutes of Health (NIH) grants HL157714 and HL173243, a University of Pennsylvania startup fund and NIH grant R01-HD110733.
