AI and DNA: A New Era of Genetic Research Is Beginning

Prefer to listen instead? Here’s the podcast version of this article.

There are roughly three billion DNA base pairs in the human genome. At each position, the existing DNA letter can theoretically be replaced by one of the other three letters.

 

Do the math and you arrive at a rather intimidating number: about nine billion possible single-letter changes.

 

Testing every one of those changes experimentally would be practically impossible. Google DeepMind’s answer is to let artificial intelligence do the first pass.

 

In September 2026, Google DeepMind introduced AlphaGenome Atlas, a massive AI-generated catalogue containing predictions for the molecular effects of approximately nine billion single-nucleotide variants—essentially every possible one-letter substitution across the human genome. Google describes it as its most comprehensive effort yet to map how genetic mutations could alter molecular biology. Readers looking for the technical announcement can explore [Google DeepMind’s AlphaGenome Atlas launch article]

 

It is an ambitious idea, even by modern AI standards. Instead of asking an AI model about one genetic variant at a time, DeepMind has essentially asked it the genomic equivalent of: “What happens if we change every letter, one by one?”

 

 

What Is Google’s AlphaGenome Atlas?

AlphaGenome Atlas builds on AlphaGenome, a genomic AI model that Google DeepMind first unveiled in 2025.

 

AlphaGenome was designed to analyze extremely long DNA sequences—up to one million base pairs—and predict thousands of molecular characteristics associated with gene regulation. These include effects involving gene expression, RNA splicing, chromatin accessibility and other mechanisms that influence how genetic information is actually used by cells. For a deeper look at the underlying model, Google DeepMind’s earlier [AlphaGenome introduction explains how the system approaches regulatory variant prediction].

 

That distinction matters.

 

When most people hear the word gene, they naturally think about DNA that directly encodes proteins. Yet only around 2% of the human genome is protein-coding. Much of the remaining genome participates in controlling when, where and how genes are activated, while other regions have functions that remain poorly understood.

 

Think of protein-coding DNA as the recipes in a cookbook. Non-coding regulatory DNA contains a huge amount of the information telling the kitchen which recipe to use, when to use it, in what quantity and under what conditions.

 

A one-letter change in the wrong regulatory location can therefore matter even when it does not directly alter a protein.

 

That difficult territory is where AlphaGenome becomes particularly interesting.

 

 

Why Nine Billion DNA Variants?

The calculation behind the headline is surprisingly straightforward.

 

The human reference genome contains roughly three billion DNA positions. DNA uses four nucleotide bases—A, C, G and T. At any given position, the letter already present could theoretically be replaced by any of the other three.

 

Three billion positions multiplied by three possible substitutions gives approximately nine billion single-nucleotide variants.

 

AlphaGenome Atlas precomputes predictions for those substitutions rather than forcing researchers to repeatedly run the model themselves.

 

That could turn an expensive computational question into something closer to a database lookup.

 

As [Nature] individual nucleotide substitutions are among the most common forms of human genetic variation. Some have little or no meaningful effect, while others can contribute to disease risk or, in rare cases, play a direct role in genetic disease.

 

The challenge is figuring out which is which.

 

Nine billion possibilities make “check them all in the lab” a fairly terrible project plan.

 

AI offers a way to prioritize.

 

 

The AlphaGenome Variant Impact Score Could Be Just as Important as the Atlas

Producing billions of predictions creates another problem: researchers still need a practical way to navigate them.

 

Google DeepMind therefore introduced the AlphaGenome Variant Impact, or AVI, score.

 

The AVI score combines information from AlphaGenome with predictions from AlphaMissense, DeepMind’s system for assessing protein-altering variants. The goal is to provide researchers with a single metric for prioritizing potentially consequential variants across both coding and non-coding regions.

 

The Atlas can also break the predicted impact down into biological categories, helping researchers investigate why a particular variant receives attention—for example, whether the model predicts disruption to RNA splicing, gene expression or chromatin-related processes.

 

This is important because useful scientific AI cannot simply shout, “Interesting mutation!” and walk out of the room.

 

 

Why AlphaGenome Could Matter for Rare Disease Research

One particularly promising application is variant prioritization in rare diseases.

 

Genome sequencing can reveal enormous numbers of differences between an individual genome and a reference genome. The hard part is often determining which differences are biologically meaningful.

 

AlphaGenome Atlas could help researchers rank candidate variants before committing scarce laboratory resources to experimental validation.

 

Google says collaborators have already used the Atlas to investigate rare-disease variants. One reported example involved researchers using its AVI score to prioritize a variant affecting the DNM1 gene, with AlphaGenome predicting the creation of an incorrect splice site.

 

That does not mean AI has “diagnosed” the disease. It means AI has helped researchers narrow the search space and generate a mechanism worth investigating experimentally.

 

 

AI Is Especially Useful in the Genome’s “Control Panel”

The most intriguing part of AlphaGenome may be its ability to interrogate non-coding DNA.

 

Traditional genetics understandably focused heavily on protein-coding genes because their biological effects were easier to connect to specific mechanisms. Regulatory DNA is messier.

 

Its influence can depend on cell type, surrounding sequence, chromatin state and complex interactions with proteins and other molecular machinery.

 

AI models are particularly attractive for this type of problem because they can search enormous datasets for context-dependent patterns that are difficult to express through simple rules.

 

But there is an important catch.

 

[Ars Technica] notes that AlphaGenome is currently based on the biological datasets and cell types available to researchers. Predictions involving poorly characterized biological contexts therefore remain an area where real-world validation will matter enormously.

 

In other words, AI can offer researchers a very good map.

 

It does not guarantee that every road on that map has been driven.

 

 

AlphaGenome Is a Research Tool, Not a Medical Diagnosis Machine

This may be the most important point for anyone reading about AlphaGenome outside the genomics community.

 

AlphaGenome Atlas is not a clinical diagnostic system.

 

Google DeepMind explicitly states that information produced by AlphaGenome Atlas is not a substitute for professional medical advice, diagnosis or treatment and that the system has not been validated or approved for clinical use.

 

Predicting a molecular effect is not the same thing as predicting whether a person will develop a disease.

 

Human health emerges from complicated interactions among multiple genes, regulatory processes, environment, lifestyle, development, age and many other factors.

 

This is why responsible AI becomes particularly important once artificial intelligence moves from web search and productivity tools into biology and medicine.

 

 

Genomic AI Creates New Ethical Questions Too

The better AI becomes at interpreting genomes, the more important governance becomes.

 

Genetic information is unusually sensitive because it can reveal information not only about one person but also, indirectly, about biological relatives. Organizations combining genomic datasets with increasingly powerful AI models therefore need strong controls covering consent, access, security, data provenance and appropriate downstream use.

 

Bias deserves equal attention.

 

Genomic datasets do not represent every population equally. A model that performs strongly against widely studied populations, tissues or cell types may not automatically generalize equally well everywhere else.

 

Scientific teams should therefore treat AI predictions as evidence to investigate—not infallible biological truth.

 

 

Conclusion

Google’s AlphaGenome Atlas shows how powerful AI is becoming in the study of human genetics. By evaluating billions of possible one-letter DNA changes, the system can help researchers quickly identify which genetic variants may be worth studying more closely.

 

This does not mean AI can replace doctors, geneticists, or laboratory research. Instead, tools like AlphaGenome can help scientists narrow down huge amounts of genetic information, develop better research questions, and focus experiments on the most promising possibilities.

 

The real value of Google’s AI genome system is speed and scale. What could take researchers an enormous amount of time to investigate manually can now be explored much more efficiently with AI. As these tools improve, they could support advances in rare disease research, drug discovery, and our broader understanding of how the human genome works.

 

At the same time, accuracy, privacy, fairness, and responsible use will remain essential. AI can point researchers in the right direction, but human expertise and scientific testing are still needed to confirm what those predictions actually mean.

 

Ultimately, AlphaGenome Atlas is another sign that AI is becoming an important partner in scientific discovery—helping researchers turn billions of possibilities into useful questions that could lead to better insights about human health.

WEBINAR

INTELLIGENT IMMERSION:

How AI Empowers AR & VR for Business

Wednesday, June 19, 2024

12:00 PM ET •  9:00 AM PT