DeepMind’s free AI DNA mutation predictions cover 9 billion genome variants

20 hours ago 28
AI DNA mutation predictions

Scientists have spent more than two decades holding the human genome’s “instruction manual” without being able to fully read it. That gap just got noticeably smaller. Google DeepMind has released a free, publicly accessible database of AI DNA mutation predictions covering every possible single-letter change that could occur across the human genome — all 9 billion of them, according to the company and reporting from Fortune.

Key takeaways

  • Google DeepMind released the AlphaGenome Atlas, a free research database of AI-generated predictions for all 9 billion possible single-letter human DNA mutations.
  • The tool interprets both protein-coding DNA and the much larger, harder-to-read regulatory portion of the genome.
  • Pushmeet Kohli, DeepMind’s vice president for research, said this marks the first time a researcher anywhere can access a comprehensive map of human genetic variation “by simply opening a browser.”
  • The Atlas is free for non-commercial research starting immediately; commercial access will follow soon through licensing on Google Cloud, and sister company Isomorphic Labs will also need that license.
  • A paper describing how the Atlas was built has been posted as a preprint on bioRxiv.

Google DeepMind launches AlphaGenome Atlas with AI-driven mutation predictions

The AlphaGenome Atlas is best described as a giant, precomputed catalogue: for every possible swap of one DNA base for another, it estimates what that single change is likely to do to the biological machinery that turns genes on and off. Instead of researchers having to run predictions one variant at a time, the answers are already sitting in the database, ready to be looked up.

Mapping all 9 billion possible single-letter mutations

DNA is built from four chemical bases — cytosine, guanine, adenine and thymine — paired together along a double helix. When one of those letters mutates, swapping an A for a G, for example, it can reshape the protein that stretch of DNA instructs a cell to build, and that reshaping can sometimes trigger disease. Google DeepMind built the Atlas by running AlphaGenome, an AI model it released last year, across a reference version of the human genome and comparing every base against each of the three possible alternative letters, generating predictions across all 9 billion combinations.

According to DeepMind’s research paper, each individual variant is tied to roughly 27,000 separate predictions covering effects on gene expression and how DNA gets transcribed into instructions for building proteins, spanning hundreds of human and mouse cell and tissue types. The team also scored more than 100 million insertions and deletions drawn from population databases including the UK Biobank and the U.S. National Institutes of Health’s All of Us project.

To make sense of that volume of data, DeepMind is also releasing a summary metric called the AlphaGenome Variant Impact score, which folds AlphaGenome’s regulatory predictions together with an earlier protein-focused model, AlphaMissense. A score of 10 places a variant among the 10% most impactful in the genome; a score of 30 puts it in the top one in a thousand.

Interpreting the genome’s protein-coding and regulatory DNA

Protein-coding DNA makes up only about 2% of the human genome. The remaining 98% is regulatory material that governs when and where genes switch on, and it has historically been far tougher for scientists to decode. That distinction is central to why the AlphaGenome Atlas matters: it doesn’t just flag mutations that alter a protein directly, it also interprets the vast regulatory landscape surrounding those genes. Žiga Avsec, DeepMind’s genomics lead, framed the split this way: AlphaMissense looks at proteins, while AlphaGenome focuses on the regulatory part of the genome. The Atlas also includes a catalogue of more than 2,500 recurring short DNA sequences, known as motifs, mapped across the genome — the segments where transcription factors physically bind.

Advancing genetic research by replacing slow previous methods

Before the Atlas existed, testing what a single mutation might do required running a model one variant at a time or physically testing it in a laboratory — a process that was, in DeepMind’s own words, painstakingly slow. Covering all 9 billion possible single-letter mutations that way would have taken many human lifetimes.

Pushmeet Kohli on finally being able to “read the book”

Kohli, who leads DeepMind’s AI for science team, told reporters on a briefing call that this is the first time any researcher in the world can reach a comprehensive map of human genetic variation “by simply opening a browser.” He tied the release to unfinished work from the Human Genome Project, which mapped the entire human DNA sequence back in 2003. “As the saying goes, we bought the book,” Kohli said, “but we did not understand how to read it.”

That framing matters beyond the marketing line. Genetic diagnosis today often stalls not because sequencing is hard — it’s cheap and fast now — but because interpreting what a given mutation actually does remains slow and uncertain, especially in the non-coding genome. Early testers gave a glimpse of what closing that gap could mean in practice. Researchers at the Broad Institute working with the GREGoR Consortium, which studies unexplained rare genetic disorders, used the Atlas’s impact score to revisit a case involving epileptic encephalopathy and were pointed toward a variant in the gene DNM1 that creates a hidden splice site in a brain-specific version of the gene. Laboratory experiments later confirmed the prediction, and the variant was reclassified as likely pathogenic. In a separate test, a researcher at the University of Exeter applied the Atlas to genome data from more than 54,000 UK Biobank participants and found 22% more associations between rare variants and blood protein levels than the same analysis produced without it.

Access and licensing: free for research now, commercial deals soon

The Atlas is available immediately for non-commercial use through a dedicated website Google DeepMind has set up. That means academic labs, university researchers and public-health scientists can start querying the database of AI DNA mutation predictions today, without paying for access.

Google Cloud licensing and Isomorphic Labs’ role

Commercial use is a different story, at least for now. DeepMind said the Atlas will become available for commercial use through a licensing arrangement on Google Cloud “soon,” though the company has not detailed what those terms will look like. Kohli confirmed that Isomorphic Labs, DeepMind’s sister company focused on using AI for drug discovery, will get access to the Atlas — but only under a commercial license, just like any other paying customer. That detail is worth noting: even a closely affiliated Alphabet company doesn’t get a free pass into the same data other biotech firms will eventually have to license.

DeepMind has been careful to frame the Atlas as a research aid, not a diagnostic replacement. Avsec said AlphaGenome performs well for variants affecting splicing or gene promoters but can miss effects in other regulatory elements, particularly enhancers, and is not as reliable overall as DeepMind’s protein-structure model AlphaFold. The predictions, he said, are accurate enough to point researchers in the right direction for follow-up studies, but they shouldn’t be treated as final answers.

Scientific validation and what happens next

A scientific paper describing how the Atlas was built and validated has been published as a preprint on bioRxiv, the biomedical repository where researchers post findings ahead of formal peer review. That gives outside scientists a way to scrutinize the methodology behind the AI DNA mutation predictions before the tool becomes embedded in broader research pipelines.

Signs of that embedding are already visible. Ewan Birney, director of EMBL’s European Bioinformatics Institute, said his organization is working to integrate the Atlas’s impact score into Ensembl’s Variant Effect Predictor, a widely used annotation tool. If that integration goes through, the reach of DeepMind’s genetic variation AI model would extend well beyond researchers who visit the Atlas website directly, folding it into infrastructure that geneticists around the world already rely on daily.

FAQ

What is the AlphaGenome Atlas?

The AlphaGenome Atlas is a free research database released by Google DeepMind encompassing computational forecasts regarding the consequences of each of the 9 billion conceivable single-letter variations throughout the human genome.

How does the AlphaGenome Atlas improve genetic research?

It replaces slow prior methods of testing variants individually or in labs by providing immediate AI-predicted impacts of every mutation on both protein-coding and regulatory DNA, accessible through a browser.

Who can use the AlphaGenome Atlas and how?

The Atlas is available immediately for non-commercial research use via a dedicated website, with commercial access coming soon through licensing via Google Cloud.

Has the AlphaGenome Atlas been scientifically validated?

A scientific paper describing its creation has been published as a preprint on bioRxiv, supporting the technology and dataset behind it.

Article produced with the assistance of artificial intelligence and reviewed by the editorial team.

Read Entire Article