Google DeepMind's Mutation Atlas Reveals Previously Unexplored Human DNA

Google DeepMind’s Mutation Atlas Reveals Previously Unexplored Human DNA

The molecular impacts of altering each DNA base pair in the human genome one by one have been cataloged in the AlphaGenome Atlas. The deep learning algorithm from Google examined the effect of changing each of the three billion DNA letters individually and evaluated each alteration based on its anticipated effect.

This encompasses what was previously deemed junk DNA – segments that do not encode proteins but are now acknowledged as essential for activating or deactivating genes or regulating their functions. Under 2% of the human genome encodes proteins, with the non-coding ‘dark genome’ comprising the remainder. The implications of variation within this dark genome are less understood. Now, a geneticist can investigate how altering a DNA base anywhere in the genome affects a gene.

‘Similar to how a topographic map illustrates elevation at various geographical points, the Atlas illustrates the consequences of mutations at different sites in the genome,’ remarks Žiga Avec, a computational biologist at Google DeepMind, which developed the Atlas using its AlphaGenome AI model. ‘It’s a series of maps because there is an individual map for each cell type and each regulatory function.’

What do individuals refer to when discussing AI in scientific contexts?

Artificial intelligence (AI) is a broad term frequently misused to signify a range of related but simpler processes.

AI refers to the capability of machines and computer programs to perform tasks that are typically human domains, such as reasoning, responding to inputs, and making decisions.

Generative AI is a recent form of AI that analyzes and identifies patterns in training datasets to create original text, images, and videos in reaction to user requests. ChatGPT, Microsoft Copilot, Google Gemini, and more recently X’s Grok are all instances of chatbots utilizing generative AI.

Neural networks consist of an interconnected network of artificial neurons, resembling biological brains, that recognize, assess, and learn from statistical patterns in data.

Machine learning is a subdivision of AI that enables machines to learn from datasets and forecast outcomes based on novel data, without explicit instructions from programmers. Machine learning models enhance their efficacy as they process more data.

Deep learning is a refined form of machine learning that implements neural networks with multiple layers to scrutinize intricate data from vast datasets. Applications of deep learning encompass speech recognition, image creation, and translation.

Large language models or LLMs are a category of deep learning trained on extensive data to comprehend and generate language. LLMs discern patterns in text by predicting the subsequent word in a sequence, and these models can now compose prose, analyze online text, and engage in conversations with users.

Testing all single-nucleotide variants in a laboratory is nearly infeasible, but the Google DeepMind team employed AlphaGenome to predict the repercussions of genetic modifications. ‘They utilized datasets from numerous experiments in various tissues and attempted to forecast the impact of a variant using those across the entire genome,’ states Caroline Wright, a geneticist at the University of Exeter, UK, who contributed to the recent Atlas preprint.

AlphaGenome, unveiled in January, can forecast which genes are expressed in different tissues, where they undergo splicing, or which DNA letters are accessible, proximally situated, or attached to proteins. The Atlas applied knowledge from this AI tool throughout the human genome for each of the three possibilities across three billion bases, as well as millions of additions and deletions documented in biobanks. It subsequently produced a variant score to rank mutations by their impact. ‘It provides a ranked list of the most significant predicted effects, but you can also select a specific cell type of interest,’ explains Carl de Boer, a genomics researcher at the University of British Columbia, Canada.

Illuminating a rare disorder

The Atlas uncovered rare non-coding variants that influence circulating protein levels. It also utilized its AlphaGenome Variant Impact to associate a variant in a gene (DNM1) with a rare condition: epileptic encephalopathy. It predicted that this variant induced an error resulting in an unusually elongated protein – subsequent tests confirmed the prediction.

‘If you identify a gene or area linked to a disease of interest, you can then investigate further,’ remarks Wright. ‘It may indicate that altering a protein level could affect a disease, which could be advantageous for identifying new pharmaceutical targets.’

‘The genome is scattered with switches known as enhancers that activate or deactivate genes, and variants there can affect diseases,’ explains Jorge Ferrer at the Centre for Genomic Regulation (CRG) in Barcelona, Spain, who studies genomes from diabetes patients. ‘The most prevalent diseases are largely affected by variants that operate not in protein-coding genes, but on these switches and enhancers. They might exert a minuscule effect, but if you accumulate the effect of numerous variants from various regions of the genome, it can significantly influence disease susceptibility.’ These include heart disease, lipid