Skip to content
Tech News
← Back to articles

Google's AI genome system evaluates every possible one-base change

read original get The Gene: An Intimate History (Siddhartha Mukherjee) → more articles
Why This Matters

Google's AlphaGenome Atlas applies its AI model to all 9 billion possible single-base variants across the human genome, aiming to predict which changes in non-coding DNA actually matter. Since over 97% of the genome doesn't encode proteins, a single unified resource for interpreting regulatory sequence could speed up genetics and disease research — though its real added value over its training data remains unproven until researchers adopt it.

Key Takeaways
Worth a Look

The Gene: An Intimate History (Siddhartha Mukherjee) — If AlphaGenome Atlas mapping all 9 billion possible single-base variants makes you want to actually understand what those bases do, Mukherjee's sweeping history of genetics is the ideal companion read. It walks through how we learned to read DNA, mutations, and gene regulation in clear, story-driven prose that makes headlines like this one click.

See The Gene: An Intimate History (Siddhartha Mukherjee) on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

On Tuesday, Google announced AlphaGenome Atlas, a resource that attempts to predict the consequences of every possible single-base variant in the human genome. The human genome is about 3 billion bases long, so trying the other three DNA bases that don’t appear in our reference genome means sending a total of 9 billion bases through AlphaGenome software.

AlphaGenome is designed to identify potential functions of non-coding DNA, which does not encode proteins but makes up the vast majority of the human genome. Some of this non-coding DNA is essential for controlling the activity of the protein-coding portion—it tells the cell where and when to make messenger RNAs, how to process them into mature protein-coding forms, and so on. But much of it appears to be little more than the remains of viruses and other molecular parasites.

Being able to identify the functional portion is very useful, as is having all the analysis done by a single software package. But until biologists start to use it heavily (assuming they do), it won’t be clear what AlphaGenome offers beyond what we could have gotten out of its training data.

Non-coding sequences

While we tend to focus on proteins, the portion of the human genome that encodes proteins is less than 3 percent. Most of the genome is non-coding and contains a mix of things, including centromeres, which help ensure chromosomes are divided evenly between cells, and caps that protect the chromosome ends. There’s also the regulatory DNA that controls gene activity, along with the signals that help determine what should and shouldn’t be included in mature messenger RNAs produced by genes. Other sequences help control how the DNA is packaged inside the cell.