How the variant-impact map was built
AlphaGenome Atlas is not a catalogue of results from one experiment. Google DeepMind says the model generated predictions for roughly nine billion possible single-letter changes in the human genome, producing a resource of about one petabyte. For each variant, the system estimates effects across thousands of molecular properties in many cell and tissue types. The AlphaGenome Variant Impact score combines these predictions with AlphaMissense outputs, allowing the atlas to cover both protein-coding regions and the much larger non-coding portion of the genome.
The method turns an experimentally impossible search into a selection problem. Instead of testing billions of variants in a laboratory, a team can use the model to create a shorter list of changes that may affect gene expression, RNA splicing or other molecular processes. This does not remove experimentation; it changes its order. The source documents the atlas, its scale and its reported results. The interpretation that its central value lies in directing scarce laboratory time is Mateusz’s editorial conclusion, not a separate clinical finding reported by Google.
What the early applications show
Google describes several applications conducted with scientific collaborators. One involved a DNM1 variant whose predicted effect was subsequently tested experimentally. Another used data from more than 54,000 UK Biobank participants to search for links between variants and traits. The company reports that incorporating AlphaGenome scores in non-coding regions revealed additional associations beyond those found with simpler annotations. These are concrete case studies, but they do not establish identical performance across every disease, population or experimental setting.
The evidence has several levels. The DNM1 validation shows that a useful prediction can lead to a testable biological mechanism. The UK Biobank analysis demonstrates statistical utility in a large cohort, but an association alone does not establish causation or determine what a variant will do in one person. The results support using the atlas to prioritize candidates, particularly outside protein-coding regions. They do not support treating every high model score as a confirmed cause of disease.
AVI spans the whole genome
Most of the genome does not code for proteins, and those regions are harder to interpret with simple rules.
- Protein-coding regions2%
- Non-coding regions98%
What the atlas cannot determine
AlphaGenome predicts molecular effects from patterns learned in data; it does not directly observe a person’s future disease course. The consequence of a variant may depend on other genes, age, environment, ancestry and the relevant cell type. Google explicitly states that the atlas is not approved for clinical diagnosis. A score should therefore not determine a diagnosis, treatment, family-risk estimate or reproductive decision on its own. Those uses require independent evidence, qualified interpretation and the safeguards of an appropriate medical process.
The atlas also cannot eliminate the representativeness problem in biological data. Even a resource of exceptional breadth may perform unevenly for underrepresented populations, uncommon tissues or molecular processes that current assays capture poorly. A model can flag a relationship without supplying the complete mechanism behind it. Nine billion variants describes computational coverage, not nine billion verified discoveries. Keeping that distinction visible prevents the scale of the resource from being mistaken for certainty about each individual prediction.
A responsible way to use the result
Mateusz’s proposed workflow has four stages: prediction, prioritization, experiment and replication. The atlas first identifies variants worth attention; a biologist then compares them with the literature, population evidence and the context of the condition under study. The strongest candidates proceed to functional testing, and an important finding should be reproduced with an independent method or dataset. This sequence is not presented by Google as a formal clinical protocol. It is Mateusz’s practical model for turning the limitations stated in the sources into a safer research process.
The next evidence to watch includes peer-reviewed applications, independent comparisons with other methods and performance across populations and cell types. Prospective experiments will be especially valuable: researchers should record the prediction before seeing the laboratory result. Those tests can reveal how often a high score leads to a genuinely new and reproducible finding. Progress should therefore be judged less by the number of predictions generated and more by the share of prioritized candidates that survive successive levels of biological validation.
The hardest genomic problem is no longer reading the letters
Sequencing can identify where one person’s genome differs from a reference, but a list of differences does not explain their meaning. A variant might alter a protein, RNA splicing, chromatin accessibility or gene expression—or produce no measurable effect in the context being studied. Non-coding regions are especially difficult. They make up roughly 98% of the genome and regulate gene activity without the relatively direct rule that maps coding triplets to amino acids.
A laboratory cannot construct and test each of roughly nine billion possible single-letter substitutions separately. AlphaGenome Atlas changes the order of work: create a computational map of possible effects first, then direct experiments to selected points. This is a move from largely blind search to prioritization. It does not turn a prediction into a biological fact. The atlas is a hypothesis layer between sequence reading and expensive experimentation, not a substitute for the experiment.
Scale of the atlas
Four figures show the computational footprint and interpretability layer.
The model and the atlas solve different problems
AlphaGenome is a model that accepts a long DNA sequence—up to one million base pairs—and predicts many forms of molecular signal, with single-base resolution for most outputs. A researcher can query a selected region, variant and tissue context. The atlas is a precomputed layer covering every possible single-nucleotide variant in the human genome. Instead of invoking the model billions of times, users retrieve prepared results and filter them at cohort scale.
The distinction guides tool choice. Atlas data suits cohort searches, ranking very large variant lists and comparing prepared attributions. Direct model access is better for a bounded set of new regions or analyses needing specific output modalities. The official repository notes that ordinary model queries are unlikely to suit projects requiring more than one million predictions; the precomputed atlas exists precisely to make this scale tractable.
A prediction compares two versions of the same sequence
For a variant, a researcher can construct reference and alternate sequences and compare the predicted signals. If one letter changes a splice site, transcription-factor binding or expression, the difference between the two tracks becomes a candidate mechanism. Long context matters because a regulatory element may affect a gene far from the variant itself. Outputs for splicing, expression, chromatin and genomic contacts support a richer question than “does anything change?”: they suggest the molecular layer in which a change may occur.
The answer still depends on the selected cell type, tissue and modality. No predicted effect in one context does not mean no effect anywhere; a strong molecular signal does not establish that a person will develop disease. Other genes, environment, development and biological compensation intervene between DNA and phenotype. A model result should therefore be read as a contrast between alleles under a represented context, not as a direct measurement of an organism.
AVI compresses thousands of signals, while attributions reopen the score
Thousands of predictions for one variant are scientifically rich but difficult to rank. The AlphaGenome Variant Impact score reduces them to one allelic-resolution measure of potential impact. In coding regions it also incorporates protein-effect information from AlphaMissense, allowing a common ranking layer to span the roughly 2% of the genome that codes for proteins and the non-coding majority. Its primary purpose is to order candidates for investigation.
A single unexplained number would be dangerously tempting, so Atlas connects AVI to additive feature attributions. They indicate how categories such as splicing, expression, chromatin accessibility, conservation and protein effect contribute to the score. An attribution does not prove a mechanism, but it suggests which experiment should come next. A researcher can test RNA, promoter activity or a particular factor’s binding instead of treating a high rank as an indivisible model verdict.
Early studies demonstrate prioritization, not clinical performance
Google DeepMind describes a rare-disease collaboration in which AVI helped prioritize a previously overlooked variant in DNM1, a gene strongly associated with epileptic encephalopathy, after which the team investigated its effect experimentally. In another project, whole-genome data from more than 54,000 UK Biobank participants was analysed by grouping rare variants according to predicted molecular effect. The project account says this surfaced 22% more non-coding associations above statistical noise.
An analysis of body mass index narrowed hundreds of millions of non-coding variants to the one percent predicted to be most impactful and identified 19 regions for follow-up. These are examples of stronger search and experiment targeting, not a prospective diagnostic study. The results are reported in the atlas creators’ official material and still need assessment in full publications, independent replication and tests across other populations. The strongest conclusion concerns candidate selection, not readiness for patient-level decisions.
Genome-wide resolution is not biological completeness
The atlas covers single-nucleotide variants, so “every possible letter change” does not mean every kind of genomic alteration. Insertions, deletions, structural variants, copy-number changes and combinations of variants can require other methods. Precomputation also depends on a reference-genome version and the cell types and training evidence available to the model. Biology that is poorly represented in those data may remain poorly predicted even when genomic coordinates are covered uniformly.
A high AVI score indicates potential molecular impact, not automatically pathogenicity, effect direction, penetrance or clinical evidence strength. Benchmark leadership is an average result on defined datasets; it does not certify every variant. For a particular disorder, phenotype match, inheritance, population frequency, literature and functional evidence still matter. AlphaGenome’s official terms state directly that predictions are for theoretical modelling and research and must not guide clinical decisions.
A practical workflow runs from question to experiment
A sound project starts with a biological question, not a one-petabyte download. Define the gene, trait, tissue and variant criteria, then use Atlas rankings and attributions. Compare candidates with population resources, known mechanisms and independent predictors. For the resulting short list, choose a functional assay matched to the suggested layer—for example splicing, expression or regulatory activity. Negative findings should return to the workflow as information about model limits.
The portal supports exploration without code, while the API and published client enable programmatic analysis. Terms matter before work scales: Atlas and standard API access are provided free for non-commercial use, while commercial use follows a separate Google Cloud route. The resource’s lasting value will depend on versioning, reproducibility and validation outside its creators’ collaborations. Atlas may dramatically shorten a candidate list, but science accelerates only when a ranking leads to an experiment capable of telling the model it was wrong.
A good experiment tests a mechanism without sidelining under-described groups
A highly ranked variant begins an experimental decision; it does not automatically justify the most expensive assay. The attribution must first be translated into a mechanism that can be falsified: altered RNA splicing, regulatory activity, expression or another measurable process. The plan should state in advance which observation would support the prediction, which result would contradict it, which cell type and biological state are appropriate, and which controls would expose a failure in the assay itself. The most informative candidate is not always the highest-scoring one. A pair of variants with similar rankings but different predicted mechanisms may teach more, as may an uncertain variant for which separate methods imply different consequences. An experiment selected this way does more than confirm an individual hypothesis. It reveals where the model distinguishes biological processes successfully and where its signal needs correction.
Representation fairness belongs in the plan before candidates enter the queue. A team should document which populations, tissues, age ranges, sexes and disease contexts are well supported by available evidence and where coverage is weak or unknown. The aim is not artificial symmetry but preventing an average result from hiding systematically lower reliability for a particular group. Where stratification is scientifically and ethically justified, results should be examined by subgroup; where evidence is sparse, the correct response is wider uncertainty and priority for independent validation rather than manufactured confidence. Some laboratory capacity should remain available for signals from under-described groups and for negative results. Otherwise, ranking can repeatedly direct resources toward the settings for which the model already has the richest evidence, deepening the existing imbalance. Experimental choice should therefore combine mechanistic plausibility, evidence quality, information value and the consequences of uneven error.
Frequently asked questions
Can AlphaGenome Atlas diagnose a patient?
No. The official terms restrict predictions to modelling and research. A result can prioritize a variant, but clinical decisions require complete genetic assessment, phenotype evidence and independent validation.
What is the AVI score?
AlphaGenome Variant Impact is a single measure of a variant’s potential effect, derived from many molecular predictions and protein-effect information. Feature attributions indicate which biological layers contributed most to the ranking.
Why do non-coding regions matter?
They comprise roughly 98% of the genome and contain elements regulating gene activity. Variants there can alter expression or splicing, yet their consequences are usually harder to interpret than a change directly altering a protein.
How should a researcher use the atlas?
Begin with a defined question and tissue, rank variants, inspect attributions, compare independent resources and choose a matching functional assay. The ranking starts validation; it does not finish it.
