Skip to main content
Genome understanding & design · Peer-reviewed capability · 2026-01-28

Million-base regulatory variant prediction

Can a single model predict how coding and non-coding variants alter expression, splicing, chromatin and regulatory binding over long genomic context?

What researchers are trying

AlphaGenome processes up to one million DNA bases at single-base resolution and predicts thousands of regulatory tracks and variant effects.

Why it matters

Most disease-associated variants are non-coding. Better regulatory predictions could connect association signals to mechanisms, targets and experiments.

Evidence boundary

The model is a research predictor, not a personal-genome or clinical diagnostic system; tissue specificity and very long-range enhancer logic remain limitations.

Organizations represented

Google DeepMind

Demonstrated evidence

What has actually been shown.

  • Reported strong performance across sequence and variant-effect benchmarks and mechanistic case studies.

Unresolved questions

  • How reliably do predictions transfer to rare cell states and patient contexts?
  • Can causal mechanisms be separated from learned correlations?
  • How should predictions be prospectively validated?

Signals to watch next

  • Independent variant benchmarks
  • Disease-specific validation
  • Cell-state conditioning
  • Clinical governance
Connected evidence graph

Related BioAtlas model passports.

AlphaGenome

Reading the genome's 'dark matter' at base-pair resolution.

2/7 evidence fields documented

Enformer

Transformers that predict gene expression from raw sequence.

2/7 evidence fields documented

AlphaMissense

Classifying which missense mutations cause disease.

2/7 evidence fields documented
Primary and evaluation sources

Inspect the evidence directly.