Skip to main content
model-family passport · Review date not recorded

Nucleotide Transformer

BERT-style foundation models for the human genome.

3/7Evidence fields documented
60-SECOND EVALUATION VIEW

What should a scientist know before using Nucleotide Transformer?

UnresolvedEvidence direction is incomplete or not yet resolved
Best suited forGeneration · Prediction
Evidence supportsPrimary links may be present, but BioAtlas does not claim a review date without a record-level timestamp.
Evidence does not establishUniversal superiority, therapeutic success, clinical utility or regulatory acceptance.
Major limitationPerformance depends on the evaluation dataset and operating conditions.
Current registry recordVersion history not yet curated1 recorded release · Review date not recorded. A newer version is not assumed to be universally better.

What it is

The Nucleotide Transformer family are large language models pretrained on thousands of genomes that transfer to many genomic prediction tasks with minimal fine-tuning. Built by InstaDeep (acquired by BioNTech), they helped establish self-supervised pretraining as a genomics workhorse.

Evidence trail

BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.

Sources3 connectedPrimary resources and normalized claims
Claims0 normalizedNo normalized claim yet
EntityNucleotide Transformermodel-family · Version history not yet curated
ReviewReview date not recordedReview date not claimed
ConclusionContext requiredAdd to an evaluation before operational use

Model passport

Entity typemodel-family
OrganizationInstaDeep (BioNTech)
Model family introducedNot normalized
AccessOpen source
Commercial useAllowed / verify checkpoint terms
DeploymentSelf-hosted
ComputeGPU recommended
Domainsgenomics
Biology → representation → computation → evidence

How Nucleotide Transformer represents biology

model-familygenomics

Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.

1 · Biological inputs
DNA sequence
2 · Input representation
Nucleotide / genomic tokens
3 · Internal representation
Genomic representation
4 · Architecture
Genomic foundation model
5 · Learning objective
Sequence modelling
6 · Output representation
Dense vectorsGenomic tracksSequence

Biological scale

Modalities & tasks

DNAGenerationPredictionRepresentation

Registry, claims and frontier intelligence

Versioned registry

Version history not yet curated

1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.

Explore version lineage →
Benchmark claim ledger

0 normalized claims

No task, dataset, split and metric claim has been normalized for this record yet.

Open claim intelligence →

Connected research frontiers

These records describe active research directions, not guaranteed capabilities of this model. Evidence stages and unresolved questions are preserved separately.

Genome understanding & design

Million-base regulatory variant prediction

Google DeepMind · 2026-01-28
Peer-reviewed capability

Can a single model predict how coding and non-coding variants alter expression, splicing, chromatin and regulatory binding over long genomic context?

Evidence boundary and unresolved questions

The model is a research predictor, not a personal-genome or clinical diagnostic system; tissue specificity and very long-range enhancer logic remain limitations.

  • How reliably do predictions transfer to rare cell states and patient contexts?
  • Can causal mechanisms be separated from learned correlations?
  • How should predictions be prospectively validated?
regulatory genomics · variant effects · non-coding DNA · splicingOpen frontier record →

Inputs and outputs

Inputs

DNA sequence

Outputs

Sequence predictionsEmbeddings or generated sequence

Scientific and technical profile

Scientific principles

Masked genomic language modelingTransfer learning

Technology

BERT-style Transformerk-mer / BPE tokenizationMulti-species pretraining
Ideas before algorithms

Scientific lineage

Explore all foundations

These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.

Genomics & cell systems

DNA as the hereditary transforming principle

Oswald Avery, Colin MacLeod & Maclyn McCarty

Genomics, variant interpretation, gene therapy and sequence foundation models depend on DNA being the durable molecular carrier of biological information.

Matched concepts: dna, genome, sequence
Genomics & cell systems

The DNA double helix and complementary base pairing

James Watson & Francis Crick

Sequence analysis, variant prediction, genome design and nucleic-acid therapeutics all rest on this structural logic.

Matched concepts: dna, genome, nucleotide
Genomics & cell systems

Reading the sequences of proteins and DNA

Frederick Sanger

Biological foundation models exist because proteins and genomes became readable, comparable and computable at scale.

Matched concepts: sequence, dna, genome
Computational intelligence

Transformer self-attention

Ashish Vaswani and colleagues

Protein, genome, molecule and single-cell foundation models use attention to learn dependencies across biological sequences and multimodal inputs.

Matched concepts: transformer, language model, sequence
Computational intelligence

Information, entropy and communication

Claude E. Shannon

Sequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.

Matched concepts: language model, sequence, token
Genomics & cell systems

X-ray evidence for the helical structure of DNA

Rosalind Franklin & Raymond Gosling

Structural genomics and sequence-to-structure reasoning began with experimentally grounded molecular geometry.

Matched concepts: dna, genome, sequence

Evaluation evidence

Dataset or evaluationNot yet curated
Task or metricNot yet extracted
Evidence statusPrimary paper linked; benchmark extraction pending
Open source ↗

BioAtlas has not yet extracted a structured benchmark claim for this record.

Known limitations

  • Performance depends on the evaluation dataset and operating conditions.
  • A structured benchmark claim has not yet been extracted for this record.
  • Outputs require task-specific scientific and experimental validation.

Milestones

Not normalized

InstaDeep was acquired by BioNTech.

Evidence

Open weights across several model sizes.