Skip to main content
model-family passport · Review date not recorded

DNABERT-2

A BERT-style genomic foundation model with efficient DNA tokenization and multi-species pretraining.

4/7Evidence fields documented
60-SECOND EVALUATION VIEW

What should a scientist know before using DNABERT-2?

SupportedEvidence supports the stated context with explicit boundaries
Best suited forRepresentation · Prediction
Evidence supportsGenomic downstream tasks: Open evaluation
Evidence does not establishUniversal superiority, therapeutic success, clinical utility or regulatory acceptance.
Major limitationPerformance depends on the evaluation dataset and operating conditions.
Current registry recordVersion history not yet curated1 recorded release · Review date not recorded. A newer version is not assumed to be universally better.

What it is

DNABERT-2 is a general genomic language model for DNA representation and downstream prediction.

Evidence trail

BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.

Sources3 connectedPrimary resources and normalized claims
Claims1 normalizedGenomic sequence modelling
EntityDNABERT-2model-family · Version history not yet curated
ReviewReview date not recordedReview date not claimed
ConclusionContext requiredAdd to an evaluation before operational use

Model passport

Entity typemodel-family
OrganizationMulti-institution research team
Model family introducedNot normalized
AccessOpen source
Commercial useAllowed / verify checkpoint terms
DeploymentSelf-hosted
ComputeGPU recommended
Domainsgenomics
Biology → representation → computation → evidence

How DNABERT-2 represents biology

model-familygenomics

Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.

1 · Biological inputs
DNA sequence
2 · Input representation
Nucleotide / genomic tokens
3 · Internal representation
Contextual genomic embeddings
4 · Architecture
Genomic sequence model
5 · Learning objective
Self-supervised genomic sequence modelling
6 · Output representation
Dense vectorsScores

Biological scale

genomeregulatory-element

Modalities & tasks

DNARepresentationPrediction

Registry, claims and frontier intelligence

Versioned registry

Version history not yet curated

1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.

Explore version lineage →

Inputs and outputs

Inputs

DNA sequence

Outputs

DNA embeddingsTask predictions

Scientific and technical profile

Scientific principles

Genomic language modellingTransfer learning

Technology

Transformer encoderDNA subword tokenization
Ideas before algorithms

Scientific lineage

Explore all foundations

These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.

Computational intelligence

Information, entropy and communication

Claude E. Shannon

Sequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.

Matched concepts: language model, sequence, representation
Computational intelligence

Transformer self-attention

Ashish Vaswani and colleagues

Protein, genome, molecule and single-cell foundation models use attention to learn dependencies across biological sequences and multimodal inputs.

Matched concepts: transformer, language model, sequence
Genomics & cell systems

DNA as the hereditary transforming principle

Oswald Avery, Colin MacLeod & Maclyn McCarty

Genomics, variant interpretation, gene therapy and sequence foundation models depend on DNA being the durable molecular carrier of biological information.

Matched concepts: dna, sequence
Genomics & cell systems

X-ray evidence for the helical structure of DNA

Rosalind Franklin & Raymond Gosling

Structural genomics and sequence-to-structure reasoning began with experimentally grounded molecular geometry.

Matched concepts: dna, sequence
Genomics & cell systems

Reading the sequences of proteins and DNA

Frederick Sanger

Biological foundation models exist because proteins and genomes became readable, comparable and computable at scale.

Matched concepts: sequence, dna

Evaluation evidence

Dataset or evaluationGenomic downstream tasks
Task or metricGeneral genomic representation
Evidence statusOpen evaluation
Open source ↗

Task-specific evidence only; not comparable as a universal leaderboard score.

Genomic sequence modelling

Genomic downstream tasks

Version history not yet curated · Split details not yet normalized
developer-reported

A structured benchmark claim is recorded; consult the linked source for numeric values and protocol details.

Claim caveats
  • Protocol, split and implementation details must match before comparing this claim with another result.

Known limitations

  • Performance depends on the evaluation dataset and operating conditions.
  • Task-specific benchmark results should not be compared across unlike domains.
  • Outputs require task-specific scientific and experimental validation.

Milestones

Not normalized

A major baseline in current genomic-FM benchmarking.