What it is
Geneformer is pretrained on tens of millions of single-cell transcriptomes to learn a context-aware model of gene networks, enabling predictions about regulation and disease with limited task-specific data. It was among the first single-cell foundation models used to nominate therapeutic targets.
Evidence trail
BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.
Model passport
How Geneformer represents biology
Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.
Biological scale
Modalities & tasks
Registry, claims and frontier intelligence
Version history not yet curated
1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.
Explore version lineage →0 normalized claims
No task, dataset, split and metric claim has been normalized for this record yet.
Open claim intelligence →1 connected frontier
Virtual cells · Recent preprint
Inspect research horizon →Connected research frontiers
These records describe active research directions, not guaranteed capabilities of this model. Evidence stages and unresolved questions are preserved separately.
Virtual cells that predict perturbation response
Arc Institute · Virtual Cell research community · 2026-04-30Can models forecast how cell populations respond to unseen drugs, gene edits, cytokines and environmental changes across biological contexts?
Evidence boundary and unresolved questions
Recent strict evaluations show marked performance drops under unseen contexts and metric-dependent rankings; simple baselines remain competitive on some global trends.
- Can models recover perturbation-specific mechanisms rather than average expression shifts?
- How should cell distributions, dose and time be represented?
- Which metrics predict prospective experimental usefulness?
virtual cells · perturbation · single cell · OOD generalization · world modelsOpen frontier record →Inputs and outputs
Inputs
DNA sequenceOutputs
Sequence predictionsEmbeddings or generated sequenceScientific and technical profile
Scientific principles
Technology
Scientific lineage
These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.
Transformer self-attention
Ashish Vaswani and colleaguesProtein, genome, molecule and single-cell foundation models use attention to learn dependencies across biological sequences and multimodal inputs.
The central dogma and directional information transfer
Francis CrickMulti-omic models and sequence foundation models connect genotype, transcript and protein through this information-flow framework.
DNA as the hereditary transforming principle
Oswald Avery, Colin MacLeod & Maclyn McCartyGenomics, variant interpretation, gene therapy and sequence foundation models depend on DNA being the durable molecular carrier of biological information.
Information, entropy and communication
Claude E. ShannonSequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.
The DNA double helix and complementary base pairing
James Watson & Francis CrickSequence analysis, variant prediction, genome design and nucleic-acid therapeutics all rest on this structural logic.
X-ray evidence for the helical structure of DNA
Rosalind Franklin & Raymond GoslingStructural genomics and sequence-to-structure reasoning began with experimentally grounded molecular geometry.
Evaluation evidence
BioAtlas has not yet extracted a structured benchmark claim for this record.
Known limitations
- Performance depends on the evaluation dataset and operating conditions.
- A structured benchmark claim has not yet been extracted for this record.
- Outputs require task-specific scientific and experimental validation.
Milestones
Trained on ~30M single-cell transcriptomes.
Used to identify candidate cardiomyopathy targets.