Skip to main content
model-family passport · Review date not recorded

ProGen / ProGen2 / ProGen3

GPT-style language models that write functional proteins.

3/7Evidence fields documented
60-SECOND EVALUATION VIEW

What should a scientist know before using ProGen / ProGen2 / ProGen3?

UnresolvedEvidence direction is incomplete or not yet resolved
Best suited forGeneration · Optimization
Evidence supportsPrimary links may be present, but BioAtlas does not claim a review date without a record-level timestamp.
Evidence does not establishUniversal superiority, therapeutic success, clinical utility or regulatory acceptance.
Major limitationPerformance depends on the evaluation dataset and operating conditions.
Current registry recordVersion history not yet curated1 recorded release · Review date not recorded. A newer version is not assumed to be universally better.

What it is

ProGen showed that an autoregressive language model, prompted with a functional 'tag', can generate artificial proteins (e.g. novel lysozymes) that are actually active in the lab. The line matured at Profluent, which used protein LMs to design OpenCRISPR-1, an open-sourced, AI-generated gene editor.

Evidence trail

BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.

Sources3 connectedPrimary resources and normalized claims
Claims0 normalizedNo normalized claim yet
EntityProGen / ProGen2 / ProGen3model-family · Version history not yet curated
ReviewReview date not recordedReview date not claimed
ConclusionContext requiredAdd to an evaluation before operational use

Model passport

Entity typemodel-family
OrganizationSalesforce Research → Profluent
Model family introducedNot normalized
AccessLimited open access
Commercial useAllowed / verify checkpoint terms
DeploymentHybrid
ComputeGPU recommended
Domainsdesign
Biology → representation → computation → evidence

How ProGen / ProGen2 / ProGen3 represents biology

model-familydesign

Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.

1 · Biological inputs
Target structure or design objective
2 · Input representation
Sequence and/or 3D geometry
3 · Internal representation
Generative design representation
4 · Architecture
Generative biological model
5 · Learning objective
Conditional generation
6 · Output representation
Sequence3D coordinates

Biological scale

Modalities & tasks

ProteinGenerationOptimization

Registry, claims and frontier intelligence

Versioned registry

Version history not yet curated

1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.

Explore version lineage →
Benchmark claim ledger

0 normalized claims

No task, dataset, split and metric claim has been normalized for this record yet.

Open claim intelligence →

Connected research frontiers

These records describe active research directions, not guaranteed capabilities of this model. Evidence stages and unresolved questions are preserved separately.

Genome understanding & design

Genome-scale generative biology

Arc Institute · Stanford · NVIDIA · 2026-03-01
Peer-reviewed capability

Can a foundation model read, predict and design biological sequence continuously from single nucleotides to megabase-scale genomes?

Evidence boundary and unresolved questions

Generative plausibility is not equivalent to biological viability, function or safety. Long generated sequences require extensive synthesis, containment and functional review.

  • What biological constraints are learned versus memorized?
  • How should whole-genome designs be evaluated before synthesis?
  • Can mechanistic interpretability keep pace with model scale?
genome foundation model · long context · sequence design · biosafetyOpen frontier record →
Generative biomolecular design

Multimodal protein programming

EvolutionaryScale · 2025-01-16
Peer-reviewed capability

Can one generative model reason jointly over protein sequence, structure and function and create functional proteins from mixed prompts?

Evidence boundary and unresolved questions

One striking protein demonstration does not establish general success across enzymes, therapeutics or complex multi-objective design tasks.

  • How frequently do generated functions survive experimental testing?
  • Can the model optimize potency, stability and safety together?
  • How should synthetic training labels affect confidence?
multimodal · protein language model · function generation · synthetic biologyOpen frontier record →
Programmable genome editing

Bridge-RNA programmable DNA recombination

Arc Institute · UC Berkeley · Stanford · 2024-06-26
Peer-reviewed capability

Can RNA programmably specify both target and donor DNA to insert, excise or invert large sequences without relying on conventional CRISPR cutting and repair?

Evidence boundary and unresolved questions

The original 2024 work was early-stage and bacterial. Efficiency, specificity, delivery and control in mammalian cells require separate validation.

  • Can the system work efficiently and specifically in human cells?
  • How are off-target recombination and repeated sequences controlled?
  • Can delivery support therapeutically relevant tissues and cargo sizes?
genome editing · bridge RNA · recombinase · large DNA editsOpen frontier record →

Inputs and outputs

Inputs

Target structure or design objective

Outputs

Designed sequencesCandidate structures

Scientific and technical profile

Scientific principles

Autoregressive language modelingControllable generationSequence-to-function

Technology

GPT-style decoderControl tags / conditioningScaled protein corpora
Ideas before algorithms

Scientific lineage

Explore all foundations

These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.

Computational intelligence

Information, entropy and communication

Claude E. Shannon

Sequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.

Matched concepts: language model, sequence
Structural biology

Anfinsen’s dogma—the thermodynamic hypothesis

Christian B. Anfinsen

Protein structure prediction, inverse folding and generative protein design all assume that sequence strongly constrains structure and function.

Matched concepts: sequence, protein
Computational intelligence

Transformer self-attention

Ashish Vaswani and colleagues

Protein, genome, molecule and single-cell foundation models use attention to learn dependencies across biological sequences and multimodal inputs.

Matched concepts: language model, sequence

Evaluation evidence

Dataset or evaluationNot yet curated
Task or metricNot yet extracted
Evidence statusPrimary paper linked; benchmark extraction pending
Open source ↗

BioAtlas has not yet extracted a structured benchmark claim for this record.

Known limitations

  • Performance depends on the evaluation dataset and operating conditions.
  • A structured benchmark claim has not yet been extracted for this record.
  • Outputs require task-specific scientific and experimental validation.

Milestones

Not normalized

ProGen lysozymes were experimentally active.

Evidence

Profluent's OpenCRISPR-1 was AI-designed and open-sourced.