What it is
Profluent trains large protein language models (the ProGen line, up to ProGen3) to design functional proteins across families. It made headlines by using AI to generate OpenCRISPR-1 — a novel, working CRISPR gene editor distinct from natural Cas9 — and releasing it openly to accelerate the field.
Evidence trail
BioAtlas keeps the path from source to decision visible. A connection records provenance; it does not imply that evidence is sufficient for every context.
Model passport
How Profluent (ProGen3 / OpenCRISPR) represents biology
Category is navigation. These fields describe the model-specific computational transformation and deliberately override broad category defaults.
Biological scale
Modalities & tasks
Registry, claims and frontier intelligence
Version history not yet curated
1 version record · release year not yet normalized. Model-family identity remains separate from capability and access changes.
Explore version lineage →0 normalized claims
No task, dataset, split and metric claim has been normalized for this record yet.
Open claim intelligence →0 connected frontiers
No frontier-research record currently connects to this model.
Inspect research horizon →Inputs and outputs
Inputs
Project-specific biological dataOutputs
Models, evidence or candidatesScientific and technical profile
Scientific principles
Technology
Scientific lineage
These are transparent concept matches—not claims that one scientist alone caused this model. Each connection is based on the model’s recorded domain, scientific principles, technical terms or an explicit lineage link.
Programmable CRISPR–Cas genome editing
Jennifer A. Doudna & Emmanuelle CharpentierCRISPR enables target validation, disease models, perturbation atlases, functional genomics and gene-editing therapeutics.
Information, entropy and communication
Claude E. ShannonSequence modelling, cross-entropy training, language models, mutual information and representation learning all use Shannon’s framework.
Rational antimetabolite drug design
Gertrude B. Elion & George H. HitchingsMechanism-based design, pathway selectivity and iterative medicinal chemistry are direct descendants of this strategy.
Reading the sequences of proteins and DNA
Frederick SangerBiological foundation models exist because proteins and genomes became readable, comparable and computable at scale.
The central dogma and directional information transfer
Francis CrickMulti-omic models and sequence foundation models connect genotype, transcript and protein through this information-flow framework.
Protein sequence databases, evolutionary substitution matrices and computational comparison
Margaret Oakley DayhoffProtein language models, homology inference, multiple-sequence alignments and evolutionary priors inherit her conversion of sequence biology into computable data.
Evaluation evidence
BioAtlas has not yet extracted a structured benchmark claim for this record.
Known limitations
- Performance depends on the evaluation dataset and operating conditions.
- A structured benchmark claim has not yet been extracted for this record.
- Outputs require task-specific scientific and experimental validation.
Milestones
OpenCRISPR-1 was AI-designed and open-sourced.
Founder authored the original ProGen work.