Skip to main content
PROTEIN / LIGAND MODEL EVALUATION · v1.0.0

Which model should I use for my protein–ligand problem?

BioAtlas finds candidate model families, pins exact versions, exposes benchmark claims, freezes one evaluation dataset and metric contract, dispatches only cleared compute adapters, and keeps uncertainty visible until the evidence supports a verdict.

01 · QUESTION

Declare the scientific problem before choosing the model.

public evaluation
02 · CANDIDATE MODELS + EXACT VERSIONS

Compare the model family that is actually relevant to protein–ligand structure.

Version selection is task-specific: for example, Chai-1 is kept for complex prediction rather than substituting Chai-2 design claims.

4 selected
no governed adapter

AlphaFold 3

All-atom biomolecular complex prediction

Registry
AlphaFold 2 / 3 · Google DeepMind
Exact version
AlphaFold 3 · 2024
Checkpoint
Server and released code/weights subject to current access terms
Compute
Evidence-only · no governed adapter
Broad molecular-entity support and strong structural benchmark lineage.BioAtlas tracks AlphaFold 3 evidence and version identity, but does not expose it as a cleared local execution adapter in the current compute ledger.
clearance required

Boltz-2

Open complex structure + affinity-capable lineage

Registry
Boltz-1 / Boltz-2 · MIT (Barzilay & Jaakkola labs)
Exact version
Boltz-2 · 2025
Checkpoint
Public project release
Compute
Boltz-2 Complex Prediction · g6.2xlarge
Commercially clearable BioAtlas adapter with joint structure/affinity workflow support.A signed project-specific Context of Use, immutable container/checkpoint identity and dataset clearance are still required before execution.
clearance required

Chai-1

Multimodal biomolecular complex prediction

Registry
Chai-1 / Chai-2 · Chai Discovery
Exact version
Chai-1 · 2024
Checkpoint
Research weights under project terms
Compute
Chai-1 Complex Prediction · g6.2xlarge
Commercially clearable BioAtlas adapter with ligand, nucleic-acid, template and restraint-aware complex prediction.The protein–ligand evaluation intentionally selects Chai-1 rather than treating Chai-2 generative-design claims as equivalent complex-prediction evidence.
no governed adapter

Protenix

Open AF3-class biomolecular complex prediction

Registry
Protenix · ByteDance
Exact version
Version history not yet curated
Checkpoint
Not normalized
Compute
Evidence-only · no governed adapter
Useful independent model family for comparative structure evaluation when an exact executable release is available.BioAtlas currently has no governed Protenix execution adapter in the compute ledger, so it must remain evidence-only until an exact artifact passes license/runtime acceptance.
03 · BENCHMARK CLAIMS

Inspect what has actually been claimed—and the caveats—before rerunning anything.

version-aware evidence

AlphaFold 3

Protein structure predictionCASP14 / complex evaluations · Structure accuracy and confidence

Top-performing CASP14 system; consult the primary paper for target-level metrics.

peer-reviewed · replication multiple-independent-usesCASP performance does not establish equal accuracy for every target class or drug-relevant complex.AlphaFold 2 and AlphaFold 3 require separate evaluation contexts.

Boltz-2

Biomolecular complex predictionPoseBusters and affinity benchmarks · Structure and affinity

Open evaluation reports structure prediction and later affinity capabilities.

developer-reported · replication partialBoltz-1 structure claims and Boltz-2 affinity claims should be separated by version.Affinity performance depends strongly on target family and split design.

Chai-1

Biomolecular complex predictionComplex and antibody-design evaluations · Structure / design performance

Developer-reported AF3-class complex-prediction performance.

developer-reported · replication partialPreprint and developer-reported comparisons require independent reproduction.Benchmark protocol and entity coverage determine comparability.

Protenix

No normalized model-specific claim is currently sufficient for this task. BioAtlas keeps this as an evidence gap.

04 · EVALUATION DATASET

Freeze one protocol for every candidate.

PLINDER · similarity-controlled protein–ligand evaluation
OOD / similarity-controlled splitPLINDER

Large protein-ligand interaction resource with similarity-aware evaluation splits for proteins, ligands, pockets and interactions.

Pose, interaction and task-specific structure metrics
Orthogonal physical-validity evaluationPoseBusters

Physical and chemical validity checks for generated or docked protein-ligand poses.

Geometry, stereochemistry, clashes and interaction-validity checks
05 · RUN COMPUTE

One governed evaluation contract, multiple eligible model adapters.

Estimate first if desired. Submission uses the existing BioAtlas compute gateway and will fail closed when clearance, immutable artifacts, checkpoint identity or dataset rights are missing.

2 executable selected

Define the scientific context, inspect versions/evidence, then estimate or run the same governed evaluation across eligible adapters.

AlphaFold 3no-governed-adapter

No execution adapter

Boltz-2clearance-required

g6.2xlarge · nominal 18 min

Chai-1clearance-required

g6.2xlarge · nominal 20 min

Protenixno-governed-adapter

No execution adapter

06 · NORMALIZED RESULTS

Metrics, failure cases, runtime, reproducibility, uncertainty and evidence quality.

Only enter or import metrics generated under the frozen protocol above. Blank fields deliberately produce “Insufficient evidence.”

6 required metrics
ModelPose successhigher · threshold 70%Physical validityhigher · threshold 90%Failure ratelower · threshold 15%Reproducibilityhigher · threshold 80%Unresolved uncertaintylower · threshold 25%Evidence qualityhigher · threshold 70%Failure cases
AlphaFold 3no-governed-adapter%%%%%%
Boltz-2clearance-required%%%%%%
Chai-1clearance-required%%%%%%
Protenixno-governed-adapter%%%%%%
07 · BIOATLAS EVALUATION

Suitability is a bounded conclusion, not a popularity ranking.

human-reviewable
AlphaFold 3Insufficient evidence

Required normalized metrics are incomplete · no-governed-adapter.

AlphaFold 3 · 2024 · PLINDER · similarity-controlled protein–ligand evaluation
Boltz-2Insufficient evidence

Required normalized metrics are incomplete · clearance-required.

Boltz-2 · 2025 · PLINDER · similarity-controlled protein–ligand evaluation
Chai-1Insufficient evidence

Required normalized metrics are incomplete · clearance-required.

Chai-1 · 2024 · PLINDER · similarity-controlled protein–ligand evaluation
ProtenixInsufficient evidence

Required normalized metrics are incomplete · no-governed-adapter.

Version history not yet curated · PLINDER · similarity-controlled protein–ligand evaluation
suitable suitable with limitations insufficient evidence failed required test