Skip to main content
Evidence explorer

Benchmarks with provenance, scope, realism and explicit gaps.

Trace structured evaluations to their task, dataset, split, biological representation, evidence strength and primary source. BioAtlas separates benchmark suites from individual model claims and distinguishes public, proprietary, retrospective, out-of-distribution, prospective, experimental and clinical evidence instead of blending everything into a single leaderboard.

Benchmark landscape

Evaluation frameworks BioAtlas tracks

Scientific AI should be compared at the benchmark-suite level as well as the model level, with pipeline stage, access status, split realism, metric family and provenance made explicit.

26benchmark suites / families tracked
OpenAI and life-science contributors

LifeSciBench

Tracked
AI scientist / reasoningProspective / blind evaluation

Expert-authored life-science workflow tasks spanning evidence handling, analysis, experimental design, reasoning, validation, translation and scientific communication.

Scientific reasoningEvidence handlingAnalysisExperimental designTranslation
Metrics
Rubric-based task performance across expert-authored scientific workflows
Access
Public benchmark description and evaluation materials
Evidence
Independent expert-authored workflow benchmark
Evidence strength
Prospective / blind evaluation
BioAtlas use
Evaluate scientific-agent usefulness on realistic life-science work, not only question answering or recall.
Insilico Medicine

ScienceAI Bench

Tracked
AI scientist / reasoningRetrospective random split

Broad scientific reasoning across biology, chemistry, longevity, materials and related scientific domains.

Scientific reasoningBiologyChemistryLongevityMaterials
Metrics
Task-dependent classification, regression and reasoning metrics
Access
Public leaderboard / explorer
Evidence
Developer-maintained benchmark portal
Evidence strength
Retrospective random split
BioAtlas use
Cross-domain scientific-reasoning reference; keep task-level results separate from model evidence records.
Research community

SDABench

Tracked
AI scientist / reasoningRetrospective random split

Scientific data-analysis reasoning across descriptive, exploratory, inferential, predictive, causal and mechanistic analysis.

Scientific reasoningData analysisCausal inferenceMechanistic reasoning
Metrics
Task-appropriate reasoning and analysis accuracy
Access
Research benchmark / paper
Evidence
Research benchmark
Evidence strength
Retrospective random split
BioAtlas use
Prevents scientific reasoning from being collapsed into one undifferentiated score.
Research community

DrugPlayGround

Tracked
AI scientist / reasoningRetrospective random split

LLM-oriented drug-discovery reasoning across molecular properties, synergy, drug-protein interactions and perturbation responses.

Drug discovery reasoningChemistryDrug-target interactionPerturbation
Metrics
Task accuracy plus explanation / reasoning quality
Access
Research benchmark / paper
Evidence
Research benchmark
Evidence strength
Retrospective random split
BioAtlas use
Separates domain reasoning quality from conventional QSAR or structure-prediction metrics.
Insilico Medicine

Drug Discovery Benchmark

Tracked
Therapeutics-wideMixed / benchmark-dependent

Unified evaluation across drug-discovery and development tasks, mixing curated public datasets with proprietary evaluations.

Target discoveryChemistryADMETBiologicsClinical development
Metrics
Representative normalized task metrics plus rank-based aggregation
Access
Public leaderboard; underlying access varies by benchmark
Evidence
Developer-maintained benchmark portal
Evidence strength
Mixed / benchmark-dependent
BioAtlas use
Reference pipeline coverage and scoring design while preserving public/proprietary provenance at the individual benchmark level.
Insilico Medicine

Insilico Bench

Tracked
Therapeutics-wideOpaque / proprietary evaluation

Proprietary internal evaluations spanning scientific and drug-discovery tasks.

BiologyChemistryClinical developmentLongevityMaterials
Metrics
Task-dependent metrics aggregated into an Insilico Score
Access
Public leaderboard; benchmark contents partly proprietary
Evidence
Developer-reported proprietary benchmark
Evidence strength
Opaque / proprietary evaluation
BioAtlas use
Track as proprietary benchmark evidence; never present opaque internal scores as independently reproduced results.
TDC community

Therapeutics Data Commons (TDC)

Tracked
Therapeutics-wideMixed: random, group and temporal splits

Open benchmark ecosystem for therapeutic ML tasks spanning ADMET, efficacy, drug-target interaction, drug combinations and related discovery problems.

ADMETProperty predictionDrug-target interactionDrug combinationDrug discovery
Metrics
Dataset-appropriate classification, regression and ranking metrics
Access
Open public datasets and leaderboards
Evidence
Independent public benchmark ecosystem
Evidence strength
Mixed: random, group and temporal splits
BioAtlas use
Independent public anchor for vendor/model claims and a source of realistic split designs including temporal generalization.
Polaris consortium / industry collaborators

Polaris

Tracked
Therapeutics-wideMixed: scaffold, group and OOD-capable splits

Reproducible drug-discovery benchmark platform that packages dataset, split, metric and evaluation logic as explicit benchmark objects.

ADMEMolecular propertyStructureDrug discovery
Metrics
Benchmark-defined metrics with standardized splits and evaluation logic
Access
Public benchmark hub
Evidence
Independent public benchmark platform
Evidence strength
Mixed: scaffold, group and OOD-capable splits
BioAtlas use
Strong architectural reference for reproducible benchmark records: dataset + split + metric + evaluation logic.
BioBenchmarks community

BioBenchmarks

Tracked
Therapeutics-wideMixed: retrospective through clinical

Registry of biological and drug-discovery benchmarks, initiatives, organizations and experimental-validation status.

Benchmark registryExperimental validationProspective evaluationClinical evidence
Metrics
Registry-level metadata rather than one universal model metric
Access
Public benchmark registry
Evidence
Independent benchmark intelligence registry
Evidence strength
Mixed: retrospective through clinical
BioAtlas use
Competitive reference for benchmark discovery; BioAtlas differentiates via provenance, leakage risk, applicability and decision relevance.
Research community

MoleculeNet

Tracked
Molecules / ADMET / chemistryRandom / scaffold split depending on task

Canonical benchmark collection for molecular machine learning across quantum, physical chemistry, biophysics and physiology tasks.

Molecular propertyQSARADMET
Metrics
ROC-AUC, RMSE, MAE and dataset-specific metrics
Access
Public datasets / benchmark definitions
Evidence
Peer-reviewed public benchmark
Evidence strength
Random / scaffold split depending on task
BioAtlas use
Baseline historical reference; flag when newer OOD or temporal benchmarks provide stronger evidence.
Research community

MuMO-Instruct

Tracked
Molecules / ADMET / chemistryConstraint-based retrospective evaluation

Multi-objective molecular optimization under structural and property constraints.

Molecular optimizationMedicinal chemistry
Metrics
Success rate plus task-specific property and similarity constraints
Access
Public research benchmark
Evidence
Research benchmark
Evidence strength
Constraint-based retrospective evaluation
BioAtlas use
Track separately from ADMET prediction because constrained molecular optimization tests a distinct capability.
Research community

FGBench

Tracked
Molecules / ADMET / chemistryRetrospective evaluation

Functional-group chemistry reasoning and quantitative comparison tasks.

Chemistry reasoningMedicinal chemistry
Metrics
Accuracy and error metrics depending on task
Access
Research benchmark / reported evaluation
Evidence
Research benchmark
Evidence strength
Retrospective evaluation
BioAtlas use
Adds chemistry-reasoning coverage that molecular property prediction alone does not capture.
ProteinGym community

ProteinGym

Tracked
Protein / biologicsCross-assay / protein generalization

Large benchmark for protein mutation-effect and fitness prediction across deep mutational scanning and clinical variants.

Protein fitnessVariant effectClinical variants
Metrics
Rank correlation and task-specific mutation-effect metrics
Access
Public datasets and leaderboard
Evidence
Independent public benchmark
Evidence strength
Cross-assay / protein generalization
BioAtlas use
Core benchmark family for protein language/foundation models; expose biological strata instead of only a global score.
Research community

ProteinBench

Tracked
Protein / biologicsMulti-dimensional retrospective evaluation

Holistic protein-foundation-model evaluation across quality, novelty, diversity and robustness.

Protein generationProtein representationRobustness
Metrics
Quality, novelty, diversity and robustness dimensions
Access
Public benchmark / leaderboard
Evidence
Research benchmark
Evidence strength
Multi-dimensional retrospective evaluation
BioAtlas use
Avoids ranking protein models on a single axis and captures robustness and novelty explicitly.
Research community

PepBenchmark

Tracked
Protein / biologicsStandardized retrospective evaluation

Peptide-therapeutics benchmark spanning canonical and noncanonical peptide prediction tasks.

Peptide therapeuticsMolecular representationProperty prediction
Metrics
Dataset-specific classification and regression metrics
Access
Research benchmark / paper
Evidence
Research benchmark
Evidence strength
Standardized retrospective evaluation
BioAtlas use
Adds peptide-specific evaluation coverage that is not captured by small-molecule or general protein benchmarks.
Critical Assessment of protein Structure Prediction

CASP

Tracked
Structure / dockingProspective / blind evaluation

Independent blind community assessment of protein-structure prediction and related structural-biology tasks.

Protein structureComplex structureBlind assessment
Metrics
Structure-accuracy metrics defined by CASP category
Access
Public challenge results and evaluation materials
Evidence
Independent blind community benchmark
Evidence strength
Prospective / blind evaluation
BioAtlas use
Treat as a high-evidence structural benchmark because targets are assessed blind rather than retrospectively selected.
Research community

PSBench

Tracked
Structure / dockingBlind-challenge-derived evaluation

Large structural-model quality benchmark derived from CASP15/CASP16 models with global, local and interface-level annotations.

Structure qualityInterface qualityModel assessment
Metrics
Global, local and interface structural-quality metrics
Access
Research dataset / paper
Evidence
CASP-derived public benchmark
Evidence strength
Blind-challenge-derived evaluation
BioAtlas use
Useful for benchmarking model-quality assessment and confidence estimation against challenging blind-prediction outputs.
PLINDER community

PLINDER

Tracked
Structure / dockingOOD / similarity-controlled split

Large protein-ligand interaction resource with similarity-aware evaluation splits for proteins, ligands, pockets and interactions.

DockingProtein-ligand interactionPose predictionAffinity
Metrics
Pose, interaction and task-specific structure metrics
Access
Open dataset and tooling
Evidence
Independent public benchmark resource
Evidence strength
OOD / similarity-controlled split
BioAtlas use
Top-priority docking benchmark because it supports novelty-aware evaluation instead of easy random splits.
PLINDER community

Runs N' Poses

Tracked
Structure / dockingOut-of-distribution target or domain

Protein-ligand co-folding benchmark designed to test true zero-shot generalization and possible training-set memorization.

DockingCo-foldingGeneralization auditLeakage audit
Metrics
Pose accuracy, pocket accuracy, training similarity and physical-validity checks
Access
Open benchmark repository
Evidence
Independent public generalization benchmark
Evidence strength
Out-of-distribution target or domain
BioAtlas use
Expose memorization/generalization risk directly alongside headline pose accuracy.
Research community

PoseBench

Tracked
Structure / dockingOOD / realistic structural evaluation

Realistic protein-ligand structure prediction including apo-to-holo, unknown-pocket and multi-ligand settings.

DockingApo-to-holoMulti-ligandUnknown pocket
Metrics
Pose accuracy, physical validity and setting-specific structure metrics
Access
Public research benchmark
Evidence
Independent research benchmark
Evidence strength
OOD / realistic structural evaluation
BioAtlas use
Separate performance on familiar holo structures from harder deployment-like structural settings.
Research community

DockGen

Tracked
Structure / dockingOut-of-distribution target or domain

Docking generalization benchmark built around previously unseen protein classes and binding pockets.

DockingOOD dockingNovel pockets
Metrics
Pose-recovery and docking-generalization metrics
Access
Public research benchmark
Evidence
Independent research benchmark
Evidence strength
Out-of-distribution target or domain
BioAtlas use
Explicitly distinguish ordinary docking accuracy from generalization to novel targets and pockets.
Research community

PoseBusters

Tracked
Structure / dockingOrthogonal physical-validity evaluation

Physical and chemical validity checks for generated or docked protein-ligand poses.

DockingPose validationPhysical plausibility
Metrics
Geometry, stereochemistry, clashes and interaction-validity checks
Access
Open benchmark / toolkit
Evidence
Independent public validation toolkit
Evidence strength
Orthogonal physical-validity evaluation
BioAtlas use
Use as an orthogonal validity layer so a low-RMSD pose is not automatically treated as physically credible.
Insilico Medicine

TargetBench 1.0

Tracked
Cell / translational / clinicalRetrospective translational benchmark

Disease-specific therapeutic target-identification benchmarking with clinical retrieval and translational-readiness dimensions.

Target identificationDisease biologyTarget validation
Metrics
Clinical-target retrieval plus druggability, structure availability, repurposing and experimental-readiness measures
Access
Public framework / published methodology
Evidence
Published / developer-maintained benchmark framework
Evidence strength
Retrospective translational benchmark
BioAtlas use
High-value target-discovery reference because retrieval and downstream actionability are scored separately.
Insilico Medicine

ClinBench

Tracked
Cell / translational / clinicalRetrospective clinical prediction

Clinical-development benchmark for predicting trial outcomes, including Phase II outcome prediction.

Clinical developmentTrial outcome prediction
Metrics
F1 and task-specific classification metrics
Access
Proprietary / reported through MMAI Gym materials
Evidence
Developer-reported proprietary benchmark
Evidence strength
Retrospective clinical prediction
BioAtlas use
Track with explicit proprietary status and without unsupported independent-reproduction claims.
Insilico Medicine and collaborators

LongevityBench

Tracked
Cell / translational / clinicalRetrospective biological evaluation

Aging-biology reasoning from low-level biodata to phenotype-level conclusions.

LongevityAging biologyPhenotype prediction
Metrics
Task-specific classification, ranking and regression metrics
Access
Publication / benchmark materials
Evidence
Research benchmark
Evidence strength
Retrospective biological evaluation
BioAtlas use
Evaluate whether scientific models can reason over aging biology rather than only recall textual facts.
Research community

scDrugMap

Tracked
Cell / translational / clinicalCross-dataset / zero-shot evaluation

Single-cell foundation-model benchmark for drug-response and perturbation prediction across pooled, cross-dataset, fine-tuned and zero-shot settings.

Single-cellDrug responsePerturbationZero-shot evaluation
Metrics
Prediction metrics across pooled, transfer and zero-shot scenarios
Access
Research benchmark / paper
Evidence
Independent research benchmark
Evidence strength
Cross-dataset / zero-shot evaluation
BioAtlas use
Adds deployment-relevant cell-response evaluation and distinguishes ordinary pooled performance from cross-dataset transfer.
Model-level evidence

Task-specific records

A curated record means a task-specific benchmark claim has been extracted with source context; it does not mean the model is universally better.

86 records
ModelDomainCurationEvaluation datasetTask / metricEvidence typeCoverageSource
AlphaFold 2 / 3Google DeepMindStructure PredictionCuratedCASP14 / complex evaluationsStructure accuracy and confidencePeer-reviewed4/7Primary ↗
RoseTTAFold / All-AtomInstitute for Protein Design, UWStructure PredictionCuratedCASP14 / complex modellingStructure predictionPeer-reviewed4/7Primary ↗
OpenFoldOpenFold Consortium / ColumbiaStructure PredictionPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
ESMFoldMeta AI (FAIR)Structure PredictionPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
Chai-1 / Chai-2Chai DiscoveryStructure PredictionCuratedComplex and antibody-design evaluationsStructure / design performancePreprint / developer-reported4/7Primary ↗
Boltz-1 / Boltz-2MIT (Barzilay & Jaakkola labs)Structure PredictionCuratedPoseBusters and affinity benchmarksStructure and affinityPreprint / open evaluation4/7Primary ↗
ProtenixByteDanceStructure PredictionPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated2/7Primary ↗
RFdiffusionInstitute for Protein Design, UWProtein & Binder DesignCuratedExperimental binder validationDe novo protein designPeer-reviewed + experimental5/7Primary ↗
ProteinMPNN / LigandMPNNInstitute for Protein Design, UWProtein & Binder DesignPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
ESM3EvolutionaryScale (now CZ Biohub)Protein & Binder DesignCuratedGenerative protein evaluationsSequence, structure and functionPeer-reviewed / developer-reported4/7Primary ↗
ChromaGenerate BiomedicinesProtein & Binder DesignPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
ProGen / ProGen2 / ProGen3Salesforce Research → ProfluentProtein & Binder DesignPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
Latent-XLatent LabsProtein & Binder DesignPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
AlphaProteoGoogle DeepMindProtein & Binder DesignPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
NVIDIA BioNeMoNVIDIASmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated2/7Primary ↗
Chemistry42 / Pharma.AIInsilico MedicineSmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
NeuralPLexer / EnchantIambic TherapeuticsSmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
GEMSGenesis TherapeuticsSmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Numerion AI Chemistry PlatformNumerion LabsSmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
DiffDockMIT (Barzilay & Jaakkola labs)Small-Molecule & ChemistryCuratedPDBBindTop-ranked docking posePeer-reviewed4/7Primary ↗
Schrödinger PlatformSchrödinger, Inc.Small-Molecule & ChemistryPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Evo / Evo 2Arc Institute + Stanford + NVIDIAGenomics, DNA & RNACuratedGenomic sequence evaluationsDNA/RNA/protein generationPeer-reviewed4/7Primary ↗
AlphaGenomeGoogle DeepMindGenomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
EnformerGoogle DeepMindGenomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
AlphaMissenseGoogle DeepMindGenomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
Nucleotide TransformerInstaDeep (BioNTech)Genomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
GeneformerBroad InstituteGenomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
HyenaDNAStanford (Hazy Research)Genomics, DNA & RNAPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
BigRNADeep GenomicsGenomics, DNA & RNAPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
State / StackArc InstituteVirtual Cells & Single-CellCuratedPerturbation prediction datasetsCell-state predictionPreprint / open evaluation4/7Primary ↗
CZI Virtual Cells (rBio, TranscriptFormer)Chan Zuckerberg InitiativeVirtual Cells & Single-CellPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
scGPTUniversity of Toronto (Bo Wang Lab)Virtual Cells & Single-CellCuratedSingle-cell downstream tasksCell representationPeer-reviewed4/7Primary ↗
Cell2Sentence (C2S-Scale)Google Research + YaleVirtual Cells & Single-CellPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
scFoundationTsinghua / BioMapVirtual Cells & Single-CellPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
Absci — Integrated Drug CreationAbsciAntibodies & BiologicsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Nabla Bio — JAMNabla BioAntibodies & BiologicsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
IgLM / AntiBERTyJohns Hopkins (Gray Lab)Antibodies & BiologicsCuratedAntibody sequence evaluationsAntibody language modellingPeer-reviewed4/7Primary ↗
BigHat Biosciences — MillinerBigHat BiosciencesAntibodies & BiologicsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
LabGenius — EVALabGeniusAntibodies & BiologicsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Recursion (LOWE + Phenomics)Recursion PharmaceuticalsPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Basecamp Research (EDEN / BaseFold)Basecamp ResearchPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Cradle BioCradlePlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Profluent (ProGen3 / OpenCRISPR)Profluent BioPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Ginkgo Bioworks (Datapoints / AI Models)Ginkgo BioworksPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
OwkinOwkinPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
BenevolentAIBenevolentAIPlatforms, Data & InfraPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Isomorphic Labs (IsoDDE)Isomorphic Labs (Alphabet)AI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Xaira TherapeuticsXaira TherapeuticsAI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Generate:BiomedicinesGenerate:BiomedicinesAI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
ExscientiaRecursion PharmaceuticalsAI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Valence Labs (MolGPS / LOWE)Recursion PharmaceuticalsAI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
insitroinsitroAI-Native Discovery Cos.PendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
BoltzGenMIT / Boltz teamProtein & Binder DesignCuratedBinder-design evaluationsDesign success and structural qualityDeveloper / experimental reports2/7Primary ↗
RFantibodyInstitute for Protein Design, UWAntibodies & BiologicsCuratedExperimental antibody designEpitope-specific de novo designPeer-reviewed + experimental4/7Primary ↗
ESM-2Meta AI (FAIR)Protein Foundation & RepresentationCuratedProtein representation and structure evaluationsTransfer / language-model representationPeer-reviewed4/7Primary ↗
SaProtWestlake / Zhejiang collaboratorsProtein Foundation & RepresentationCuratedProtein understanding benchmarksStructure-aware representationPeer-reviewed4/7Primary ↗
ProtT5RostlabProtein Foundation & RepresentationPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
AnkhProteinea / academic collaboratorsProtein Foundation & RepresentationPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
xTrimoPGLMBioMap / Tsinghua collaboratorsProtein Foundation & RepresentationPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
Uni-Mol / Uni-Mol2DP Technology / DeepModelingSmall-Molecule & ChemistryCuratedMolecular representation/property benchmarks3D molecular representationOpen evaluation3/7Primary ↗
MolMIMNVIDIASmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
MegaMolBARTNVIDIASmall-Molecule & ChemistryPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
DNABERT-2Multi-institution research teamGenomics, DNA & RNACuratedGenomic downstream tasksGeneral genomic representationOpen evaluation4/7Primary ↗
CaduceusCornell / Tri Dao collaboratorsGenomics, DNA & RNACuratedLong-range genomic benchmarksVariant / sequence predictionPreprint / open evaluation4/7Primary ↗
GROVERPoetsch Lab / collaboratorsGenomics, DNA & RNAPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
BorzoiCalico / academic collaboratorsGenomics, DNA & RNACuratedFunctional-genomics and RNA-seq evaluationsSequence-to-RNA predictionPeer-reviewed4/7Primary ↗
RiNALMoUniversity of Zagreb / A*STAR collaboratorsRNA Models & DesignCuratedRNA downstream and structure tasksRNA representation / structure predictionPeer-reviewed4/7Primary ↗
RNA-FMML4Bio / academic collaboratorsRNA Models & DesignPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated2/7Primary ↗
UNI-RNAAcademic RNA foundation-model collaboratorsRNA Models & DesignPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
LucaOneBioMap / collaboratorsGenomics, DNA & RNACuratedDNA/RNA/protein downstream tasksCross-domain biological representationPeer-reviewed3/7Primary ↗
Universal Cell Embeddings (UCE)Stanford / collaboratorsVirtual Cells & Single-CellCuratedCross-dataset single-cell evaluationsUniversal cell representationPeer-reviewed / open evaluation2/7Primary ↗
TranscriptFormerChan Zuckerberg InitiativeVirtual Cells & Single-CellPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
SCimilarityGenentechVirtual Cells & Single-CellPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending3/7Primary ↗
HelixFold3PaddleHelix / BaiduStructure PredictionPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
UNI / UNI2Mahmood Lab / Brigham and Women’s HospitalTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Virchow / Virchow2Paige / Microsoft Research collaboratorsTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
Prov-GigaPathMicrosoft Research / ProvidenceTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
CONCHMahmood LabTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
TITANMahmood Lab / collaboratorsTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedPrimary paper linked; benchmark extraction pending2/7Primary ↗
H-OptimusBioptimusTissue, Pathology & ImagingPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
NicheformerSpatial biology research teamSpatial & Tissue SystemsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
NovaeSpatial-omics research communitySpatial & Tissue SystemsPendingNot yet extractedNot yet extractedNo task-specific benchmark record curated1/7Primary ↗
AIDO Cell 1.0GenBio AIVirtual Cells & Single-CellCuratedVirtual Cell Benchmark 1.031 metrics across five task familiesDeveloper-reported; independent reproduction pending2/7Primary ↗
AIDO.TissueGenBio AISpatial & Tissue SystemsCuratedSpatial transcriptomics downstream evaluationsCell/tissue representation and predictionPreprint / open evaluation4/7Primary ↗
GenBio-PathFMGenBio AITissue, Pathology & ImagingCuratedTHUNDER / HEST / PathoROBHistopathology representation and downstream performanceTechnical report / open evaluation4/7Primary ↗
Tahoe-x1Tahoe TherapeuticsVirtual Cells & Single-CellCuratedFour disease-relevant single-cell evaluation groupsEssentiality, cancer hallmarks, cell type and perturbation responsePreprint / open evaluation4/7Primary ↗

Coverage shows how many evidence fields are documented in BioAtlas. It is not a model-performance or readiness score. External suite listings are provenance records, not reproduced leaderboards.