Pith. sign in

REVIEW 2 cited by

OpenProteinSet: Training data for structural biology at scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.05326 v1 pith:RDGRD6JI submitted 2023-08-10 q-bio.BM cs.LG

classification q-bio.BMcs.LG
keywords proteinalphafold2msasopenproteinsetdatastructurebeendesign
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multiple sequence alignments (MSAs) of proteins encode rich biological information and have been workhorses in bioinformatic methods for tasks like protein design and protein structure prediction for decades. Recent breakthroughs like AlphaFold2 that use transformers to attend directly over large quantities of raw MSAs have reaffirmed their importance. Generation of MSAs is highly computationally intensive, however, and no datasets comparable to those used to train AlphaFold2 have been made available to the research community, hindering progress in machine learning for proteins. To remedy this problem, we introduce OpenProteinSet, an open-source corpus of more than 16 million MSAs, associated structural homologs from the Protein Data Bank, and AlphaFold2 protein structure predictions. We have previously demonstrated the utility of OpenProteinSet by successfully retraining AlphaFold2 on it. We expect OpenProteinSet to be broadly useful as training and validation data for 1) diverse tasks focused on protein structure, function, and design and 2) large-scale multimodal machine learning research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters

    cs.DC 2026-04 conditional novelty 6.0 of 10

    Minos groups GPU apps by power-spike histograms and SM/DRAM utilization, transferring frequency-cap behavior to unseen workloads with ~4% power and ~3% performance error while cutting profiling cost ~89%.

  2. Evolution-Aware MSA Reasoning for Subsampling via Factor Graphs

    cs.LG 2026-07 conditional novelty 5.0 of 10

    AP-REASONER, an affinity-propagation factor-graph sampler with alpha/beta knobs, improves MSA-based protein LM pretraining for contact prediction and conformational sampling, though gains are modest and partly baselin...

Pith tools