Protein language models trained on increasingly large UniRef100 snapshots show non-monotonic fitness prediction performance, with no evidence of data saturation.
Training compute-optimal protein language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-bio.QM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Scaling and Data Saturation in Protein Language Models
Protein language models trained on increasingly large UniRef100 snapshots show non-monotonic fitness prediction performance, with no evidence of data saturation.