Pith. sign in

REVIEW 3 major objections 5 minor 63 references

Exploring zero-shot structure-based protein fitness prediction

T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Disordered protein regions drag down zero-shot fitness prediction accuracy

desk verdict A clean, reproducible benchmark paper whose residue-level disorder analysis is a genuinely useful new measurement, but the headline claim would be stronger with paired significance tests and validated experimental structures. read the letter →

arxiv 2504.16886 v1 pith:3BLEYNCY submitted 2025-04-23 q-bio.QM cs.LGq-bio.BM

classification q-bio.QMcs.LGq-bio.BM
keywords zero-shotfitnesspredictionProteinGymintrinsicallydisorderedregionsstructure-basedmodelsAlphaFold2deepmutationalscanningensembles
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Zero-shot fitness prediction models that score how mutations affect protein function perform worse when a mutation falls in a disordered region—a stretch of the protein with no fixed 3D structure. This paper shows that such regions appear in 29% of ProteinGym substitution assays and that the performance drop occurs across sequence-based, MSA-based, structure-based, multi-modal, and ensemble models, for most measured function types. The finding matters because disordered regions are common in real proteins, and fitness predictors are used to interpret genetic variants and guide protein engineering. The paper also examines how the choice of input structure (predicted versus experimental) affects predictions, and it proposes simple multi-modal ensembles as strong baselines for this task.

What carries the argument

The central objects are the 217 deep mutational scanning substitution assays in ProteinGym, the matched AlphaFold 2 predicted structures and PDB experimental structures, and DisProt annotations of intrinsic disorder. The key comparison is per-mutation: separating each assay's mutations into those in disordered versus ordered regions and comparing model Spearman correlations on those two subsets. For the structure comparison, the machinery is the paired difference ρ_pred − ρ_exp for each assay, split into monomers and multimers, plus case studies on α-synuclein (P37840) and NKX3-1 (Q99801) that visually align experimental and predicted structures to explain extreme differences.

What would settle it

If one repeated the 65-assay structure comparison using manually validated experimental structures that exactly match the functional assay condition (for example, the micelle-bound helical α-synuclein instead of a fibril), and the advantage of AlphaFold 2 predicted structures over experimental structures disappeared, the claim that predicted structures are generally as good or better for fitness prediction would be falsified. This test is feasible because the authors identify the problematic case and the required structure is known.

Watch

Extended reading notes

Core claim

The paper establishes that intrinsically disordered regions systematically degrade zero-shot protein fitness prediction: in 43 ProteinGym DMS assays that contain mutations in both ordered and disordered regions, Spearman correlations between predicted and measured fitness are consistently lower for disordered-region mutations across models of all input modalities, including structure-based ESM-IF1, multi-modal ProtSSN, SaProt, TranceptEVE L, sequence-based ESM2, MSA-based GEMME, and several simple ensembles. The disorder effect holds for most function types but is less pronounced for stability assays. Separately, the paper finds that AlphaFold 2 predicted structures often yield higher Spearman correlations than experimental structures selected from the PDB, cautioning that structure–assay matching is critical and that predicted structures in disordered regions can be misleading—illustrated with α-synuclein, where the predicted structure adopts a fibril-like conformation while the experimental structure is helical, and the assay's function depends on the membrane-bound helical state.

Load-bearing premise

The paper assumes that the experimental structures retrieved from the PDB by sequence identity and full-length coverage represent the functional conformation tested in each deep mutational scanning assay; the authors did not manually verify this, so some structure-choice comparisons may reflect mismatched conformations rather than a general property of experimental versus predicted structures.

Editorial extensions

If this is right

  • Model developers should treat disordered regions as a distinct failure mode, possibly by masking low-confidence predicted structure coordinates or by conditioning on conformational ensembles rather than a single static structure.
  • Users of zero-shot fitness predictors for clinical variant interpretation should be cautious when a variant falls in a documented disordered region, since prediction quality there is systematically lower.
  • Simple ensembles that combine structure-based ESM-IF1 with sequence- or MSA-based models (e.g., StructSeq) can match or exceed more sophisticated multi-modal architectures, so such ensembles should be standard baselines in future ProteinGym-style benchmarks.
  • Predicted structures from AlphaFold 2, being trained on a distribution that often excludes the functional conformation (as for α-synuclein), can mislead fitness prediction in disordered regions; matching the structure to the assay's actual conformation is essential.
  • Stability assays stand apart: structure-based models do not show the same disorder-related drop, likely because the measured property (proteolytic stability) is more directly tied to a single folded state.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The disorder penalty may be partly an artifact of conservation: disordered regions evolve faster and are less conserved, so evolutionary-signal models (PLMs trained on sequence databases, MSA-based models) will naturally assign them flatter likelihoods; this suggests the drop may be informative rather than purely noise, and one could test whether models trained on disorder-aware data recover predi
  • A direct extension would be to use the same protocol to quantify disorder penalties on the larger set of DMS assays in MaveDB and ProtaBank, which would test how generalizable the 29% prevalence and the per-function trends are.
  • Because the paper used only one experimental structure per protein, a natural next study is to sample multiple experimental conformations or use ensemble docking to see whether supplying a more functional conformation (for α-synuclein, the micelle-bound helix) restores ESM-IF1 performance to the level seen for ordered regions.
  • The finding that simple ensembles outperform joint architectures suggests a 'disorder-aware ensemble' that down-weights structure-based models inside disordered regions could improve overall Spearman correlation; this is a concrete, testable design that the paper does not evaluate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper benchmarks zero-shot structure-based protein fitness predictors on ProteinGym substitution assays, comparing predicted (AlphaFold 2) structures against automatically matched experimental PDB structures, and analyzes how intrinsically disordered regions affect prediction quality across structure-, sequence-, MSA-, and multimodal models. It reports that disordered regions are widespread in ProteinGym, that structure-based models and simple multimodal ensembles perform competitively, and that ensemble methods are strong baselines. The main empirical claims are that predicted structures often outperform automatically matched experimental structures and that mutations in disordered regions are predicted worse than mutations in ordered regions across most model families and function types.

Significance. If the disorder-related claim survives statistical scrutiny, this is a useful and broadly relevant observation for the protein fitness prediction community, because it connects a biological property (disorder) to the behavior of many distinct model families and to the choice of structural input. The study's strengths include the use of standard external benchmarks (ProteinGym, DisProt, PDB), the breadth of models and ensembles considered, and the public release of code and data for reproduction. The multi-modal ensemble results are a useful baseline reference even though the primary novelty lies in the disorder analysis.

major comments (3)
  1. [Results, 'Intrinsically disordered regions...' and Figure 3; Tables 4-5] The central claim that disordered regions systematically reduce zero-shot fitness prediction quality across model families is supported only by descriptive per-function means in Figure 3; no confidence intervals, error bars, or paired significance tests are reported for the 43 assays. Since Table 4 includes disordered regions as short as a single residue (e.g., RCD1_ARATH_Tsuboyama_2023_5OAO, region 0-0), the per-assay Spearman correlations within disordered regions are computed on very few mutations and are highly variable. Please add a paired test across the 43 assays (e.g., Wilcoxon signed-rank on per-assay ordered-vs-disordered Spearman differences) with effect sizes and confidence intervals, and perform the same analysis per function type where sample sizes permit.
  2. [Table 1, Figure 1, and Appendix 'Selection process for experimental structures'] Because experimental structures were selected by sequence identity and full-length coverage without manual validation (as stated in the Appendix), and because ESM-IF1 is explicitly trained on AlphaFold 2 predicted structures (as noted in the Discussion), Table 1's claim that predicted structures outperform experimental structures is not established as a general property of structure-based fitness prediction. The alpha-synuclein example (Figure 5) shows that the automatically selected experimental structure can represent a different conformation from the functional one tested in the assay. Please restrict the claim to the 65 automatically matched PDB structures used here, or validate a subset of structures against the functional conformations described in the DMS assay literature.
  3. [Appendix, 'Selection process for disordered proteins', Stage 2] The disorder-annotation mapping treats same-length target sequences as referring to the same sequence even though the authors state that ProteinGym target sequences can differ from UniProt sequences at arbitrary positions, and subset sequences are located within UniProt without a described alignment algorithm. Misplacement of disordered annotations would directly bias the ordered-versus-disordered comparison in Figure 3. Please specify the alignment method precisely and provide a sensitivity analysis restricted to assays with exact sequence matching or a conservative alignment confidence threshold.
minor comments (5)
  1. [Full text, header] There is a typo in the header: 'Code a v ailability' should be 'Code availability'.
  2. [Table 6 caption] The caption says 'Ensembled 3' while Table 2 and the methods use 'Ensemble 3'; please make the naming consistent.
  3. [Table 3 and Table 4] The manuscript reports 216 assays used in the study (Table 3 and the Appendix) but also states that 63 of 217 ProteinGym assays have disordered regions; the excluded assay BRCA2_HUMAN_Erwood_2022_HEK293T still appears in Table 4. Please reconcile the denominators and clarify which assays are included in each count.
  4. [Figure 3 caption] The caption says scores are 'averaged across UniProtID and then function type' while the text describes 43 assays; please clarify the exact aggregation hierarchy and how many assays contribute to each point.
  5. [Table 1 header] The column labels 'Count of DMS assays with positive and negative ρpred−ρexp' are unclear; please use explicit labels such as 'Number of assays with ρpred − ρexp ≥ 0' and 'Number of assays with ρpred − ρexp < 0'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an empirical evaluation against external benchmarks (ProteinGym, DisProt, PDB) with no fitted parameters renamed as predictions and no load-bearing self-citation.

full rationale

The paper's claims are empirical comparisons of published zero-shot model predictions against external fitness assay data, not derivations from assumed conclusions. No parameter is fitted to a subset and then reported as a prediction; model scores marked with an asterisk are taken directly from ProteinGym, and the remaining scores are generated by running published models on the same external benchmark. The disorder analysis uses DisProt annotations and ProteinGym mutation sets, and the central ordered-versus-disordered comparison is a descriptive assessment of external model outputs, so it cannot reduce to its inputs by construction. Ensembles are explicitly arithmetic combinations of published log-probabilities with no post-hoc training, so their performance is an independent empirical result. The paper's acknowledged limitation that experimental structures were not manually confirmed to match each DMS assay is a validity caveat, not a circularity: it weakens the structure-matching comparison but does not make any claim equivalent to an input by definition. Self-citation is minimal and not load-bearing: the one citation involving the authors (Gelman et al. 2025) is mentioned only to note an excluded model modality, and the central claims do not rest on it. The lack of significance tests for Figure 3 is a statistical-rigor concern, not evidence of circularity. Under the required standard, where circularity must be demonstrated by quoting the paper and exhibiting a specific reduction, no such reduction exists here.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claims rest on the standard benchmark interpretation of ProteinGym DMS scores as fitness ground truth, on DisProt annotations for disorder, and on the suitability of AlphaFold2 predicted structures. No free parameters are fitted in this study; the ensemble combination is arithmetic with no learned weights. The only assumption specific to this paper is that PDB structures selected by sequence identity are relevant to the assayed function, which the authors explicitly flag as unvalidated. No invented entities are introduced.

assumptions (5)
  • domain assumption ProteinGym DMS assay scores are valid ground-truth measurements of protein fitness.
    All evaluations use Spearman correlation and Top-10 Recall against DMS scores as the fitness measure, following ProteinGym conventions (Methods).
  • domain assumption DisProt disorder annotations, after sequence alignment, correctly identify disordered residues in each ProteinGym target sequence.
    The central disorder analysis relies on the 2024_12 DisProt release and an alignment procedure that handles sequence differences between UniProt and ProteinGym (Appendix, 'Selection process for disordered proteins').
  • domain assumption AlphaFold2 predicted structures from ProteinGym are suitable inputs for evaluating structure-based fitness models.
    All predicted structures were generated with AlphaFold2 v2.3.1 default parameters (Methods) and are used without additional validation.
  • ad hoc to paper Experimental structures selected from PDB using sequence identity and full-length coverage represent the functional conformation in each DMS assay.
    The authors state they did not manually confirm the structures match the functional assays, which makes the predicted-versus-experimental comparison assumption-specific to this study (Appendix, 'Selection process for experimental structures').
  • domain assumption Inverse folding and masked language model likelihoods are indicative of protein fitness.
    The models evaluated (ESM-IF1, ESM2, VESPA, SaProt) are used in the zero-shot fitness setting based on prior work cited in the Introduction and Methods.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploring zero-shot structure-based protein fitness prediction." pith.science (2026). https://pith.science/paper/3BLEYNCY

@misc{pith2026250416886,
  author       = {Pith},
  title        = {Pith review of: Exploring zero-shot structure-based protein fitness prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3BLEYNCY}},
  note         = {Machine review of arXiv:2504.16886}
}
read the original abstract

The ability to make zero-shot predictions about the fitness consequences of protein sequence changes with pre-trained machine learning models enables many practical applications. Such models can be applied for downstream tasks like genetic variant interpretation and protein engineering without additional labeled data. The advent of capable protein structure prediction tools has led to the availability of orders of magnitude more precomputed predicted structures, giving rise to powerful structure-based fitness prediction models. Through our experiments, we assess several modeling choices for structure-based models and their effects on downstream fitness prediction. Zero-shot fitness prediction models can struggle to assess the fitness landscape within disordered regions of proteins, those that lack a fixed 3D structure. We confirm the importance of matching protein structures to fitness assays and find that predicted structures for disordered regions can be misleading and affect predictive performance. Lastly, we evaluate an additional structure-based model on the ProteinGym substitution benchmark and show that simple multi-modal ensembles are strong baselines.

Figures

Figures reproduced from arXiv: 2504.16886 by the authors.

Figure 1
Figure 1. Difference in Spearman correlation between predictions made using predicted and experi [PITH_FULL_IMAGE:figures/full_fig_p018_1.png] view at source ↗
Figure 2
Figure 2. Protein-level Spearman ρ predicting function of proteins comprising some form of disordered regions in the ProteinGym assay target sequence according to the DisProt database. 19 [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗
Figure 3
Figure 3. Spearman correlation averaged across UniProtID and then function type for disordered [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Predictions made by ESM2 650M on the P53_HUMAN_Giacomelli_2018_WT_Nutlin [PITH_FULL_IMAGE:figures/full_fig_p022_4.png]
Figure 5
Figure 5. Figure 5: Protein structure alignment between the experimental (orange, PDB 1XQ8 (Ulmer et al., [PITH_FULL_IMAGE:figures/full_fig_p022_5.png]
Figure 6
Figure 6. Figure 6: Protein structure alignment between the experimental (orange, PDB 2L9R) and predicted [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Distribution of Spearman correlation of various models being considered across 216 DMS assays from [PITH_FULL_IMAGE:figures/full_fig_p023_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

63 extracted references · 35 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    DisProt in 2024: improving function annotation of intrinsically disordered proteins

    Maria Cristina Aspromonte, Maria Victoria Nugnes, Federica Quaglia, Adel Bouharoua, DisProt Consortium , Silvio C E Tosatto, and Damiano Piovesan. DisProt in 2024: improving function annotation of intrinsically disordered proteins. Nucleic Acids Research, 52 0 (D1): 0 D434--D441, January 2024. ISSN 0305-1048. doi:10.1093/nar/gkad928. URL https://doi.org/1...

  3. [3]

    RCSB Protein Data Bank : exploring protein 3D similarities via comprehensive structural alignments

    Sebastian Bittrich, Joan Segura, Jose M Duarte, Stephen K Burley, and Yana Rose. RCSB Protein Data Bank : exploring protein 3D similarities via comprehensive structural alignments. Bioinformatics, 40 0 (6): 0 btae370, June 2024. ISSN 1367-4811. doi:10.1093/bioinformatics/btae370. URL https://doi.org/10.1093/bioinformatics/btae370

  4. [4]

    Rapid protein stability prediction using deep learning representations

    Lasse M Blaabjerg, Maher M Kassem, Lydia L Good, Nicolas Jonsson, Matteo Cagiada, Kristoffer E Johansson, Wouter Boomsma, Amelie Stein, and Kresten Lindorff-Larsen. Rapid protein stability prediction using deep learning representations. eLife, 12: 0 e82593, May 2023. ISSN 2050-084X. doi:10.7554/eLife.82593. URL https://doi.org/10.7554/eLife.82593. Publish...

  5. [5]

    Blaabjerg, Nicolas Jonsson, Wouter Boomsma, Amelie Stein, and Kresten Lindorff-Larsen

    Lasse M. Blaabjerg, Nicolas Jonsson, Wouter Boomsma, Amelie Stein, and Kresten Lindorff-Larsen. SSEmb : A joint embedding of protein sequence and structure enables robust variant effect predictions. Nature Communications, 15 0 (1): 0 9646, November 2024. ISSN 2041-1723. doi:10.1038/s41467-024-53982-z. URL https://www.nature.com/articles/s41467-024-53982-z...

  6. [6]

    Predicting protein variants with equivariant graph neural networks

    Antonia Boca and Simon Mathis. Predicting protein variants with equivariant graph neural networks. arXiv:2306.12231, July 2023. doi:10.48550/arXiv.2306.12231. URL http://arxiv.org/abs/2306.12231

  7. [7]

    Brown, Sachiko Takayama, Andrew M

    Celeste J. Brown, Sachiko Takayama, Andrew M. Campen, Pam Vise, Thomas W. Marshall, Christopher J. Oldfield, Christopher J. Williams, and A. Keith Dunker. Evolutionary Rate Heterogeneity in Proteins with Long Disordered Regions . Journal of Molecular Evolution, 55: 0 104--110, July 2002. ISSN 1432-1432. doi:10.1007/s00239-001-2309-6. URL https://doi.org/1...

  8. [8]

    Center for High Throughput Computing , 2006

    Center for High Throughput Computing . Center for High Throughput Computing , 2006. URL https://chtc.cs.wisc.edu/

Show all 63 references
  1. [9]

    Schneider, Andrew W

    Jun Cheng, Guido Novati, Joshua Pan, Clare Bycroft, Akvil \.e Z emgulyt \.e , Taylor Applebaum, Alexander Pritzel, Lai Hong Wong, Michal Zielinski, Tobias Sargeant, Rosalia G. Schneider, Andrew W. Senior, John Jumper, Demis Hassabis, Pushmeet Kohli, and Z iga Avsec. Accurate p...

  2. [10]

    Dauparas, I

    J. Dauparas, I. Anishchenko, N. Bennett, H. Bai, R. J. Ragotte, L. F. Milles, B. I. M. Wicky, A. Courbet, R. J. de Haas, N. Bethel, P. J. Y. Leung, T. F. Huddy, S. Pellock, D. Tischer, F. Chan, B. Koepnick, H. Nguyen, A. Kang, B. Sankaran, A. K. Bera, N. P. King, and D. Baker....

  3. [11]

    Mohamed Fawzy and Joseph A. Marsh. Assessing variant effect predictors and disease mechanisms in intrinsically disordered proteins. bioRxiv, pp.\ 2025.04.01.646619, April 2025. doi:10.1101/2025.04.01.646619. URL https://www.biorxiv.org/content/10.1101/2025.04.01.646619v1

  4. [12]

    Protein Interface Prediction using Graph Convolutional Networks

    Alex Fout, Jonathon Byrd, Basir Shariat, and Asa Ben-Hur. Protein Interface Prediction using Graph Convolutional Networks . In Advances in Neural Information Processing Systems , Long Beach, CA, USA, 2017. Curran Associates, Inc. URL https://proceedings.neurips.cc/paper/2017/h...

  5. [13]

    Hierarchical graph learning for protein protein interaction

    Ziqi Gao, Chenran Jiang, Jiawen Zhang, Xiaosen Jiang, Lanqing Li, Peilin Zhao, Huanming Yang, Yong Huang, and Jia Li. Hierarchical graph learning for protein protein interaction. Nature Communications, 14 0 (1): 0 1093, February 2023. ISSN 2041-1723. doi:10.1038/s41467-023-367...

  6. [14]

    Sam Gelman, Bryce Johnson, Chase Freschlin, Arnav Sharma, Sameer D Costa, John Peters, Anthony Gitter, and Philip A. Romero. Biophysics-based protein language models for protein engineering. bioRxiv, pp.\ 2024.03.15.585128, January 2025. doi:10.1101/2024.03.15.585128. URL http...

  7. [15]

    Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Berenberg, Tommi Vatanen, Chris Chandler, Bryn C

    Vladimir Gligorijevi \'c , P. Douglas Renfrew, Tomasz Kosciolek, Julia Koehler Leman, Daniel Berenberg, Tommi Vatanen, Chris Chandler, Bryn C. Taylor, Ian M. Fisk, Hera Vlamakis, Ramnik J. Xavier, Rob Knight, Kyunghyun Cho, and Richard Bonneau. Structure-based protein function...

  8. [16]

    Learning inverse folding from millions of predicted structures

    Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. In Proceedings of the 39th International Conference on Machine Learning , pp.\ 8946--8970. PMLR, June 2022. ...

  9. [17]

    Schief, and David Baker

    Po-Ssu Huang, Yih-En Andrew Ban, Florian Richter, Ingemar Andre, Robert Vernon, William R. Schief, and David Baker. RosettaRemodel : A Generalized Framework for Flexible Backbone Protein Design . PLOS ONE, 6 0 (8): 0 e24109, August 2011. ISSN 1932-6203. doi:10.1371/journal.pon...

  10. [18]

    Yufei Huang, Siyuan Li, Lirong Wu, Jin Su, Haitao Lin, Odin Zhang, Zihan Liu, Zhangyang Gao, Jiangbin Zheng, and Stan Z. Li. Protein 3D Graph Structure Learning for Robust Structure - Based Protein Property Prediction . Proceedings of the AAAI Conference on Artificial Intellig...

  11. [19]

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stani...

  12. [20]

    Kulikova, Daniel J

    Anastasiya V. Kulikova, Daniel J. Diaz, Tianlong Chen, T. Jeffrey Cole, Andrew D. Ellington, and Claus O. Wilke. Two sequence- and two structure-based ML models have learned different aspects of protein biochemistry. Scientific Reports, 13 0 (1): 0 13280, August 2023. ISSN 204...

  13. [21]

    GEMME : A Simple and Fast Global Epistatic Model Predicting Mutational Effects

    Elodie Laine, Yasaman Karami, and Alessandra Carbone. GEMME : A Simple and Fast Global Epistatic Model Predicting Mutational Effects . Molecular Biology and Evolution, 36 0 (11): 0 2604--2619, November 2019. ISSN 0737-4038. doi:10.1093/molbev/msz179. URL https://doi.org/10.109...

  14. [22]

    Weitzner, Steven M

    Julia Koehler Leman, Brian D. Weitzner, Steven M. Lewis, Jared Adolf-Bryfogle, Nawsad Alam, Rebecca F. Alford, Melanie Aprahamian, David Baker, Kyle A. Barlow, Patrick Barth, Benjamin Basanta, Brian J. Bender, Kristin Blacklock, Jaume Bonet, Scott E. Boyken, Phil Bradley, Chri...

  15. [23]

    Protein Expansion Is Primarily due to Indels in Intrinsically Disordered Regions

    Sara Light, Rauan Sagit, Oxana Sachenkova, Diana Ekman, and Arne Elofsson. Protein Expansion Is Primarily due to Indels in Intrinsically Disordered Regions . Molecular Biology and Evolution, 30 0 (12): 0 2645--2653, December 2013. ISSN 0737-4038. doi:10.1093/molbev/mst157. URL...

  16. [24]

    Evolutionary-scale prediction of atomic-level protein structure with a language model

    Zeming Lin, Halil Akin, Roshan Rao, Brian Hie, Zhongkai Zhu, Wenting Lu, Nikita Smetanin, Robert Verkuil, Ori Kabeli, Yaniv Shmueli, Allan dos Santos Costa, Maryam Fazel-Zarandi, Tom Sercu, Salvatore Candido, and Alexander Rives. Evolutionary-scale prediction of atomic-level p...

  17. [25]

    Kragelund

    Kresten Lindorff-Larsen and Birthe B. Kragelund. On the Potential of Machine Learning to Examine the Relationship Between Sequence , Structure , Dynamics and Function of Intrinsically Disordered Proteins . Journal of Molecular Biology, 433 0 (20): 0 167196, October 2021. ISSN ...

  18. [26]

    Freed, and Carrie A

    Rongming Liu, Liya Liang, Maria Priscila Lacerda, Emily F. Freed, and Carrie A. Eckert. Chapter 10 - Advances in protein engineering and its application in synthetic biology. In Vijai Singh (ed.), New Frontiers and Applications of Synthetic Biology , pp.\ 147--158. Academic Pr...

  19. [27]

    Marks, Lucy J

    Debora S. Marks, Lucy J. Colwell, Robert Sheridan, Thomas A. Hopf, Andrea Pagnani, Riccardo Zecchina, and Chris Sander. Protein 3D Structure Computed from Evolutionary Sequence Variation . PLOS ONE, 6 0 (12): 0 e28766, December 2011. ISSN 1932-6203. doi:10.1371/journal.pone.00...

  20. [28]

    Embeddings from protein language models predict conservation and variant effects

    Céline Marquet, Michael Heinzinger, Tobias Olenyi, Christian Dallago, Kyra Erckert, Michael Bernhofer, Dmitrii Nechaev, and Burkhard Rost. Embeddings from protein language models predict conservation and variant effects. Human Genetics, 141 0 (10): 0 1629--1647, October 2022. ...

  21. [29]

    ColabFold : making protein folding accessible to all

    Milot Mirdita, Konstantin Schütze, Yoshitaka Moriwaki, Lim Heo, Sergey Ovchinnikov, and Martin Steinegger. ColabFold : making protein folding accessible to all. Nature Methods, 19 0 (6): 0 679--682, June 2022. ISSN 1548-7105. doi:10.1038/s41592-022-01488-1. URL https://www.nat...

  22. [30]

    Newberry, Taylor Arhar, Jean Costello, George C

    Robert W. Newberry, Taylor Arhar, Jean Costello, George C. Hartoularos, Alison M. Maxwell, Zun Zar Chi Naing, Maureen Pittman, Nishith R. Reddy, Daniel M. C. Schwarz, Douglas R. Wassarman, Taia S. Wu, Daniel Barrero, Christa Caggiano, Adam Catching, Taylor B. Cavazos, Laurel S...

  23. [31]

    Newberry, Jaime T

    Robert W. Newberry, Jaime T. Leong, Eric D. Chow, Martin Kampmann, and William F. DeGrado. Deep mutational scanning reveals the structural basis for -synuclein activity. Nature Chemical Biology, 16 0 (6): 0 653--659, June 2020 b . ISSN 1552-4469. doi:10.1038/s41589-020-0480-6....

  24. [32]

    GitHub issue: AlphaFold2 structure generation details, 2025

    Pascal Notin. GitHub issue: AlphaFold2 structure generation details, 2025. URL https://github.com/OATML-Markslab/ProteinGym/issues/65#issuecomment-2640138972

  25. [33]

    Gomez, Debora Marks, and Yarin Gal

    Pascal Notin, Mafalda Dias, Jonathan Frazer, Javier Marchena Hurtado, Aidan N. Gomez, Debora Marks, and Yarin Gal. Tranception: Protein Fitness Prediction with Autoregressive Transformers and Inference -time Retrieval . In Proceedings of the 39th International Conference on Ma...

  26. [34]

    Kollasch, Daniel Ritter, Yarin Gal, and Debora Susan Marks

    Pascal Notin, Lood Van Niekerk, Aaron W. Kollasch, Daniel Ritter, Yarin Gal, and Debora Susan Marks. TranceptEVE : Combining Family -specific and Family -agnostic Models of Protein Sequences for Improved Fitness Prediction . In NeurIPS 2022 Workshop on Learning Meaningful Repr...

  27. [35]

    ProteinGym : Large - Scale Benchmarks for Protein Fitness Prediction and Design

    Pascal Notin, Aaron Kollasch, Daniel Ritter, Lood van Niekerk, Steffanie Paul, Han Spinner, Nathan Rollins, Ada Shaw, Rose Orenbuch, Ruben Weitzman, Jonathan Frazer, Mafalda Dias, Dinko Franceschi, Yarin Gal, and Debora Marks. ProteinGym : Large - Scale Benchmarks for Protein ...

  28. [36]

    Combining Structure and Sequence for Superior Fitness Prediction

    Steffanie Paul, Aaron W Kollasch, Pascal Notin, and Debora S Marks. Combining Structure and Sequence for Superior Fitness Prediction . In NeurIPS 2023 Generative AI and Biology (GenBio) Workshop, 2023

  29. [37]

    Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives

    Roshan M. Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives. MSA Transformer . In Proceedings of the 38th International Conference on Machine Learning , pp.\ 8844--8856. PMLR, July 2021. URL https://proceedings.mlr.press/v1...

  30. [38]

    Lawrence Zitnick, Jerry Ma, and Rob Fergus

    Alexander Rives, Joshua Meier, Tom Sercu, Siddharth Goyal, Zeming Lin, Jason Liu, Demi Guo, Myle Ott, C. Lawrence Zitnick, Jerry Ma, and Rob Fergus. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proceedings of the...

  31. [39]

    Romero, Andreas Krause, and Frances H

    Philip A. Romero, Andreas Krause, and Frances H. Arnold. Navigating the protein fitness landscape with Gaussian processes. Proceedings of the National Academy of Sciences of the United States of America, 110 0 (3): 0 E193--E201, January 2013. ISSN 0027-8424. doi:10.1073/pnas.1...

  32. [40]

    Rubin, Jeremy Stone, Aisha Haley Bianchi, Benjamin J

    Alan F. Rubin, Jeremy Stone, Aisha Haley Bianchi, Benjamin J. Capodanno, Estelle Y. Da, Mafalda Dias, Daniel Esposito, Jonathan Frazer, Yunfan Fu, Sally B. Grindstaff, Matthew R. Harrington, Iris Li, Abbye E. McEwen, Joseph K. Min, Nick Moore, Olivia G. Moscatelli, Jesslyn Ong...

  33. [41]

    High- Resolution Comparative Modeling with RosettaCM

    Yifan Song, Frank DiMaio, Ray Yu-Ruei Wang, David Kim, Chris Miles, TJ Brunette, James Thompson, and David Baker. High- Resolution Comparative Modeling with RosettaCM . Structure, 21 0 (10): 0 1735--1742, October 2013. ISSN 0969-2126. doi:10.1016/j.str.2013.08.005. URL https:/...

  34. [42]

    MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets

    Martin Steinegger and Johannes S \"o ding. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology, 35 0 (11): 0 1026--1028, November 2017. ISSN 1546-1696. doi:10.1038/nbt.3988. URL https://www.nature.com/articles/nbt.39...

  35. [43]

    SaProt : Protein Language Modeling with Structure -aware Vocabulary

    Jin Su, Chenchen Han, Yuyang Zhou, Junjie Shan, Xibin Zhou, and Fajie Yuan. SaProt : Protein Language Modeling with Structure -aware Vocabulary . In The Twelfth International Conference on Learning Representations, October 2023. URL https://openreview.net/forum?id=6MRm3G4NiU

  36. [44]

    Semantical and Geometrical Protein Encoding Toward Enhanced Bioactivity and Thermostability

    Yang Tan, Bingxin Zhou, Lirong Zheng, Guisheng Fan, and Liang Hong. Semantical and Geometrical Protein Encoding Toward Enhanced Bioactivity and Thermostability . eLife, 13, June 2024. doi:10.7554/eLife.98033.1. URL https://elifesciences.org/reviewed-preprints/98033. Publisher:...

  37. [45]

    UniProt : the Universal Protein Knowledgebase in 2023

    The UniProt Consortium . UniProt : the Universal Protein Knowledgebase in 2023. Nucleic Acids Research, 51 0 (D1): 0 D523--D531, January 2023. ISSN 0305-1048. doi:10.1093/nar/gkac1052. URL https://doi.org/10.1093/nar/gkac1052

  38. [46]

    Intrinsically disordered proteins: An overview

    Rakesh Trivedi and Hampapathalu Adimurthy Nagarajaram. Intrinsically disordered proteins: An overview. International Journal of Molecular Sciences, 23 0 (22), 2022. ISSN 1422-0067. doi:10.3390/ijms232214050. URL https://www.mdpi.com/1422-0067/23/22/14050

  39. [47]

    Weinstein, Niall M

    Kotaro Tsuboyama, Justas Dauparas, Jonathan Chen, Elodie Laine, Yasser Mohseni Behbahani, Jonathan J. Weinstein, Niall M. Mangan, Sergey Ovchinnikov, and Gabriel J. Rocklin. Mega-scale experimental analysis of protein folding stability in biology and design. Nature, 620 0 (797...

  40. [48]

    Tuttle, Gemma Comellas, Andrew J

    Marcus D. Tuttle, Gemma Comellas, Andrew J. Nieuwkoop, Dustin J. Covell, Deborah A. Berthold, Kathryn D. Kloepper, Joseph M. Courtney, Jae K. Kim, Alexander M. Barclay, Amy Kendall, William Wan, Gerald Stubbs, Charles D. Schwieters, Virginia M. Y. Lee, Julia M. George, and Cha...

  41. [49]

    Ulmer, Ad Bax, Nelson B

    Tobias S. Ulmer, Ad Bax, Nelson B. Cole, and Robert L. Nussbaum. Structure and Dynamics of Micelle -bound Human - Synuclein . Journal of Biological Chemistry, 280 0 (10): 0 9595--9603, March 2005. ISSN 0021-9258. doi:10.1074/jbc.M411805200. URL https://www.sciencedirect.com/sc...

  42. [50]

    Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron L

    Michel van Kempen, Stephanie S. Kim, Charlotte Tumescheit, Milot Mirdita, Jeongjae Lee, Cameron L. M. Gilchrist, Johannes Söding, and Martin Steinegger. Fast and accurate protein structure search with Foldseek . Nature Biotechnology, 42 0 (2): 0 243--246, February 2024. ISSN 1...

  43. [51]

    AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences

    Mihaly Varadi, Damian Bertoni, Paulyna Magana, Urmila Paramval, Ivanna Pidruchna, Malarvizhi Radhakrishnan, Maxim Tsenkov, Sreenath Nair, Milot Mirdita, Jingi Yeo, Oleg Kovalevskiy, Kathryn Tunyasuvunakool, Agata Laydon, Augustin Z \'i dek, Hamish Tomlinson, Dhavanthi Harihara...

  44. [52]

    Wang, Paul M

    Connie Y. Wang, Paul M. Chang, Marie L. Ary, Benjamin D. Allen, Roberto A. Chica, Stephen L. Mayo, and Barry D. Olafson. ProtaBank : A repository for protein design and engineering data. Protein Science, 27 0 (6): 0 1113--1124, 2018. ISSN 1469-896X. doi:10.1002/pro.3406. URL h...

  45. [53]

    GraphBind : protein structural context embedded rules learned by hierarchical graph neural networks for recognizing nucleic-acid-binding residues

    Ying Xia, Chun-Qiu Xia, Xiaoyong Pan, and Hong-Bin Shen. GraphBind : protein structural context embedded rules learned by hierarchical graph neural networks for recognizing nucleic-acid-binding residues. Nucleic Acids Research, 49 0 (9): 0 e51, May 2021. ISSN 0305-1048. doi:10...

  46. [54]

    Lal, James C

    Jason Yang, Ravi G. Lal, James C. Bowden, Raul Astudillo, Mikhail A. Hameedi, Sukhvinder Kaur, Matthew Hill, Yisong Yue, and Frances H. Arnold. Active learning-assisted directed evolution. Nature Communications, 16 0 (1): 0 714, January 2025. ISSN 2041-1723. doi:10.1038/s41467...

  47. [55]

    Yang, Zachary Wu, and Frances H

    Kevin K. Yang, Zachary Wu, and Frances H. Arnold. Machine-learning-guided directed evolution for protein engineering. Nature Methods, 16 0 (8): 0 687--694, August 2019. ISSN 1548-7105. doi:10.1038/s41592-019-0496-6. URL https://www.nature.com/articles/s41592-019-0496-6. Publis...

  48. [56]

    Masked inverse folding with sequence transfer for protein representation learning

    Kevin K Yang, Niccol \`o Zanichelli, and Hugh Yeh. Masked inverse folding with sequence transfer for protein representation learning. Protein Engineering, Design and Selection, 36: 0 gzad015, January 2023. ISSN 1741-0126. doi:10.1093/protein/gzad015. URL https://doi.org/10.109...

  49. [57]

    Yang, Nicolo Fusi, and Alex X

    Kevin K. Yang, Nicolo Fusi, and Alex X. Lu. Convolutions are competitive with transformers for protein sequence pretraining. Cell Systems, 15 0 (3): 0 286--294.e2, March 2024. ISSN 2405-4712, 2405-4720. doi:10.1016/j.cels.2024.01.008. URL https://www.cell.com/cell-systems/abst...

  50. [58]

    Goodsell, Maria Voigt, and Stephen K

    Christine Zardecki, Shuchismita Dutta, David S. Goodsell, Maria Voigt, and Stephen K. Burley. RCSB Protein Data Bank : A Resource for Chemical , Biochemical , and Structural Explorations of Large and Small Biomolecules . Journal of Chemical Education, 93 0 (3): 0 569--575, Mar...

  51. [59]

    TM -align: a protein structure alignment algorithm based on the TM -score

    Yang Zhang and Jeffrey Skolnick. TM -align: a protein structure alignment algorithm based on the TM -score. Nucleic Acids Research, 33 0 (7): 0 2302--2309, April 2005. ISSN 0305-1048. doi:10.1093/nar/gki524. URL https://doi.org/10.1093/nar/gki524

  52. [60]

    Wayment-Steele, Garyk Brixi, Haobo Wang, Dorothee Kern, and Sergey Ovchinnikov

    Zhidian Zhang, Hannah K. Wayment-Steele, Garyk Brixi, Haobo Wang, Dorothee Kern, and Sergey Ovchinnikov. Protein language models learn evolutionary statistics of interacting sequence motifs. Proceedings of the National Academy of Sciences, 121 0 (45): 0 e2406285121, November 2...

  53. [61]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  54. [62]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  55. [63]

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.