REVIEW 4 major objections 5 minor 19 references
Into the Unknown: From Structure to Disorder in Protein Function Prediction
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read General protein function predictors fail to produce specific Gene Ontology predictions on disordered proteins, while disorder-trained FAIDR does.
desk verdict A useful review of IDR function prediction whose central empirical claim is weaker than the prose suggests; the dis-CAFA proposal is the real contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison is carried by tinfo, the information content of a predicted GO term (a measure of how specific the term is within the GO hierarchy), applied to predictions from four general predictors and FAIDR on proteins stratified into fully structured, IDR-containing or conditionally folding, and fully disordered groups. FAIDR is the contrasting machinery: it represents each IDR by evolutionary signatures, Z-scores of compositional, physicochemical, repeat, and motif features computed over evolutionarily related IDRs, and uses regularized logistic regression to assign GO terms at both protein and IDR level. This representation is what lets FAIDR capture sequence-function signals that multiple sequence alignments miss.
What would settle it
Run DeepGOPlus, DeepFRI, Sprof-GO, StarFunc, and FAIDR on a pre-registered set of fully disordered proteins with experimentally validated GO annotations from DisProt and compare precision and recall at matched specificity levels; if general predictors retrieve correct specific terms as often as FAIDR, the paper's central claim would be contradicted.
Extended reading notes
Core claim
According to the authors, the sequence-to-function link in IDRs is real but encoded differently from folded domains: through short linear motifs, molecular recognition features, post-translational modifications, and bulk properties such as charge patterning and composition, not through a stable fold. The central empirical claim is their Figure 2 comparison: for proteins grouped by disorder content, the cumulative median information content (tinfo) of Gene Ontology predictions from DeepGOPlus, DeepFRI, Sprof-GO, and StarFunc falls as disorder increases, and for fully disordered proteins the predictors return few confident terms, typically GO terms with little or no information content. FAIDR, trained on IDR features, yields markedly more informative predictions on these same proteins. From this the authors conclude that general-purpose models remain biased toward structured regions, that reliance on structure and homology explains the failure, and that IDR-specific representations rather than larger models are the path forward.
Load-bearing premise
The load-bearing premise is that the information content of a GO prediction is a valid proxy for prediction quality; tinfo says how specific a term is, not whether it is the right term, and the disorder comparison rests on a small set of test proteins.
Editorial extensions
If this is right
- Current benchmarks such as CAFA overstate the practical value of general predictors for disordered proteins, because IDR-containing proteins and IDPs are largely absent from their evaluation sets.
- For proteins with substantial disorder, disorder-trained tools like FAIDR are a better starting point than general predictors, with more specific GO annotations across DNA-, RNA-, and membrane-related functions.
- Model evaluation should stratify test proteins by disorder content, since strong performance on folded proteins does not transfer to disordered proteins.
- A formal community challenge for IDR function prediction and a dedicated disordered-function ontology would provide the standardized benchmarks and training labels needed to make progress.
Reading between the lines
- Editorial inference: If the information-content gap reflects a real accuracy gap, then incorporating disordered regions into protein language model training corpora, rather than scaling model size, is the most direct route to better predictions for IDRs.
- Editorial inference: tinfo measures specificity, not correctness, so a decisive test would compare the same predictors on IDPs with experimentally confirmed DisProt annotations using precision and recall, not just information content.
- Editorial inference: The one-to-many mapping between IDR sequences and functions implies that evaluation schemes should reward predicting multiple coexisting functions per region, rather than a single best GO term.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a review of computational methods for predicting the functions of intrinsically disordered regions (IDRs) and intrinsically disordered proteins (IDPs). It surveys alignment-free sequence comparison (SHARK, SHARK-dive), IDR-specific functional annotation (FAIDR), binding-site and function predictors for disordered regions, protein language model-based function prediction, and the role of community challenges such as CAFA, CAID, and CASP. The review also proposes a 'dis-CAFA' community challenge and a future framework for residue-level GO annotation across the order-disorder continuum. In addition to the review material, the authors present an original analysis (Figure 2) that compares four general-purpose protein function predictors (DeepGOPlus, DeepFRI, StarFunc, Sprof-GO) and FAIDR on a set of structured, IDR-containing, and fully disordered proteins, using the cumulative median information content (tinfo) of GO predictions as the evaluation metric.
Significance. The review is timely and broad, and the proposal of a dedicated community evaluation for IDR function prediction ('dis-CAFA') is constructive. The authors correctly identify a real gap: general-purpose function predictors are rarely trained or evaluated on disordered regions, and the review's synthesis of recent alignment-free and IDR-aware methods is useful to the community. The descriptions of individual methods are generally accurate and well-cited. However, the original empirical claim in Figure 2—that general predictors underperform on IDPs/IDRs and that FAIDR yields 'markedly more informative predictions'—is not supported by the evidence as presented. The chosen metric, tinfo, measures GO-term specificity, not prediction correctness, and the test set is small and incompletely described. Because this claim is central to the abstract and conclusions, the manuscript needs substantial revision before the empirical assertion can be accepted.
major comments (4)
- [Section 'Out There' (p. 12) and Figure 2] The central empirical claim that 'all predictors returned few confident predictions' for IDPs and that 'FAIDR yields markedly more informative predictions on IDRs and IDPs' rests entirely on comparing cumulative median tinfo. As defined in the Glossary, tinfo is the negative logarithm of a GO term's annotation frequency or descendant count; it quantifies how specific a term is in the ontology, not whether the term is correctly assigned to the protein. A method that emits rare but incorrect terms can arbitrarily achieve high tinfo, while a method that correctly assigns broad terms receives low tinfo. This is not a minor nuance: Figure 2 includes a row for experimentally determined GO annotations, and if those annotations for IDPs are themselves broad (low tinfo), then the low-tinfo predictions of DeepGOPlus, DeepFRI, StarFunc, and Sprof-GO may be well calibrated, and FAIDR's high tinfo may reflect an over-specific label distribution rather than better function prediction. The authors need to validate the metric against correctness (e.g., precision/recall against experimental annotations, or F-max) or reframe Figure 2 as a comparison of output specificity. Without this, the conclusion that general predictors underperform on IDRs is not established.
- [Section 'Out There' (p. 12) and Figure 2] The figure is not reproducible from the manuscript. The text does not state how many proteins are in each of the three categories, how they were selected, or which disorder classification method was used. The caption says 'The Uniprot IDs of tested proteins are listed at the bottom of the figure' but the full list is not in the text, and only three examples (P62942, P37840, P0C671) are discussed. No prediction confidence thresholds are defined, so 'few confident predictions' cannot be verified. No error bars or statistical tests are provided, so the difference in median tinfo between predictors could be within sampling noise. Please provide the complete protein list, category definitions, the version and run date of each predictor, the thresholds used, and per-protein tinfo values, or a link to a data/code repository.
- [Section 'Out There' (p. 12) and Figure 2] The comparison is not like-for-like. FAIDR is trained on evolutionary signatures of IDRs and outputs GO terms from an IDR-relevant restricted vocabulary, while the four general predictors are trained on full-GO annotations from large protein databases. Because tinfo depends on the annotating term's position in the GO hierarchy, a predictor that only outputs a curated set of specific IDR terms will, on average, have higher tinfo than a predictor that must also output broad parent terms for proteins with no specific annotation. The 'markedly more informative predictions' may therefore reflect a difference in vocabulary design rather than in functional insight. The authors should restrict all predictors to a common set of GO terms (or to the FAIDR vocabulary), or use metrics that normalize for label prior, and discuss this limitation explicitly.
- [Section 'Out There' (p. 12) and Figure 2] The benchmarking is not independent: the authors develop FAIDR (and SHARK, SHARK-dive, SHARK-capture) and also select the test proteins and the evaluation metric. While this does not by itself invalidate the claim, it raises the bar for the evidence required. Given the problems with the metric and the incomplete description of the test set, the lack of external validation is a substantive limitation. We recommend either removing the original comparison from the review or substantially strengthening it with an established benchmark (e.g., CAFA-style evaluation with accuracy-based metrics) conducted or confirmed by an independent group.
minor comments (5)
- [Glossary and Figure 2 caption] The Glossary defines tinfo as a measure of the specificity of a GO term, but the Figure 2 caption says 'Higher tinfo reflects more specific and complete functional annotation.' A specific term can be incomplete or incorrect; we suggest using 'specific' consistently and avoiding 'complete'.
- [Table 1] In Table 1, the Sprof-GO row contains an uppercase 'X' in the homology-info column while other entries use lowercase 'x'; also, the 'Other input type' column mixes labels such as 'homology info' and 'textual description of proteins' without a legend. Please make the symbols and column headers consistent.
- [Section 'Out There' (p. 12)] The text says 'we observed' and 'we evaluated' without stating that Figure 2 is an original analysis performed for this review. Please state explicitly in the text that Figure 2 is a new analysis and provide the date and versions of the tools used.
- [Section 'Out There' (p. 12)] The sentence 'the overall information content of the predictions decreased as disorder content increased' is presented as a general trend, but Figure 2 reports a median over a small set of proteins without error bars. Please qualify this as applying to the small set of proteins analyzed here.
- [Abstract and Conclusions] The abstract and conclusions state that 'many predictors neglect or underperform on IDRs' as an established fact, but the only direct evidence in this manuscript is Figure 2. Please either cite prior systematic evaluations for this broad claim or soften the wording to reflect the preliminary nature of the analysis presented.
Circularity Check
No significant circularity: Figure 2 is an empirical benchmark, and the tinfo metric raises validity concerns rather than a circular derivation.
full rationale
This manuscript is a review with one original empirical contribution, the benchmark in Figure 2. The central claim that general-purpose predictors underperform on IDRs/IDPs and that FAIDR yields more informative predictions is based on a direct comparison of cumulative median tinfo across predictors. This is a reported measurement, not a derivation in which the input already contains the conclusion. The authors cite their own tools and prior work (SHARK, SHARK-capture, FAIDR; references 23, 26, 28, 30, 49, 50, 58, 64), but those citations are used to describe published methods and results, and the Figure 2 comparison is newly computed here rather than inherited from those citations. The tinfo metric is defined in the glossary as GO-term specificity; treating tinfo as 'informativeness' and inferring that low tinfo means predictors 'struggle' is a metric-validity concern because tinfo does not measure prediction correctness, but this is not a circularity. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from self-citations, and no ansatz is smuggled in via citation. The main evaluative weakness is the unvalidated assumption that higher tinfo equals better function prediction, which belongs under correctness risk rather than circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption GO term information content (tinfo) is a meaningful proxy for functional prediction quality.
- domain assumption The proteins selected for Figure 2 are representative of fully structured, IDR-containing, and fully disordered protein classes.
- domain assumption Experimental GO annotations used as the reference in the benchmark are complete and accurate.
Cite this review
Pith. "Pith review of Into the Unknown: From Structure to Disorder in Protein Function Prediction." pith.science (2026). https://pith.science/paper/GNYPHFTU
@misc{pith2026250606004,
author = {Pith},
title = {Pith review of: Into the Unknown: From Structure to Disorder in Protein Function Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/GNYPHFTU}},
note = {Machine review of arXiv:2506.06004}
}
read the original abstract
Intrinsically disordered regions (IDRs) account for one-third of the human proteome and play essential biological roles. However, predicting the functions of IDRs remains a major challenge due to their lack of stable structures, rapid sequence evolution, and context-dependent behavior. Many predictors of protein function neglect or underperform on IDRs. Recent advances in computational biology and machine learning, including protein language models, alignment-free approaches, and IDR-specific methods, have revealed conserved bulk features and local motifs within IDRs that are linked to function. This review highlights emerging computational methods that map the sequence-function relationship in IDRs, outlines critical challenges in IDR function annotation, and proposes a community-driven framework to accelerate interpretable functional predictions for IDRs.
Figures
Reference graph
Works this paper leans on
-
[1]
1 Forman-Kay, J. D. and Mittag, T. (2013) From sequence and forces to structure, function, and evolution of intrinsically disordered proteins. Structure, Elsevier BV 21, 1492–1499 https://doi.org/10.1016/j.str.2013.08.001 2 Lemke, E. A., Babu, M. M., Kriwacki, R. W., Mittag, T., Pappu, R. V., Wright, P. E., et al. (2024) Intrinsic disorder: A term to defi...
arXiv 2013
-
[6]
bioRxiv https://doi.org/10.1101/2025.01.31.635962 80 Buchan, D
Predicting molecular recognition features in protein sequences with MoRFchibi 2.0. bioRxiv https://doi.org/10.1101/2025.01.31.635962 80 Buchan, D. W. A., Moffat, L., Lau, A., Kandathil, S. M. and Jones, D. T. (2024) Deep learning for the PSIPRED Protein Analysis Workbench. Nucleic Acids Res., Oxford University Press (OUP) 52, W287–W293 https://doi.org/10....
-
[7]
Evolutionary analyses of IDRs reveal widespread signals of conservation. bioRxiv https://doi.org/10.1101/2023.12.05.570250 60 Strome, B., Elemam, K., Pritisanac, I., Forman-Kay, J. D. and Moses, A. M. (2023, April
-
[10]
bioRxiv https://doi.org/10.1101/2024.03.25.586696 89 Garg, S
Foldclass and Merizo-search: embedding-based deep learning tools for protein domain segmentation, fold recognition and comparison. bioRxiv https://doi.org/10.1101/2024.03.25.586696 89 Garg, S. G. and Hochberg, G. K. A. (2025) A general substitution matrix for structural phylogenetics. Mol Biol Evol https://doi.org/10.1093/molbev/msaf124 90 Gligorijević, V...
-
[11]
bioRxiv https://doi.org/10.1101/2022.02.10.480018 34 King, M
Sequence- and chemical specificity define the functional landscape of intrinsically disordered regions. bioRxiv https://doi.org/10.1101/2022.02.10.480018 34 King, M. R., Ruff, K. M., Lin, A. Z., Pant, A., Farag, M., Lalmansingh, J. M., et al. (2024) Macromolecular condensation organizes nucleolar sub-phases to set up a pH gradient. Cell 187, 1889–1906.e24...
-
[12]
Protein language models are biased by unequal sequence sampling across the tree of life. bioRxiv https://doi.org/10.1101/2024.03.07.584001 125 Unsal, S., Atas, H., Albayrak, M., Turhan, K., Acar, A. C. and Doğan, T. (2022) Learning functional properties of proteins with language models. Nat. Mach. Intell., Springer Science and Business Media LLC 4, 227–24...
-
[13]
MSA Transformer. bioRxiv, bioRxiv https://doi.org/10.1101/2021.02.12.430858 114 Elnaggar, A., Heinzinger, M., Dallago, C., Rehawi, G., Wang, Y., Jones, L., et al. (2022) ProtTrans: Toward understanding the language of life through self-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell., Institute of Electrical and Electronics Engineers (IEEE) 44...
arXiv 2022
-
[16]
Deep learning of proteins with local and global regions of disorder. arXiv [q-bio.BM] 141 Liu, Z. H., Tsanai, M., Zhang, O., Forman-Kay, J. and Head-Gordon, T. (2024) Computational methods to investigate intrinsically disordered proteins and their complexes. ArXiv 142 Home - CASP15 https://predictioncenter.org/casp15/index.cgi 143 CAPRI Docking. CAPRI Doc...
work page 2024
Show all 19 references
-
[17]
bioRxiv https://doi.org/10.1101/2024.03.15.585291 31 González-Foutel, N
A functional map of the human intrinsically disordered proteome. bioRxiv https://doi.org/10.1101/2024.03.15.585291 31 González-Foutel, N. S., Glavina, J., Borcherds, W. M., Safranchik, M., Barrera-Vilarmau, S., Sagar, A., et al. (2022) Conformational buffering underlies functi...
2022 doi
-
[18]
bioRxiv https://doi.org/10.1101/2024.05.15.594113 96 Zhu, Y.-H., Zhang, C., Yu, D.-J
StarFunc: fusing template-based and deep learning approaches for accurate protein function prediction. bioRxiv https://doi.org/10.1101/2024.05.15.594113 96 Zhu, Y.-H., Zhang, C., Yu, D.-J. and Zhang, Y. (2022) Integrating unsupervised language model with triplet neural network...
-
[19]
bioRxiv https://doi.org/10.1101/2023.08.09.552672 156 Ginell, G
MELISSA: Semi-supervised embedding for protein function prediction across multiple networks. bioRxiv https://doi.org/10.1101/2023.08.09.552672 156 Ginell, G. M., Emenecker, R. J., Lotthammer, J. M., Usher, E. T. and Holehouse, A. S. (2024) Direct prediction of intermolecular i...
2024 doi
-
[21]
bioRxiv https://doi.org/10.1101/2024.12.18.629275 74 Pang, Y
Probabilistic annotations of protein sequences for intrinsically disordered features. bioRxiv https://doi.org/10.1101/2024.12.18.629275 74 Pang, Y. and Liu, B. (2022) DMFpred: Predicting protein disorder molecular functions based on protein cubic language model. PLoS Comput. B...
2022 doi
-
[22]
bioRxiv https://doi.org/10.1101/2023.02.22.529574 53 Mollaei, P., Sadasivam, D., Guntuboina, C
DR-BERT: A protein language model to annotate disordered regions. bioRxiv https://doi.org/10.1101/2023.02.22.529574 53 Mollaei, P., Sadasivam, D., Guntuboina, C. and Farimani, A. B. (2024) IDP-Bert: Predicting properties of intrinsically Disordered Proteins (IDP) using large l...
-
[23]
Bio Function Prediction https://biofunctionprediction.org/cafa/ 147 Rostam, N., Ghosh, S., Chow, C
CAFA. Bio Function Prediction https://biofunctionprediction.org/cafa/ 147 Rostam, N., Ghosh, S., Chow, C. F. W., Hadarovich, A., Landerer, C., Ghosh, R., et al. (2023) CD- CODE: crowdsourcing condensate database and encyclopedia. Nat. Methods 20, 673–676 https://doi.org/10.103...
2023 doi
-
[26]
bioRxivorg https://doi.org/10.1101/2023.11.26.568742 109 Tran, N
Learning sequence, structure, and function representations of proteins with language models. bioRxivorg https://doi.org/10.1101/2023.11.26.568742 109 Tran, N. C. and Gao, J. X. (2023) Integrating heterogeneous biological networks and ontologies for improved protein function pr...
2023
-
[28]
bioRxiv, bioRxiv https://doi.org/10.1101/2020.06.26.174417 124 Ding, F
BERTology meets biology: Interpreting attention in protein language models. bioRxiv, bioRxiv https://doi.org/10.1101/2020.06.26.174417 124 Ding, F. and Steinhardt, J. (2024, March
2020 doi
-
[29]
bioRxiv https://doi.org/10.1101/2023.04.28.538739 61 Martínez-Pérez, E., Pajkos, M., Tosatto, S
Computational design of intrinsically disordered protein regions by matching bulk molecular properties. bioRxiv https://doi.org/10.1101/2023.04.28.538739 61 Martínez-Pérez, E., Pajkos, M., Tosatto, S. C. E., Gibson, T. J., Dosztanyi, Z. and Marino-Buslje, C. (2023) Pipeline fo...
2023
-
[30]
arXiv [cs.CL] 155 Wu, K., Zhou, D., Slonim, D., Hu, X
Domain-specific language model pretraining for biomedical natural language processing. arXiv [cs.CL] 155 Wu, K., Zhou, D., Slonim, D., Hu, X. and Cowen, L. (2023, August
2023
-
[2025]
and Buday, L
Nucleic Acids Res., Oxford University Press (OUP) 53, D444–D456 https://doi.org/10.1093/nar/gkae1082 43 Tompa, P., Szász, C. and Buday, L. (2005) Structural disorder throws new light on moonlighting. Trends Biochem. Sci., Elsevier BV 30, 484–489 https://doi.org/10.1016/j.tibs....
2005 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.