{"id":"4fcb695c-0567-4a37-9905-31e3710de229","arxiv_id":"2504.16886","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Disordered regions are common in ProteinGym and consistently degrade zero-shot fitness prediction across model classes, while simple multi-modal ensembles remain the strongest baselines.","lead":"This paper tests how structure choice, disordered regions, and model combinations affect zero-shot protein fitness prediction on the ProteinGym benchmark. It finds that disordered regions hurt most models, that predicted structures often beat experimental ones for one inverse folding model, and that simple sequence-plus-structure ensembles are strong baselines.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Figure 3 disorder-vs-ordered comparison lacks significance tests, so the central claim of a systematic disorder penalty is not statistically secured.","rationale":"The paper's main novel contribution is the disorder analysis, and its central claim is that disordered regions systematically reduce zero-shot fitness prediction quality. The supporting evidence is descriptive: Figure 3 shows mean Spearman correlations for ordered vs disordered mutations with no error bars, confidence intervals, or significance tests. This is a genuine gap because per-assay Spearman correlations for disordered regions are computed on small numbers of mutations (many disordered regions in Table 4 are just a few residues), making them noisy. A few large assays could dominate the aggregate. The reader's chosen weakest assumption about experimental structure matching is real but targets the secondary predicted-versus-experimental structure comparison, not the disorder claim. The paper itself acknowledges the lack of manual validation in the appendix, so that concern is already admitted and does not undermine the central contribution. In contrast, the statistical robustness of the disorder effect is unaddressed and directly determines whether the headline finding holds. Since the paper's code and data are available, the proposed paired Wilcoxon test is immediately runnable and would settle the question. I therefore keep the reader's CONDITIONAL verdict: the concern is addressable with additional analysis rather than fatal. The reader did identify the missing uncertainty quantification in their rationale, so agreement is partial, but their formal weakest-assumption field points elsewhere.","tokens_in":20225,"tokens_out":4496,"duration_ms":41783,"concrete_test":"For each model shown in Figure 3, compute the per-assay difference in Spearman correlation Δρ = ρ_ordered − ρ_disordered across the 43 assay set; run a one-sided Wilcoxon signed-rank test against Δρ = 0 and report bootstrap 95% confidence intervals for the median difference. If the median Δρ is not significantly positive for most models, the systematic-disorder claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, that disordered regions systematically reduce zero-shot fitness prediction quality across model families, rests on Figure 3, which compares mean Spearman correlations for mutations in ordered versus disordered regions across 43 assays. The analysis is purely descriptive: no error bars, confidence intervals, or significance tests are reported. This is a critical gap because many disordered regions in Table 4 are extremely short (e.g., a single residue, as in RCD1_ARATH_Tsuboyama_2023_5OAO with region 0-0), so the per-assay Spearman correlation within disordered regions is computed on very few mutations and is highly variable. The consistent downward direction across models could be driven by a few high-leverage assays with large disordered regions (e.g., BRCA1, p53) rather than a systematic effect. Without a paired significance test across the 43 assays that have mutations in both regions, the claim that disorder 'affect[s]' predictions across multiple model types is not established. This is the most load-bearing weakness because it directly targets the main novel contribution of the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper benchmarks zero-shot structure-based protein fitness predictors on ProteinGym substitution assays, comparing predicted (AlphaFold 2) structures against automatically matched experimental PDB structures, and analyzes how intrinsically disordered regions affect prediction quality across structure-, sequence-, MSA-, and multimodal models. It reports that disordered regions are widespread in ProteinGym, that structure-based models and simple multimodal ensembles perform competitively, and that ensemble methods are strong baselines. The main empirical claims are that predicted structures often outperform automatically matched experimental structures and that mutations in disordered regions are predicted worse than mutations in ordered regions across most model families and function types.","tokens_in":20416,"tokens_out":2991,"duration_ms":29036,"significance":"If the disorder-related claim survives statistical scrutiny, this is a useful and broadly relevant observation for the protein fitness prediction community, because it connects a biological property (disorder) to the behavior of many distinct model families and to the choice of structural input. The study's strengths include the use of standard external benchmarks (ProteinGym, DisProt, PDB), the breadth of models and ensembles considered, and the public release of code and data for reproduction. The multi-modal ensemble results are a useful baseline reference even though the primary novelty lies in the disorder analysis.","major_comments":[{"comment":"The central claim that disordered regions systematically reduce zero-shot fitness prediction quality across model families is supported only by descriptive per-function means in Figure 3; no confidence intervals, error bars, or paired significance tests are reported for the 43 assays. Since Table 4 includes disordered regions as short as a single residue (e.g., RCD1_ARATH_Tsuboyama_2023_5OAO, region 0-0), the per-assay Spearman correlations within disordered regions are computed on very few mutations and are highly variable. Please add a paired test across the 43 assays (e.g., Wilcoxon signed-rank on per-assay ordered-vs-disordered Spearman differences) with effect sizes and confidence intervals, and perform the same analysis per function type where sample sizes permit.","section":"Results, 'Intrinsically disordered regions...' and Figure 3; Tables 4-5"},{"comment":"Because experimental structures were selected by sequence identity and full-length coverage without manual validation (as stated in the Appendix), and because ESM-IF1 is explicitly trained on AlphaFold 2 predicted structures (as noted in the Discussion), Table 1's claim that predicted structures outperform experimental structures is not established as a general property of structure-based fitness prediction. The alpha-synuclein example (Figure 5) shows that the automatically selected experimental structure can represent a different conformation from the functional one tested in the assay. Please restrict the claim to the 65 automatically matched PDB structures used here, or validate a subset of structures against the functional conformations described in the DMS assay literature.","section":"Table 1, Figure 1, and Appendix 'Selection process for experimental structures'"},{"comment":"The disorder-annotation mapping treats same-length target sequences as referring to the same sequence even though the authors state that ProteinGym target sequences can differ from UniProt sequences at arbitrary positions, and subset sequences are located within UniProt without a described alignment algorithm. Misplacement of disordered annotations would directly bias the ordered-versus-disordered comparison in Figure 3. Please specify the alignment method precisely and provide a sensitivity analysis restricted to assays with exact sequence matching or a conservative alignment confidence threshold.","section":"Appendix, 'Selection process for disordered proteins', Stage 2"}],"minor_comments":[{"comment":"There is a typo in the header: 'Code a v ailability' should be 'Code availability'.","section":"Full text, header"},{"comment":"The caption says 'Ensembled 3' while Table 2 and the methods use 'Ensemble 3'; please make the naming consistent.","section":"Table 6 caption"},{"comment":"The manuscript reports 216 assays used in the study (Table 3 and the Appendix) but also states that 63 of 217 ProteinGym assays have disordered regions; the excluded assay BRCA2_HUMAN_Erwood_2022_HEK293T still appears in Table 4. Please reconcile the denominators and clarify which assays are included in each count.","section":"Table 3 and Table 4"},{"comment":"The caption says scores are 'averaged across UniProtID and then function type' while the text describes 43 assays; please clarify the exact aggregation hierarchy and how many assays contribute to each point.","section":"Figure 3 caption"},{"comment":"The column labels 'Count of DMS assays with positive and negative ρpred−ρexp' are unclear; please use explicit labels such as 'Number of assays with ρpred − ρexp ≥ 0' and 'Number of assays with ρpred − ρexp < 0'.","section":"Table 1 header"}],"recommendation":"major_revision","confidential_remarks":"The paper is a workshop-length empirical study, and the main revision should focus on adding statistical support to the disorder analysis, which is the most novel and load-bearing contribution. The experimental-structure comparison is adequately framed as exploratory, but the abstract's wording ('confirm the importance of matching protein structures to fitness assays') overstates what the automatic PDB matching procedure can establish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a competent, reproducible benchmark study, and the residue-level disorder analysis is the real contribution. It maps DisProt disorder annotations onto 63 ProteinGym assays, then compares Spearman correlations for mutations in ordered vs disordered regions across 12 models including structure-based, sequence, MSA, multi-modal, and simple ensembles. For the 43 assays with mutations in both region types, the ordered-region correlations are consistently higher across almost every model and function type. That is a new, measurement-like result; how to model or mask disordered regions is something the field still does badly.\n\nThe paper also does a few things well: the SSEmb evaluation is a useful addition to the ProteinGym zoo, the simple-ensembles-are-strong point is a good reminder for people building fancier architectures, and the alpha-synuclein case study is a nice concrete illustration of why structure matching matters. Code and data are on GitHub/Zenodo, which counts for a lot.\n\nThe soft spots are real but addressable. The stress-test worry about Figure 3 is on target: no confidence intervals, no paired significance test across the 43 assays, and some disordered regions are a single residue, so the per-assay disordered-set Spearman is noisy. The direction is consistent across models, but the claim that disorder 'systematically' hurts prediction is not statistically pinned down. A paired Wilcoxon signed-rank test or bootstrap CIs would likely settle it. Second, the predicted-vs-experimental structure comparison (Table 1) rests on PDB entries selected by sequence identity and full-length coverage, without manual validation against the DMS assay conditions. The authors admit this openly, and they give a mitigating case study, but the comparison remains vulnerable to mismatched baselines. That's a moderate weakness, not a fatal one.\n\nThe paper is a workshop paper with workshop-sized ambitions; it doesn't introduce a new method or mechanism. Who is it for? Anyone working on structure-based fitness prediction or variant interpretation in disordered proteins. It deserves a serious referee: with paired significance tests and a better-described experimental structure curation it could be a solid benchmark contribution. I'd accept it for peer review with revisions recommended.","headline":"A clean, reproducible benchmark paper whose residue-level disorder analysis is a genuinely useful new measurement, but the headline claim would be stronger with paired significance tests and validated experimental structures.","tokens_in":20875,"tokens_out":2663,"would_cite":true,"duration_ms":24384,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Disordered protein regions drag down zero-shot fitness prediction accuracy","keywords":["zero-shot fitness prediction","ProteinGym","intrinsically disordered regions","structure-based models","AlphaFold 2","deep mutational scanning","ensembles"],"falsifier":"If one repeated the 65-assay structure comparison using manually validated experimental structures that exactly match the functional assay condition (for example, the micelle-bound helical α-synuclein instead of a fibril), and the advantage of AlphaFold 2 predicted structures over experimental structures disappeared, the claim that predicted structures are generally as good or better for fitness prediction would be falsified. This test is feasible because the authors identify the problematic case and the required structure is known.","tokens_in":20033,"feed_emoji":"🧬","tokens_out":1834,"duration_ms":18254,"temperature":0.7,"pith_summary":"Zero-shot fitness prediction models that score how mutations affect protein function perform worse when a mutation falls in a disordered region—a stretch of the protein with no fixed 3D structure. This paper shows that such regions appear in 29% of ProteinGym substitution assays and that the performance drop occurs across sequence-based, MSA-based, structure-based, multi-modal, and ensemble models, for most measured function types. The finding matters because disordered regions are common in real proteins, and fitness predictors are used to interpret genetic variants and guide protein engineering. The paper also examines how the choice of input structure (predicted versus experimental) affects predictions, and it proposes simple multi-modal ensembles as strong baselines for this task.","feed_headline":"Disordered regions blunt protein fitness predictors","feed_subtitle":"Across model types, mutations in unstructured protein stretches score with far lower accuracy—a systematic blind spot for variant…","key_machinery":"The central objects are the 217 deep mutational scanning substitution assays in ProteinGym, the matched AlphaFold 2 predicted structures and PDB experimental structures, and DisProt annotations of intrinsic disorder. The key comparison is per-mutation: separating each assay's mutations into those in disordered versus ordered regions and comparing model Spearman correlations on those two subsets. For the structure comparison, the machinery is the paired difference ρ_pred − ρ_exp for each assay, split into monomers and multimers, plus case studies on α-synuclein (P37840) and NKX3-1 (Q99801) that visually align experimental and predicted structures to explain extreme differences.","core_discovery":"The paper establishes that intrinsically disordered regions systematically degrade zero-shot protein fitness prediction: in 43 ProteinGym DMS assays that contain mutations in both ordered and disordered regions, Spearman correlations between predicted and measured fitness are consistently lower for disordered-region mutations across models of all input modalities, including structure-based ESM-IF1, multi-modal ProtSSN, SaProt, TranceptEVE L, sequence-based ESM2, MSA-based GEMME, and several simple ensembles. The disorder effect holds for most function types but is less pronounced for stability assays. Separately, the paper finds that AlphaFold 2 predicted structures often yield higher Spearman correlations than experimental structures selected from the PDB, cautioning that structure–assay matching is critical and that predicted structures in disordered regions can be misleading—illustrated with α-synuclein, where the predicted structure adopts a fibril-like conformation while the experimental structure is helical, and the assay's function depends on the membrane-bound helical state.","pith_inferences":["The disorder penalty may be partly an artifact of conservation: disordered regions evolve faster and are less conserved, so evolutionary-signal models (PLMs trained on sequence databases, MSA-based models) will naturally assign them flatter likelihoods; this suggests the drop may be informative rather than purely noise, and one could test whether models trained on disorder-aware data recover predi","A direct extension would be to use the same protocol to quantify disorder penalties on the larger set of DMS assays in MaveDB and ProtaBank, which would test how generalizable the 29% prevalence and the per-function trends are.","Because the paper used only one experimental structure per protein, a natural next study is to sample multiple experimental conformations or use ensemble docking to see whether supplying a more functional conformation (for α-synuclein, the micelle-bound helix) restores ESM-IF1 performance to the level seen for ordered regions.","The finding that simple ensembles outperform joint architectures suggests a 'disorder-aware ensemble' that down-weights structure-based models inside disordered regions could improve overall Spearman correlation; this is a concrete, testable design that the paper does not evaluate."],"forward_implications":["Model developers should treat disordered regions as a distinct failure mode, possibly by masking low-confidence predicted structure coordinates or by conditioning on conformational ensembles rather than a single static structure.","Users of zero-shot fitness predictors for clinical variant interpretation should be cautious when a variant falls in a documented disordered region, since prediction quality there is systematically lower.","Simple ensembles that combine structure-based ESM-IF1 with sequence- or MSA-based models (e.g., StructSeq) can match or exceed more sophisticated multi-modal architectures, so such ensembles should be standard baselines in future ProteinGym-style benchmarks.","Predicted structures from AlphaFold 2, being trained on a distribution that often excludes the functional conformation (as for α-synuclein), can mislead fitness prediction in disordered regions; matching the structure to the assay's actual conformation is essential.","Stability assays stand apart: structure-based models do not show the same disorder-related drop, likely because the measured property (proteolytic stability) is more directly tied to a single folded state."],"supporting_citations":[{"why":"Provides the ProteinGym benchmark, including DMS substitution assays, AlphaFold 2 predicted structures, and zero-shot model scores used throughout.","marker":"Notin et al. 2023"},{"why":"Supplies the DisProt database used to define intrinsically disordered regions in the target sequences.","marker":"Aspromonte et al. 2024"},{"why":"Provides the RCSB PDB resource from which experimental structures were queried and selected.","marker":"Zardecki et al. 2016"},{"why":"Describes ESM-IF1, the structure-based inverse folding model whose predictions are the primary subject of the structure-choice analysis.","marker":"Hsu et al. 2022"},{"why":"Defines the StructSeq ensemble (Ensemble 1) that combines ESM-IF1 and TranceptEVE, and is a key baseline showing simple ensembles outperform joint models.","marker":"Paul et al. 2023"},{"why":"Describes SSEmb, a structure-plus-MSA model that the paper generates fresh MSAs for; its performance underscores strengths of multi-modal input.","marker":"Blaabjerg et al. 2024"},{"why":"Introduces ProtSSN, a structure-plus-PLM ensemble model whose disorder-related performance drop is analyzed in Figure 3b.","marker":"Tan et al. 2024"},{"why":"Introduces SaProt, a structure-token PLM whose disorder sensitivity is compared; it uses pLDDT masking like ESM-IF1.","marker":"Su et al. 2023"},{"why":"Introduces TranceptEVE, a sequence-plus-MSA model whose disorder drop shows the effect is not exclusive to structure-aware models.","marker":"Notin et al. 2022b"},{"why":"Describes AlphaFold 2, the source of predicted structures and the pLDDT confidence scores used by ESM-IF1 and SaProt for masking.","marker":"Jumper et al. 2021"}],"fun_headline_variants":["Disordered regions blight protein fitness prediction","Zero-shot fitness prediction blinds to disorder","Protein fitness models stumble on disordered regions","Disorder dims zero-shot protein fitness scores","Structure-based fitness prediction misses disorder"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the experimental structures retrieved from the PDB by sequence identity and full-length coverage represent the functional conformation tested in each deep mutational scanning assay; the authors did not manually verify this, so some structure-choice comparisons may reflect mismatched conformations rather than a general property of experimental versus predicted structures.","fun_headline_variants_meta":{"raw":{"variants":["Disordered regions blight protein fitness prediction","Zero-shot fitness prediction blinds to disorder","Protein fitness models stumble on disordered regions","Disorder dims zero-shot protein fitness scores","Structure-based fitness prediction misses disorder"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000145,"raw_usage":{"total_tokens":1151,"prompt_tokens":888,"completion_tokens":263,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":504,"tokens_out":263,"duration_ms":3174,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:52:43.674084+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If one repeated the 65-assay structure comparison using manually validated experimental structures that exactly match the functional assay condition (for example, the micelle-bound helical α-synuclein instead of a fibril), and the advantage of AlphaFold 2 predicted structures over experimental structures disappeared, the claim that predicted structures are generally as good or better for fitness prediction would be falsified. This test is feasible because the authors identify the problematic case and the required structure is known.","supporting_citations":[{"cited_title":"DisProt in 2024: improving function annotation of intrinsically disordered proteins","cited_arxiv_id":null,"evidence_quote":"Supplies the DisProt database used to define intrinsically disordered regions in the target sequences."},{"cited_title":"Goodsell, Maria Voigt, and Stephen K","cited_arxiv_id":null,"evidence_quote":"Provides the RCSB PDB resource from which experimental structures were queried and selected."},{"cited_title":"Learning inverse folding from millions of predicted structures","cited_arxiv_id":null,"evidence_quote":"Describes ESM-IF1, the structure-based inverse folding model whose predictions are the primary subject of the structure-choice analysis."}],"review_version":1}