{"id":"5fa950f7-5bab-44b3-8659-84ff4d1a6ec0","arxiv_id":"1908.07662","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"CASP13 assembly predictions improved by roughly 5 to 15 percent over CASP12 when matched by percentiles, driven largely by homology modeling with human intervention, while homomeric interface contacts were predicted well but not used.","lead":"This paper presents the official assessment of protein assembly prediction in CASP13, comparing 45 predictor groups across 42 oligomeric targets. It finds increased participation, modest quality improvement over CASP12, and evidence that ignoring quaternary structure penalizes tertiary prediction.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5–15% progress claim in Section 3.1 rests on an untested target-difficulty equivalence assumption, supported only by an unresolved citation; percentile-matched gains may reflect target selection rather than improved prediction.","rationale":"The assessment is otherwise well-constructed: participation growth is documented (45 groups, 17 submitting >10 targets vs 10 in CASP12; >5000 models), homomeric contact findings are specific and falsifiable, and the baseline comparison is a useful control. The main abstract claim of 'consistent, albeit modest, improvement' is the only part that depends on cross-edition comparability, and the paper does not currently establish it. My concern is not that the conclusion is wrong but that the quantitative basis is not yet demonstrated. This matches the reader's weakest_assumption; the reader also flagged the over-strong 'guarantees' wording, which I agree is secondary. Since the conditional verdict already represents this uncertainty, I do not move the verdict. A revision that supplies the difficulty-matched analysis or qualifies the improvement claim would settle the issue; if the analysis fails, the verdict should move toward reject or 'unverified' for the progress claim.","tokens_in":15583,"tokens_out":5751,"duration_ms":55819,"concrete_test":"Recompute the Section 3.1 percentile-matched comparison within difficulty-matched subsets. Use the authors' own Easy/Medium/Difficult classifications (Table 3.1 and its CASP12 equivalent) or, if those are not published, template-availability features (HHpred probability/coverage and oligomeric stoichiometry) to stratify targets. For each difficulty class, redo the percentile matching and calculate the median score shift with a permutation test that swaps edition labels (10,000 resamples, 95% CI). If the 5-15% shift is not reproduced within each stratum, or if the CI includes 0 for any score, the improvement claim should be reworded to an 'uncontrolled comparison' rather than a quantitative finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 (Performance) makes the central quantitative claim: '5-15% improvement for all scores across the board' via percentile matching of CASP12 and CASP13 predictions. The validity of this comparison hinges on the assumption that 'the difficulty of the assembly targets in CASP12 and CASP13 has roughly the same distribution.' The only support offered is the unresolved placeholder citation '(EVIDENCE IN; REFTHIS YEAR'S DOMAIN PREDICTION ASSESSMENT=)', which is not a bibliographic entry and, even if completed, refers to domain prediction rather than assembly-specific difficulty. No distribution of the three difficulty classes across editions is shown, and no bootstrap or confidence interval is given for the percentile-matched shift. If CASP13's 42 targets included more homomeric, more template-accessible, or smaller assemblies than CASP12's 30, the same prediction methods would produce higher scores at matched percentiles without any real progress. The paper's own solved-target comparison (9/42 vs 6/30) is essentially flat, so it does not independently corroborate improvement. This is the load-bearing weak point because the abstract's 'modest improvement' claim and the concluding 'quality consistently increased' both depend on it. A secondary overstatement—'absence of detectable assembly templates with near-complete coverage guarantees absence of good models'—is template-centric and contradicted by target 40976's successful model from a monomeric template, but it is not central to the progress comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript presents the CASP13 assembly category assessment. The authors evaluate predictions for 42 oligomeric targets using four scores (ICS, IPS, oligomeric lDDT and GDT), define a 'solved' criterion requiring all four scores to exceed 0.5, compare CASP13 with CASP12 by percentile matching of scores, and rank groups using leave-one-out Z-scores. They report increased participation, a modest 5–15% improvement over CASP12, an unchanged solved-target proportion (9/42 vs 6/30), a dominant role for human-assisted homology modeling, successful prediction of homomeric interface contacts that was nevertheless not used in assembly modeling, and no systematic benefit from data-assisted targets except target 80957.","tokens_in":15876,"tokens_out":5118,"duration_ms":51326,"significance":"If the cross-edition comparison is valid, this assessment provides a useful benchmark for protein assembly prediction and a clear statement of the field's state. The paper's strengths include a transparent evaluation protocol, a baseline comparison, the use of four complementary scores, and the novel observation that homomeric interface contacts are predictable but currently unused. However, the headline quantitative claim of 5–15% improvement rests on an untested comparability assumption and lacks uncertainty estimates; the paper's own solved-target comparison is essentially flat and does not independently support the improvement claim. The manuscript would be substantially strengthened by reporting the difficulty-class distributions for CASP12 and CASP13 and by adding confidence intervals or bootstrap estimates for the percentile-matched differences.","major_comments":[{"comment":"The '5–15% improvement for all scores across the board' claim is load-bearing and currently rests on the unsupported assumption that the difficulty of CASP12 and CASP13 assembly targets has roughly the same distribution. The only support offered is the unresolved placeholder citation '(EVIDENCE IN; REFTHIS YEAR'S DOMAIN PREDICTION ASSESSMENT=)', which is not a completed bibliographic entry and, even if completed, refers to domain prediction rather than assembly-specific difficulty. The manuscript itself defines three difficulty classes in §2.2 but never reports their distributions in the two editions, so the assumption cannot be checked from the presented data. Moreover, no confidence intervals or bootstrap estimates accompany the percentile-matched values, and the paper's own solved-target comparison (9/42 vs 6/30) is essentially flat and therefore does not independently corroborate improvement. Please either provide difficulty-class distributions and uncertainty estimates, or weaken the progress claim to something like 'observed score gains are consistent with modest improvement, assuming comparable target difficulty.'","section":"§3.1, Figure 3"},{"comment":"The statement in §3.1 that 'absence of detectable assembly templates with near-complete coverage guarantees absence of good models' is too strong and is internally contradicted by the paper's own example of target 40976 in §3.2, where successful dimeric models were produced from a monomeric template whose interdomain interfaces resembled the dimeric interface. As written, the 'guarantee' is false; if the intended meaning is that this is the dominant pattern for most targets, the sentence should be revised to state that pattern and to specify how templates were defined and detected.","section":"§3.1 vs §3.2"}],"minor_comments":[{"comment":"The placeholder citation '(EVIDENCE IN; REFTHIS YEAR'S DOMAIN PREDICTION ASSESSMENT=)' must be resolved to a proper bibliographic entry before publication; a citation to the domain prediction assessment, even if completed, would not by itself establish equivalence of assembly-target difficulty.","section":"§3.1"},{"comment":"The statement that the maximum and minimum leave-one-out total scores 'can be used to assess the significance of the differences between closely ranked groups' is not a valid significance test; leave-one-out variation measures sensitivity to individual targets, not sampling uncertainty in group ability. The ranking itself is fine as a descriptive result, but the significance language should be removed or accompanied by an appropriate test.","section":"§2.3, Figure 4"},{"comment":"The conclusion that data-assisted scores are 'not significantly different' from regular predictions is made without any statistical test or confidence interval; given the small number of targets (7), the wording should be descriptive rather than inferential, or the relevant test should be reported.","section":"§3.5"},{"comment":"The notation '13-score algorithm' is unclear; the text appears to intend a standard structural alignment score such as TM-score, and the acronym should be written out consistently with the cited reference.","section":"§2.3"},{"comment":"The phrase '6(EASY) TARGETS OUT OF 30' is ambiguous: it is not clear whether all six solved CASP12 targets were in the EASY difficulty class or whether '(EASY)' is a typographical artifact. Please clarify.","section":"§3.1"},{"comment":"The captions describe rich information, but the figures as provided in this manuscript version do not show the target identifiers or score distributions in a way that a reader can use to verify the claims about individual targets; please ensure the published figures include clearly legible labels.","section":"Figures 2 and 3"}],"recommendation":"major_revision","confidential_remarks":"The central quantitative claim is plausible in direction, but the current evidence does not support it as strongly as the abstract and conclusions state. The main revision needed is to make the cross-edition comparison defensible: report the difficulty-class distributions, add uncertainty estimates, and either corroborate or soften the improvement claim. Please also ensure that all placeholders, especially the unresolved citation in §3.1, are replaced with proper references before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good to see this paper. The CASP13 assembly assessment is a solid, transparent follow-up to the CASP12 assessment by the same group, and it does give us genuinely new information. The homomeric-contact finding is the real news: contact prediction groups are already ranking interfacial contacts well, but those predictions are treated as false positives and not used in assembly modelling. That observation alone should push the field to start evaluating homomeric contacts, and it is well supported by the examples and by the ranking change for target 40968S2.\n\nWhat the paper does well is the craft of assessment. The evaluation pipeline is clearly specified: chain mapping for lDDT/GDT, the Seok naive baseline, Z-score outlier removal, leave-one-out ranking, and honest discussion of the target difficulty classes. The case studies (40976, 41001, H0953) are informative and show real thought about why some targets succeed and others fail. The data-assisted section is also measured and does not oversell.\n\nWhere the soft spots are: the central quantitative claim—5–15% improvement across all scores—rests on an assumption that the difficulty distributions of CASP12 and CASP13 assembly targets are roughly the same. That assumption is not tested, and the only citation is an unresolved placeholder pointing to the domain prediction assessment. There are no confidence intervals on the percentile-matched differences, and the paper's own solved-target proportion is flat (6/30 vs 9/42, same fraction), so the improvement claim is not independently corroborated. I would not call this a fatal flaw—the upstream participation increase and the flat solved proportion are consistent with modest, real progress—but the abstract and conclusion say \"consistent improvement\" as if it were established. That needs to be reworded or backed with a proper comparability analysis.\n\nThe minor overstatement is the sentence saying absence of detectable assembly templates \"guarantees\" absence of good models. Target 40976, built from a monomeric template, shows that guarantee is too strong. It is a side remark, not load-bearing.\n\nWho this is for: anyone working on quaternary structure prediction, contact prediction evaluation, or CASP methodology. It deserves a serious referee and likely publication after revision. I would want the authors to either provide evidence of comparable target difficulty or soften the progress claim to match what is actually shown.\n\nMy recommendation: send it to peer review. The assessment is valuable, the methods are transparent, and the homomeric-contact point is worth bringing to a wider audience.","headline":"A transparent, valuable CASP13 assembly assessment whose headline 5–15% improvement claim rests on a comparability assumption the paper itself does not substantiate; the homomeric-contact finding is the strongest new result.","tokens_in":16407,"tokens_out":1833,"would_cite":true,"duration_ms":100326,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CASP13 assembly assessment reports clear uptake in participation, more oligomeric targets, and a consistent, albeit modest, 5–15 percent improvement over CASP12 across all four scoring measures.","keywords":["CASP13","protein assembly prediction","quaternary structure","oligomeric state","interface prediction","homology modelling","contact prediction","structural biology assessment"],"falsifier":"Recompute the percentile-matched comparison with targets stratified by the paper's own easy/medium/difficult classes, or on a subset of targets matched for template availability and interface complexity. If the 5–15 percent improvements shrink to zero or reverse within matched difficulty strata, the claimed progress is explained by target selection rather than method improvement.","tokens_in":15416,"feed_emoji":"🧬","tokens_out":5713,"duration_ms":481142,"temperature":0.7,"pith_summary":"This paper assesses the second dedicated protein assembly category in CASP13, comparing predictions from 45 groups on 42 oligomeric targets against the CASP12 assembly round. It claims consistent, albeit modest, progress: 5–15 percent improvement across all four evaluation scores, with 9 of 42 targets solved on all metrics. It also finds that ignoring the oligomeric state harms even tertiary-structure predictions, that homomeric interfacial contacts are already predictable but were not used for assembly modelling, and that human-assisted homology modelling dominates the field. For a general reader, the takeaway is that quaternary structure is intrinsic to protein modelling and must be built into prediction pipelines from the start.","feed_headline":"Protein assembly prediction improves 5–15% in CASP13","feed_subtitle":"Nine of 42 oligomeric targets solved; homomeric interface contacts are predictable but go unused.","key_machinery":"The machinery is the scoring and ranking protocol. Four scores—ICS (F1) and IPS (Jaccard) for interfaces, plus lDDT_O and GDT_O for the whole assembly—require mapping chains between model and target, done with the 13-score algorithm (13-align for the largest target). Per-target z-scores with outlier removal and leave-one-out ranking guard against inflated scores on hard targets. For the CASP12-vs-CASP13 comparison, scores are matched by percentiles under the assumption of similar target difficulty, with a naive assembly method as the per-target baseline.","core_discovery":"The central discovery is that protein assembly prediction improved measurably but not dramatically between CASP12 and CASP13. Using four scores—interface contact similarity (F1), interface patch similarity (Jaccard), and oligomeric versions of lDDT and GDT—the paper finds a consistent 5–15 percent improvement across score percentiles, and identifies 9 of 42 assembly targets as solved by all four scores above 0.5. The gains track human-assisted homology modelling; fully automated servers rank near the naive baseline. A second finding is that homomeric interfacial contacts are already predicted well by some contact prediction groups, yet these predictions are currently counted as false positives and were not fed into assembly modelling. The paper also shows that treating chains as independent folding units degrades even tertiary predictions for targets with intertwined interfaces.","pith_inferences":["If homomeric interface contacts were added to contact-prediction evaluation, some group rankings would change (the paper notes H0968S2 as an example); a future CASP round that scores them explicitly would likely reward groups that already predict them.","The dominance of human-assisted homology modelling suggests the modest gains may saturate as structural templates become the limiting factor; a direct test is whether deep-learning pipelines trained on oligomeric targets can beat the homology baseline on difficult targets with no assembly templates.","The data-assisted comparison is currently confounded because the best non-assisted groups did not join the assisted category; a cleaner test would compare the same groups' assisted and non-assisted predictions on identical targets."],"forward_implications":["Predictors that ignore the oligomeric state will fail on intertwined interfaces: targets whose evaluation unit was the monomer received poor tertiary predictions despite good subunit templates.","Contact prediction groups already produce accurate homomeric interface contacts, so the next step is to fold multiple chains simultaneously from contact matrices rather than discarding interfacial contacts as false positives.","The 5–15 percent improvement is real but modest and largely attributable to human-assisted homology modelling; automated servers rank close to the naive baseline.","Data-assisted predictions using SAXS, crosslinking, or NMR showed no systematic improvement over regular predictions in CASP13, apart from one target with favorable crosslinks.","Only two servers participate in the fully automated multimeric CAMEO experiment, indicating that automation in assembly modelling lags behind tertiary modelling."],"supporting_citations":[{"why":"Supplies the CASP12 assembly assessment that defined the four-score evaluation and the comparison baseline.","marker":"[21]"},{"why":"Source of the CASP12 target statistics (30 of 71 targets oligomeric) used to frame the increase in CASP13.","marker":"[15]"},{"why":"Provided the naive assembly method used as a per-target baseline for score interpretation.","marker":"[27]"},{"why":"Defines the lDDT score used for local model quality in the four-score pool.","marker":"[22]"},{"why":"Defines the GDT score adapted here to oligomeric models for global fold similarity.","marker":"[23]"},{"why":"Supplies the 13-score chain-mapping algorithm needed to score multi-chain models.","marker":"[24]"},{"why":"Provides the 13-align variant used for chain mapping on the largest target H1021.","marker":"[25]"},{"why":"Documents z-score inflation on difficult targets, motivating the leave-one-out ranking used for group scores.","marker":"[26]"},{"why":"Shows contact prediction matured in CASP12, supporting the finding that homomeric interface contacts are now predictable.","marker":"[14]"}],"fun_headline_variants":["CASP13: Assembly prediction up 5–15%, but servers lag","Protein assembly gains modest in CASP13, humans beat servers","CASP13 assembly: 9 of 42 solved, contacts go unused","Homology with human touch wins CASP13 assembly round","CASP13: Assembly quality rises, automation stalls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The progress claim stands on the assumption that CASP12 and CASP13 assembly targets have roughly the same difficulty distribution; if the 2019 targets are easier, the 5–15 percent gains are an artifact of target selection.","fun_headline_variants_meta":{"raw":{"variants":["CASP13: Assembly prediction up 5–15%, but servers lag","Protein assembly gains modest in CASP13, humans beat servers","CASP13 assembly: 9 of 42 solved, contacts go unused","Homology with human touch wins CASP13 assembly round","CASP13: Assembly quality rises, automation stalls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1188,"prompt_tokens":843,"completion_tokens":345,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":256}},"tokens_in":459,"tokens_out":345,"duration_ms":160703,"temperature":1.0,"reasoning_tokens":256,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:59:52.205880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the percentile-matched comparison with targets stratified by the paper's own easy/medium/difficult classes, or on a subset of targets matched for template availability and interface complexity. If the 5–15 percent improvements shrink to zero or reverse within matched difficulty strata, the claimed progress is explained by target selection rather than method improvement.","supporting_citations":[{"cited_title":"Assessment of protein assembly prediction in CASP12","cited_arxiv_id":null,"evidence_quote":"Supplies the CASP12 assembly assessment that defined the four-score evaluation and the comparison baseline."},{"cited_title":"Critical assessment of methods of protein structure prediction (CASP)—Round XII","cited_arxiv_id":null,"evidence_quote":"Source of the CASP12 target statistics (30 of 71 targets oligomeric) used to frame the increase in CASP13."},{"cited_title":"The challenge of modeling protein assemblies: the CASP12-CAPRI experiment","cited_arxiv_id":null,"evidence_quote":"Provided the naive assembly method used as a per-target baseline for score interpretation."},{"cited_title":"lDDT: a local superposition-free score for comparing protein structures and models using distance difference tests","cited_arxiv_id":null,"evidence_quote":"Defines the lDDT score used for local model quality in the four-score pool."},{"cited_title":"Processing and analysis of CASP3 protein structure predictions","cited_arxiv_id":null,"evidence_quote":"Defines the GDT score adapted here to oligomeric models for global fold similarity."},{"cited_title":"Modeling protein quaternary structure of homo-and hetero-oligomers beyond binary interactions by homology","cited_arxiv_id":null,"evidence_quote":"Supplies the 13-score chain-mapping algorithm needed to score multi-chain models."},{"cited_title":"BioJava 5: A community driven open-source bioinformatics library","cited_arxiv_id":null,"evidence_quote":"Provides the 13-align variant used for chain mapping on the largest target H1021."},{"cited_title":"Evaluation of template-based models in CASP8 with standard measures","cited_arxiv_id":null,"evidence_quote":"Documents z-score inflation on difficult targets, motivating the leave-one-out ranking used for group scores."},{"cited_title":"Assessment of contact predictions in CASP12: Co-evolution and deep learning coming of age","cited_arxiv_id":null,"evidence_quote":"Shows contact prediction matured in CASP12, supporting the finding that homomeric interface contacts are now predictable."}],"review_version":1}