{"id":"d894294e-5ccc-4bfa-a80d-22eec7a4f2b2","arxiv_id":"2606.03178","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":2.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A methodological review of post-2018 evidence-aware protein complex detection methods concludes that transparent graph-based approaches offer the best plausibility-reproducibility tradeoff while calling for harmonized benchmarks and overlap-aware metrics.","lead":"This paper reviews methods for finding protein complexes in noisy interaction networks by combining topology with other biological evidence like gene annotations and expression data. It concludes that simple evidence-aware graph approaches currently balance plausibility and reliability best, but better standardized testing is the real next step.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of post-2018 method selection is the load-bearing assumption for claiming evaluation protocols are now the primary bottleneck","rationale":"The reader's weakest_assumption directly identifies the same point. With full text now available the concern remains testable via the review's own methodology rather than abstract-level inference; no other internal inconsistency (e.g., in metric definitions or GO circularity handling) appears more load-bearing for the headline synthesis.","tokens_in":1689,"tokens_out":298,"duration_ms":12524,"concrete_test":"Extract the explicit search strategy, inclusion/exclusion criteria, and full list of reviewed methods from the paper's methods or supplementary section; independently run a PubMed/arXiv search for 'protein complex detection' 2019-2024 limited to deep learning or hypergraph papers and check whether at least three high-citation examples with strong benchmark sections were omitted.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the reviewed post-2018 methods plus selected baselines are sufficiently representative to generalize that 'the central bottleneck is no longer only the lack of algorithms.' If the selection omitted recent deep or hypergraph methods that already incorporate overlap-aware or uncertainty-aware evaluation, or if inclusion criteria favor transparent graph methods, the synthesis that evaluation protocols are the dominant remaining issue does not follow. The paper's emphasis on evidence-aware graph methods as offering the strongest tradeoff rests on this sampling being unbiased.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"This manuscript is a focused methodological review of evidence-aware protein complex detection from PPI networks. It surveys methods that integrate network topology with GO annotations, expression profiles, subcellular localization, sequence/domain evidence, temporal data, and representation learning, with particular attention to post-2018 approaches and selected earlier baselines. The central synthesis states that transparent evidence-aware graph methods currently provide the strongest tradeoff between biological plausibility and reproducibility, whereas deep, hypergraph, and dynamic heterogeneous models increase realism at the cost of requiring tighter benchmark controls. The paper concludes that the primary remaining bottleneck is the absence of harmonized, overlap-aware, and reproducible evaluation protocols, and therefore advocates unified benchmark versions, explicit GO-circularity controls, overlap-aware metrics, uncertainty estimates, and executable software packages.","tokens_in":1787,"tokens_out":420,"duration_ms":18862,"significance":"If the synthesis holds, the review would be useful in redirecting community effort from isolated algorithmic novelty toward standardized, reproducible evaluation practices. The explicit recommendations for overlap-aware metrics and GO-circularity controls, together with the call for executable packages, constitute concrete, actionable guidance that could improve comparability across studies. The paper also usefully distinguishes the strengths of simpler evidence-aware graph methods from the added complexity of newer architectures.","major_comments":[{"comment":"Abstract: The claim that 'the central bottleneck is no longer only the lack of algorithms, but the lack of harmonized, overlap-aware, and reproducible evaluation protocols' is load-bearing for the central synthesis. This generalization requires that the reviewed post-2018 methods plus selected baselines are representative of the broader literature; without an explicit statement of literature search strategy, inclusion/exclusion criteria, or a systematic sampling frame (none provided in the abstract or visible in the synthesis), it remains possible that recent deep or hypergraph methods already incorporating overlap-aware or uncertainty-aware evaluation were omitted, which would undermine the conclusion that evaluation protocols are now the dominant issue.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. The major comment concerns the need for explicit documentation of literature selection to support the central claim in the abstract. We address this below and will revise the manuscript accordingly.","responses":[{"response":"We acknowledge the referee's point. The manuscript is presented as a focused methodological review of evidence-aware methods (explicitly combining PPI topology with GO, expression, localization, sequence/domain, temporal, or representation-learning evidence), not a systematic review. Selection was guided by coverage of post-2018 approaches meeting these criteria plus key baselines, drawn from recent surveys and field knowledge. To strengthen the paper, we will add an explicit 'Scope and Selection Criteria' paragraph (in the Introduction or a new subsection) describing the search strategy (PubMed, arXiv, Google Scholar; keywords combining 'protein complex detection', 'PPI network', 'evidence integration' or specific evidence types; post-2018 filter; inclusion of methods reporting benchmark performance). This will clarify the scope without altering the synthesis for the reviewed class of methods. We maintain that the central claim holds within this focused scope but agree explicit documentation is warranted.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that 'the central bottleneck is no longer only the lack of algorithms, but the lack of harmonized, overlap-aware, and reproducible evaluation protocols' is load-bearing for the central synthesis. This generalization requires that the reviewed post-2018 methods plus selected baselines are representative of the broader literature; without an explicit statement of literature search strategy, inclusion/exclusion criteria, or a systematic sampling frame (none provided in the abstract or visible in the synthesis), it remains possible that recent deep or hypergraph methods already incorporating overlap-aware or uncertainty-aware evaluation were omitted, which would undermine the conclusion that evaluation protocols are now the dominant issue."}],"tokens_in":1379,"tokens_out":399,"duration_ms":16169,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper is a methodological review of protein complex detection from PPI networks. It focuses on evidence-aware methods that layer in GO annotations, expression, localization, and other signals on top of topology, with special attention to post-2018 work. The central claim is that straightforward evidence-aware graph approaches currently give the best mix of biological plausibility and reproducibility, while deeper, hypergraph, or dynamic models add realism at the cost of needing tighter benchmark controls. The authors conclude that the field has largely moved past a shortage of algorithms and now needs harmonized, overlap-aware, and reproducible evaluation practices, along with explicit circularity controls and released code.\n\nIt does a clear job summarizing the practical difficulties with noisy and incomplete interactomes and spelling out concrete recommendations: unified benchmark versions, overlap-aware metrics, uncertainty estimates, and executable packages. Those suggestions are useful for anyone running these experiments.\n\nThe soft spot is the representativeness of the post-2018 methods chosen. The argument that evaluation protocols are now the dominant bottleneck only follows if the selected papers give a balanced picture of recent work. If the sample leans toward graph methods or leaves out recent deep or hypergraph approaches that already incorporate better evaluation, the synthesis weakens. The abstract does not detail inclusion criteria, so that needs checking in the full text.\n\nThis is for computational biologists who work on PPI network analysis and want a current snapshot plus practical advice on benchmarks. A reader in that area would get value from the recommendations. It deserves peer review because a careful synthesis with actionable suggestions on evaluation standards can help standardize work in the field.","headline":"A review of evidence-aware protein complex detection that flags evaluation protocols as the main remaining bottleneck, but whose claim rests on the reviewed methods being a fair sample.","tokens_in":2242,"tokens_out":396,"would_cite":false,"duration_ms":17520,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Evidence-aware graph methods best balance plausibility and reproducibility in protein complex detection, but evaluation protocols are now the main limit.","keywords":["protein complex detection","PPI networks","evidence-aware methods","reproducibility","benchmarks","evaluation protocols","Gene Ontology","graph methods"],"falsifier":"A broader survey that includes many post-2018 methods omitted here and finds that non-evidence-aware or highly complex models achieve superior reproducible performance would undermine the tradeoff claim; implementing the recommended unified benchmarks and seeing no shift in which methods rank highest would falsify the bottleneck claim.","tokens_in":2601,"feed_emoji":"🧬","tokens_out":708,"duration_ms":19052,"temperature":0.7,"pith_summary":"This review surveys methods that detect protein complexes by combining protein-protein interaction networks with additional evidence such as Gene Ontology terms, expression data, and localization. It concludes that transparent graph-based approaches that integrate this evidence deliver the strongest practical tradeoff between biological realism and the ability to reproduce results across studies. More elaborate deep learning, hypergraph, and dynamic models can capture richer biology but depend on tighter control of benchmarks to be reliable. The paper identifies inconsistent testing procedures, especially around overlapping complexes and circular use of annotations, as the current central obstacle rather than a shortage of new algorithms. It calls for shared benchmark sets, overlap-aware metrics, and full software releases to move the field forward.","feed_headline":"Evaluation protocols now limit protein complex detection more than algorithms","feed_subtitle":"Review finds transparent evidence-aware graph methods give best plausibility-reproducibility balance but inconsistent testing blocks progres","key_machinery":"Evidence-aware approaches that combine PPI topology with Gene Ontology annotations, expression profiles, subcellular localization, and other supporting data sources.","core_discovery":"Transparent evidence-aware graph methods currently offer the strongest tradeoff between biological plausibility and reproducibility, while deep, hypergraph, and dynamic heterogeneous models expand biological realism but require stronger benchmark control. The central bottleneck is no longer only the lack of algorithms, but the lack of harmonized, overlap-aware, and reproducible evaluation protocols.","pith_inferences":["Harmonized protocols developed here could be adapted to improve reproducibility in related tasks such as protein function prediction from networks.","If evaluation standards tighten, researchers might test whether evidence-aware methods maintain their edge when applied to context-specific or tissue-specific interaction maps.","The emphasis on transparent methods suggests that future work could prioritize interpretable models over black-box ones when integrating new data types like single-cell expression.","Adopting the recommended controls might narrow the gap between computational predictions and what can be validated in targeted experiments."],"forward_implications":["Unified benchmark versions would enable direct comparison of methods without hidden differences in data processing.","Explicit controls for circular use of Gene Ontology information would reduce inflated performance scores.","Overlap-aware metrics would produce rankings that better match the biological reality of shared subunits in complexes.","Routine reporting of uncertainty estimates would make performance claims more trustworthy across independent runs.","Releasing executable software packages would allow other groups to verify and extend reported results."],"fun_headline_variants":["Benchmarks bottleneck protein complex detection","Evidence-aware methods balance complex plausibility and reproducibility","Deep complex models require better benchmark control","Lack of harmonized protocols hinders complex detection progress"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The post-2018 methods and selected historical baselines reviewed are representative enough of the literature to establish that evaluation protocols rather than new algorithms are the primary limiting factor.","fun_headline_variants_meta":{"raw":{"variants":["Benchmarks bottleneck protein complex detection","Evidence-aware methods balance complex plausibility and reproducibility","Deep complex models require better benchmark control","Lack of harmonized protocols hinders complex detection progress"]},"model":"grok-4.3","cost_usd":0.0105,"raw_usage":{"total_tokens":4609,"prompt_tokens":603,"num_sources_used":0,"completion_tokens":52,"cost_in_usd_ticks":104999500,"prompt_tokens_details":{"text_tokens":603,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3954,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":603,"tokens_out":52,"duration_ms":28006,"temperature":1.0,"reasoning_tokens":3954,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T08:12:41.424626+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A broader survey that includes many post-2018 methods omitted here and finds that non-evidence-aware or highly complex models achieve superior reproducible performance would undermine the tradeoff claim; implementing the recommended unified benchmarks and seeing no shift in which methods rank highest would falsify the bottleneck claim.","supporting_citations":[],"review_version":1}