{"id":"6dd20c4a-bd31-4a08-b702-355c9f6a035e","arxiv_id":"2607.08203","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"Published SOTA gains in colonoscopy polyp segmentation are largely non-comparable because of omitted clinical metrics, incompatible splits, and missing significance tests, as shown by a 27-paper audit and uniform re-evaluation.","lead":"An audit of 27 polyp-segmentation papers finds that leaderboard progress is hard to trust: most omit boundary and recall metrics, mix incompatible train/test splits, and skip significance tests. Controlled re-runs of five models show that Dice hides large failures and that near-tied rankings flip with metric or split.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged sample-representativeness caveat; the dual audit-plus-re-evaluation design still supports the central claim.","rationale":"The central claim is dual: (1) structural evaluation failures are widespread in the audited literature, and (2) those failures are not merely cosmetic because a uniform re-evaluation exposes clinically relevant failures and non-robust rankings. The reader's weakest assumption correctly flags that the 27-paper cohort is filtered (still-image, fully-supervised, PraNet-template-adjacent) and therefore the exact omission fractions may not generalize to every polyp paper ever written. That is a fair caveat for treating the rates as definitive field standards, which is why CONDITIONAL is appropriate. However, it does not undercut the re-evaluation half of the claim: the five-model, three-protocol results stand on their own and already demonstrate that Dice conceals large boundary/recall failures, that the winner changes with the metric, and that near-tied rankings reverse across random splits. Threats to validity (self-trained weights, single-seed P1, HD95 convention differences) are explicitly quantified in the manuscript and do not reverse the qualitative findings. No stronger load-bearing flaw (e.g., train/test leakage, metric mis-implementation that collapses the gaps, or circular selection of only failing models) is evident. Therefore the reader's CONDITIONAL verdict and high confidence remain the right call; no adjustment is required.","tokens_in":15012,"tokens_out":597,"duration_ms":6209,"concrete_test":"Independently re-score the three released-weight models (PraNet, Polyp-PVT, SANet) on the P1 unseen sets with the official MetricsReloaded package for Dice/Recall/NSD and both medpy and MetricsReloaded HD95 conventions; if the ranking reshuffles or the HD95/Recall gaps shrink below the Dice gaps, the 'conceals failures' claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (representativeness of the 27-paper eligibility filter) is real but secondary. The paper's strongest claim is not solely the 25/27 and 26/27 rates; it is that those rates, together with the controlled re-evaluation of five models under three protocols (P1 fixed, P2 random, P3 OOD), show Dice-only fixed-split leaderboards are fragile. Tables 3–6 and Figure 4 supply independent empirical support: HD95 gaps far larger than Dice gaps, ranking reshuffles by metric and by seed, lesion-level reordering, and uniform absolute OOD degradation. Self-trained checkpoints and single-seed P1 runs are already disclosed and bounded by the P2 variance numbers. No internal inconsistency or hidden assumption that would overturn the qualitative conclusion was found.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper audits evaluation practices in colonoscopy polyp segmentation across 27 fully-supervised still-image papers (2015–2026) and argues that reported leaderboard progress is hard to verify. It documents three structural issues: omission of Hausdorff/surface-distance metrics in 25/27 papers, co-existence of at least five incompatible train/test split protocols on Kvasir-SEG and CVC-ClinicDB, and absence of statistical significance testing in 26/27 papers (including four works after Metrics Reloaded). To show these are not merely cosmetic, the authors re-evaluate five representative models under three controlled protocols (PraNet fixed split, random 80/20 splits, and PolypGen OOD) with a single uniform scorer, reporting omitted boundary, recall, and lesion-level metrics. They find that Dice conceals large HD95 and recall failures, that model ranking depends on the metric, and that near-tied rankings reverse across random splits. They propose a five-point Polyp Segmentation Reporting Checklist (PSRC).","tokens_in":15256,"tokens_out":1539,"duration_ms":24758,"significance":"If the audit rates and re-evaluation results hold, the paper is a useful, domain-specific corrective for a sub-community that has largely frozen its evaluation template since PraNet (2020). The dual design—literature meta-audit plus multi-protocol re-measurement under one scorer—is the right form of evidence, and the work gives concrete credit where it is due: MetricsReloaded-validated NSD, primary-source transcription checks (Table 2), FDR-corrected Wilcoxon tests, lesion-level matching under multiple criteria, and an explicit threats-to-validity section with P2 seed-variance bounds. The PSRC is lightweight and actionable. The contribution is primarily diagnostic and standard-setting rather than architectural; its value is in making published Dice comparisons more honest and clinically grounded.","major_comments":[{"comment":"Section 4 (Meta-dataset compilation and paper eligibility) and the headline 25/27 and 26/27 rates: the eligibility filter requires reporting on at least one of the five PraNet-template datasets and excludes video, weak/semi-supervised, SAM-style, and non-template-only works. That filter is reasonable for studying the PraNet-era leaderboard, but it selects precisely the literature most likely to copy the PraNet metric set. The manuscript should either (i) quantify how sensitive the omission rates are to relaxing criteria (a) and (c), or (ii) reframe the claim more narrowly as “within the PraNet-template still-image literature” rather than as a field-wide structural diagnosis of “the community.” Without that, the strongest numerical claims risk over-generalization from a path-dependent cohort.","section":"Section 4, Table 1"},{"comment":"Section 6.1 and Threats to Validity: PVT-CASCADE and G-CASCADE are self-trained because no polyp checkpoints are released. In-distribution reproduction for G-CASCADE is within 0.018–0.035 Dice of published numbers (Kvasir 0.893 vs 0.927; ClinicDB 0.929 vs 0.947). That gap is comparable to multi-year claimed SOTA margins on the same split (~3%). Absolute P1 rankings and HD95/recall for G-CASCADE on unseen sets (Table 3) therefore rest on a weaker fidelity assumption than for the three released-weight models. The paper already anchors metric-disagreement claims on released weights where possible; it should either multi-seed the self-trained runs, report the same P2-style variance for them, or demote G-CASCADE absolute scores more clearly to secondary evidence so that Table 3 is not read as a definitive five-model ranking.","section":"Section 6.1, Table 3, Threats to Validity"},{"comment":"Abstract / Introduction clinical framing of HD95: the paper repeatedly states that Hausdorff distance has “direct clinical relevance for detecting flat or small polyps.” HD95 measures boundary localization error on detected tissue; lesion-level sensitivity/recall (which the re-evaluation does report well in Table 5) is the metric that more directly addresses missed polyps. The clinical motivation for boundary metrics (resection margin, sizing) is sound and should be kept, but the wording should not equate HD95 with detection of flat/small lesions. Align the abstract claim with the Metrics Reloaded problem fingerprint and with the paper’s own lesion-level analysis.","section":"Abstract, Section 1, Figure 1"}],"minor_comments":[{"comment":"Table 1 includes PolyMamba-Net (Front., 2026) and other 2025–2026 entries; given the arXiv stamp (Jul 2026), briefly state the cutoff date and preprint inclusion rule so readers do not misread the timeline as post-dated.","section":"Table 1, Section 4"},{"comment":"NSD@3px is used throughout; state the rationale for τ = 3 px (pixel size / clinical tolerance) and whether results are stable under nearby tolerances, even if only in the supplement.","section":"Section 3, Evaluation metrics; Section 6"},{"comment":"Figure 4 is informative but dense; ensure that Dice/Rec/HD95 values under each panel remain legible in print and that the blue/orange border legend is repeated in the caption.","section":"Figure 4"},{"comment":"Section 5 notes that secondary survey tables contain transcription errors; consider releasing the primary-source extraction sheet with the open-source toolkit so the Table 2 discrepancies are independently checkable.","section":"Section 5, Table 2"},{"comment":"Minor wording: “Title Suppressed Due to Excessive Length” appears as a running header in the provided text; fix for camera-ready. Also standardize “CVC-ClinicDB/CVC-612” naming on first use.","section":"Headers / Section 1"},{"comment":"P2 uses three seeds and an 80/20 ratio; report the exact seed values and whether the 20% test set is stratified by dataset source (Kvasir vs ClinicDB), which affects reproducibility of Table 4.","section":"Section 6.1, Table 4"}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid audit-and-reproducibility contribution appropriate for a methods/reproducibility track or a medical imaging journal that values evaluation standards. The main risk is over-claiming “community-wide” rates from a PraNet-template-selected cohort; that is fixable with reframing and does not undermine the re-evaluation core. I would not require new architectures or a larger model zoo. Fit is good for venues that published Metrics Reloaded-style work or MICCAI/MIDL reproducibility tracks; less so for pure architecture venues."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is not another architecture paper. It is a careful audit of how 27 polyp-segmentation papers actually evaluate, paired with a uniform re-scoring of five models under fixed, random, and OOD protocols. The central claim holds: published Dice progress is hard to trust because boundary metrics and recall are almost never reported, incompatible splits sit in the same leaderboard columns, and significance testing is essentially absent.\n\nWhat is new is the combination. The 27-paper table (metric profile, split codes, external validation, stats) is transparent, the primary-source transcription check catches real secondary-table errors, and the re-evaluation is the load-bearing part. Under one scorer they show HD95 gaps far larger than Dice gaps, ranking reshuffles by metric and by random seed, lesion-level reordering that differs from pixel Dice, and a uniform absolute drop on PolypGen. NSD is checked against Metrics Reloaded; they report FDR-corrected Wilcoxon tests and multiple lesion-matching criteria. That is real work, not just complaint.\n\nSoft spots exist but are already mostly disclosed. Two of five models are self-trained (no public polyp weights); they bound the issue with in-distribution reproduction numbers and P2 seed variance. P1 is single-seed. The 27-paper cohort is filtered (still-image, fully supervised, PraNet-template interface), so the exact 25/27 and 26/27 rates are sample rates, not a census of every related paper. That does not overturn the qualitative conclusion that Dice-only fixed-split leaderboards are fragile. HD95 implementation differences across libraries are noted; absolute numbers should only be compared inside one scorer. The PSRC checklist is sensible and lightweight; the claimed toolkit is mentioned but not fully detailed in the text we have.\n\nThis is for people who care about medical-imaging evaluation practice, not for someone hunting a new backbone. Citation pattern is appropriate (Metrics Reloaded, PolypGen, the PraNet lineage). Math and data handling look solid enough for the claims being made. I would send it to referees; the field needs this kind of corrective more than another 0.01 Dice paper. Engage with it, and cite the re-evaluation tables if you work in the area.","headline":"Solid audit-plus-re-evaluation paper: Dice-only fixed-split leaderboards in polyp segmentation are fragile, and the controlled multi-protocol evidence makes that claim stick.","tokens_in":15860,"tokens_out":556,"would_cite":true,"duration_ms":6433,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Apparent progress on colonoscopy polyp segmentation leaderboards is hard to verify because of omitted boundary metrics, incompatible train/test splits, and claims made without significance tests.","keywords":["evaluation metrics","polyp segmentation","reproducibility","benchmark audit","clinical metrics","Hausdorff distance","train-test splits","statistical significance"],"falsifier":"Expand the audit to a substantially larger, independently sampled set of still-image polyp papers and re-run the three-protocol re-evaluation on a broader model set; if most papers already report Hausdorff and significance tests, or if rankings remain stable across metrics and random splits, the structural-failure claim fails.","tokens_in":15874,"feed_emoji":"🩺","tokens_out":958,"duration_ms":19751,"temperature":0.7,"pith_summary":"This paper argues that the steady state-of-the-art gains reported for deep models that segment polyps in colonoscopy images cannot be trusted as true progress. An audit of twenty-seven papers finds that almost all skip a clinically relevant boundary-error metric, that at least five incompatible data-split recipes are used on the same public datasets so published Dice numbers are not comparable, and that nearly every paper asserts superiority without any statistical test. A controlled re-scoring of five representative models under one uniform evaluator then shows that the usual region-overlap score conceals large boundary and missed-lesion failures, that the winning model changes when a different metric is used, and that near-tied rankings reverse across random splits. The authors offer a short five-point reporting checklist so future papers can make claims that are clinically meaningful and actually comparable.","feed_headline":"Polyp AI leaderboards rest on non-comparable scores","feed_subtitle":"Omitted boundary metrics, mixed splits, and no significance tests hide ranking flips and clinical failures","key_machinery":"The four-dimension audit (metric profile, data partitioning, generalization scope, statistical rigor) of the twenty-seven-paper meta-dataset, followed by a uniform re-evaluation of five models under three protocols (fixed historical split, random splits, out-of-distribution multi-center data) that surfaces the omitted boundary, recall, and lesion-level scores and tests ranking stability.","core_discovery":"A systematic audit of twenty-seven fully supervised polyp-segmentation papers from 2015 to 2026 documents three structural evaluation failures: twenty-five omit Hausdorff or any surface-distance metric, at least five incompatible train/test split protocols coexist on the same two public datasets, and twenty-six make performance claims without statistical significance testing. Re-evaluating five models under three controlled protocols with a single scorer confirms that these failures are not cosmetic: Dice conceals large boundary and recall errors, the best model depends on which metric is chosen, and near-tied rankings flip across random splits.","pith_inferences":["The same freeze-around-an-early-template pattern likely exists in other medical segmentation sub-fields that copy a single popular six-metric table for years.","Journals and conferences could treat the five-point checklist as a mandatory reporting item, converting an optional community norm into an enforceable standard.","If secondary survey tables already contain transcription errors larger than claimed gains, automated primary-source verification should become part of any future leaderboard.","A public multi-center hold-out set that is never used for architecture search would make out-of-distribution claims falsifiable rather than optional."],"forward_implications":["Published Dice numbers on the two main public datasets cannot be placed in the same leaderboard column unless the exact split protocol is identical and declared.","Boundary accuracy and lesion-level recall must be reported alongside Dice for any claim about clinical utility or edge-aware architectures.","Near-tied model rankings on a single fixed split are not robust evidence of superiority.","Future papers that follow the five-point checklist will produce scores that are both clinically interpretable and statistically comparable.","In-distribution leaderboard gains can coexist with large absolute drops on multi-center out-of-distribution data that remain invisible under current reporting."],"fun_headline_variants":["Polyp leaderboards hide non-comparable Dice and omitted boundary metrics","Audit of 27 papers finds mixed splits make polyp scores non-comparable","Colonoscopy AI rankings flip when metrics and splits are standardized","Most polyp papers skip Hausdorff, mix protocols, and skip significance tests","Flawed polyp benchmarks conceal boundary failures and ranking reversals"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim that the twenty-seven selected papers fairly represent community practice, so the high omission rates are structural rather than an artifact of how the cohort was filtered.","fun_headline_variants_meta":{"raw":{"variants":["Polyp leaderboards hide non-comparable Dice and omitted boundary metrics","Audit of 27 papers finds mixed splits make polyp scores non-comparable","Colonoscopy AI rankings flip when metrics and splits are standardized","Most polyp papers skip Hausdorff, mix protocols, and skip significance tests","Flawed polyp benchmarks conceal boundary failures and ranking reversals"]},"model":"grok-4.5","effort":"low","cost_usd":0.00415,"raw_usage":{"total_tokens":1373,"prompt_tokens":926,"num_sources_used":0,"completion_tokens":91,"cost_in_usd_ticks":41500000,"prompt_tokens_details":{"text_tokens":926,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":356,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":926,"tokens_out":91,"duration_ms":4034,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T11:22:35.169406+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Expand the audit to a substantially larger, independently sampled set of still-image polyp papers and re-run the three-protocol re-evaluation on a broader model set; if most papers already report Hausdorff and significance tests, or if rankings remain stable across metrics and random splits, the structural-failure claim fails.","supporting_citations":[],"review_version":1}