{"id":"df17d8ed-cc91-4586-9249-52a2b2f4e03c","arxiv_id":"2608.13223","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On BraTS-GoAT, a 3-seed deep ensemble's inter-member disagreement rises steeply under synthetic image corruption while single-model confidence stays flat, making disagreement the more shift-sensitive uncertainty signal.","lead":"This paper tests whether a deep learning segmentation model's uncertainty signals notice when brain MRI images are degraded, using synthetic corruptions as a stand-in for real-world scanner changes. A three-model ensemble's disagreement rises sharply with corruption while a single model's confidence stays flat, suggesting ensembles offer a better early warning for clinical deployment.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Relative-rise metric undermines the central sensitivity claim: disagreement has a small baseline, so percentage rises do not establish better case-level shift detection.","rationale":"The reader's weakest assumption is that synthetic corruptions reproduce real acquisition shift. That is an external-validity concern. The concern I identify is more load-bearing because it attacks the comparison metric itself: if relative rise is not a valid measure of case-level sensitivity, the headline claim fails even under the paper's own controlled study. The paper's Section 4 defense of relative rise makes this a deliberate analytical choice, not an oversight, but it does not resolve the scale-dependence problem. Since the goal is a case-level indicator, the appropriate test is whether the signal separates shifted from clean cases, e.g., via AUROC or standardized effect size, not merely whether its mean rises by a larger percentage. The concrete test requires no retraining and can be run from existing predictions, so it is a quick way to settle the concern. The reader already issued a conditional verdict asking for repeated corruptions and a second split; my concern adds a more direct validation of the sensitivity metric but does not move the verdict, so no adjustment is needed.","tokens_in":8982,"tokens_out":7125,"duration_ms":80213,"concrete_test":"Recompute Section 3.4 using per-case signal values already available for the 271 held-out cases: for clean and each severe corruption, compute (i) single-model mean uncertainty, (ii) 3-seed mean disagreement, (iii) 3-seed mean uncertainty. Then (a) compute the paired AUROC for classifying clean versus each severe corruption using each signal; (b) compute standardized mean differences (Cohen's d) of per-case changes with bootstrap 95% CIs; (c) repeat over at least five independent corruption realizations and, if compute permits, over the other folds. If disagreement's AUROC or standardized separation does not exceed single-model confidence's, the central claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 3.4, that ensemble disagreement is a more sensitive case-level indicator of acquisition shift than single-model confidence, is operationalized entirely through percentage rise from the clean condition (Fig. 4). This metric is scale-dependent. Section 2.3 notes that disagreement D, being a variance of probabilities, is 'inherently small', while single-model uncertainty 1-c has a much larger baseline. A small baseline can produce a large relative change from a modest absolute increase, and a signal with a larger absolute increase can show a smaller relative rise. The paper reports no absolute signal values, no per-case distributions, no confidence intervals, and no ROC-style analysis separating clean from corrupted cases. The limitation paragraph in Section 4 explicitly leans on the relative rise, but relative rise is not a measure of case-level discriminative utility: a shift indicator is useful only if its value can separate shifted from unshifted cases. Thus, even within the synthetic proxy, the reported sensitivity order does not establish the headline conclusion.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-14T15:57:21.611639+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}