{"id":"aa82554b-d12b-48ff-a4ce-4c12cfb77761","arxiv_id":"2506.12006","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The 2022 and 2023 crossMoDA challenge results show that training on heterogeneous multi-institutional data reduces segmentation outliers on homogeneous test sets, while cochlea Dice declined in 2023.","lead":"This paper reports the results of the 2022 and 2023 crossMoDA challenges, which test algorithms that segment vestibular schwannoma tumors and cochleas on T2 MRI using models trained on contrast-enhanced T1 images. It finds that models trained on more diverse multi-institutional data produce fewer outliers on older homogeneous test sets, while cochlea accuracy dipped in 2023.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The data-heterogeneity claim in §6.2 is confounded by winner-model differences; a fixed-architecture training-set ablation is needed before the causal interpretation can stand.","rationale":"The reader's weakest_assumption correctly identifies the uncontrolled winner-model comparison as the load-bearing premise of the paper's central causal claim. My review agrees: the conclusion that 'increased data heterogeneity can enhance segmentation performance even on homogeneous data' is presented as a demonstrated finding, but the evidence in Fig. 12 and Table 8 compares three different methods, not three training-data configurations under a fixed method. This is not an internal inconsistency in the reported challenge results, which appear carefully collected and analyzed, but it is a validity threat to the interpretive claim in the abstract, §6.2, and the conclusion. The paper remains valuable as a challenge report with public benchmarks, stable ranking analyses, and detailed method summaries; the issue is about the causal inference drawn from those benchmarks. A controlled ablation as proposed would settle the question, and in its absence the correct framing is a hypothesis or observation, not a demonstrated causal effect. The reader's conditional verdict is therefore appropriate, and I would not move it to accept or reject on the basis of this concern. No ad hominem or methodological dishonesty is implied; the concern is purely about the strength of the evidence for the stated conclusion.","tokens_in":32558,"tokens_out":2696,"duration_ms":33771,"concrete_test":"Run a controlled experiment that fixes the 2023 winner's complete pipeline (3D QS-Attn synthesis, nnUNet segmentation, self-training, 11-model ensemble) and trains it on three training-data configurations: (i) London SC-GK only, (ii) London SC-GK + Tilburg SC-GK, and (iii) all three datasets including UK MC-RC. Evaluate each configuration on the 2021 London test set and the 2022 London + Tilburg test set, reporting median DSC and a pre-specified outlier definition (e.g., cases below median minus 1.5 IQR). If configuration (iii) alone reduces outliers, the data-heterogeneity claim is supported; if configurations (i) or (ii) already match it, the observed gains are due to architecture or ensembling rather than data diversity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, stated in the abstract and in §6.2, is that the 2023 winner's reduced outlier count on 2021 and 2022 test data demonstrates that increased data heterogeneity improves segmentation performance even on homogeneous data. This claim is used to explain Fig. 12 and Table 8, but the comparison is not a controlled experiment: the 2021, 2022, and 2023 winners differ simultaneously in image-translation architecture (2D CycleGAN, NiceGAN, 3D QS-Attn), segmentation backbone, augmentation strategy, self-training details, and ensemble size (5 models vs. 11 models). For example, on the London SC-GK test set, the 2023 winner differs from the 2022 winner not only in training-data composition but also in using a 3D synthesis network and an 11-model ensemble, so the observed gain cannot be uniquely attributed to data heterogeneity. The same confound applies to the comparison on the combined London + Tilburg test set, where the 2023 winner was trained on additional UK MC-RC data while also using a different method. Furthermore, the paper does not define what counts as an outlier, and the p-values reported in §6.2 are not tied to a stated statistical test, so the quantitative evidence for 'fewer outliers' and 'significant improvement' is not fully specified. The observed trends may well be directionally correct, but the causal conclusion as written is not justified by the presented comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the crossMoDA 2022 and 2023 challenge editions and retrospectively analyzes the 2021-2023 series. The challenge task is unsupervised cross-modality segmentation of vestibular schwannoma and cochlea from ceT1 to T2 MRI. The paper describes the datasets (London SC-GK, Tilburg SC-GK, UK MC-RC), the segmentation and Koos classification tasks, the metrics and ranking scheme, and the methods used by participating teams. It then presents the 2022 and 2023 results, a bootstrap-based ranking stability analysis, and a cross-edition comparison of the winning models. The central claim is that the 2023 winner, trained on the most diverse multi-institutional dataset, reduced the number of outliers on the 2021 and 2022 test sets, demonstrating that increased data heterogeneity improves segmentation even on homogeneous data; a secondary claim is that cochlea Dice declined in 2023 due to added sub-annotation complexity.","tokens_in":32926,"tokens_out":4308,"duration_ms":50494,"significance":"If the central claims held, this would be a valuable benchmark paper: it delivers openly available multi-institutional datasets, a reproducible evaluation pipeline (rank-then-aggregate with bootstrapped stability analysis), detailed method summaries from all participating teams, and a longitudinal view of how unsupervised domain-adaptation techniques evolved over three years. The ranking stability analysis in §5.6 is carefully executed and the data-sharing and evaluation-code links are concrete assets. However, the paper's headline causal inference—that data heterogeneity is the driver of the 2023 winner's improved robustness—is not identified by the presented comparison, because the winning models differ across multiple design dimensions simultaneously. The paper's descriptive and benchmarking content is solid, but the causal framing needs either a controlled analysis or a substantial weakening.","major_comments":[{"comment":"The central claim that increased data heterogeneity reduces outliers is confounded by winner-model differences. The 2021, 2022, and 2023 winners differ simultaneously in image-translation architecture (2D CycleGAN/PAST, NiceGAN, 3D QS-Attn), segmentation backbone, augmentation strategy, self-training details, and ensemble size (5-model vs. 11-model ensembles). For example, the 2023 winner Vandy365 uses an 11-model ensemble and site-specific style synthesis, whereas the 2022 winner ne2e uses NiceGAN with a 5-model nnUNet ensemble. The observed improvement on the London SC-GK test set could arise from any of these differences, so this is not a controlled test of data heterogeneity. A fixed-architecture ablation varying only the training-data composition—e.g., the 2023 winner's pipeline trained on 2021-only, 2022-only, and 2023 data—would be needed to support the causal reading; absent that, the abstract's 'demonstrating' should be softened to 'consistent with' or 'suggesting.'","section":"§6.2, Fig. 12, Table 8"},{"comment":"The statistical evidence behind the p-values is unspecified. The text reports p<0.01 and p<0.005 for several comparisons in Fig. 12, but no test name, test statistic, pairing scheme, or multiple-comparison correction is given. Additionally, 'outliers' and 'outlier spread' are not defined in this section; Section 5.5 refers informally to box-plot outliers, but the reader cannot verify the claim that 'the number of outliers decreases with an expanding dataset' without a concrete outlier rule (e.g., Tukey fences) and outlier counts. Please specify the exact test used (e.g., paired Wilcoxon signed-rank test on per-case Dice values), state whether the test is paired across winners on the same test cases, report effect sizes, and correct for the multiple pairwise comparisons implied by Fig. 12.","section":"§6.2, Fig. 12"},{"comment":"The attribution of the 2023 cochlea decline to sub-segmentation complexity is speculative. The abstract says the decline is 'likely due to the introduction of a subdivided tumour annotation' and §6.2 says it is 'likely due to the increased complexity introduced by multi-institutional training data and increased challenge posed by the new intra/extra-VS sub-segmentation requirements.' However, the 2023 winner differs from the 2022 winner in architecture, training data, and task objective (the 2023 ranking includes intra-meatal, extra-meatal, and cochlea, rather than combined VS and cochlea), so the decline cannot be uniquely attributed to sub-annotation complexity. At minimum, the paper should test or explicitly acknowledge alternative explanations, such as the different ranking objective or the different synthesis/ensemble design.","section":"§6.2 and Abstract"},{"comment":"The cross-test evaluation of older winners may be affected by task mismatch. The 2021 and 2022 winners were trained for two-class segmentation (VS and cochlea), while the 2023 winner was trained with intra-meatal and extra-meatal sub-labels. When the 2021/2022 winners are evaluated on the 2023 test set, it is not stated how the combined VS ground truth is derived from the intra/extra masks or whether the older models' predictions were post-processed to match the combined evaluation. If the maps are simply merged, the comparison may penalize older models for a task they were not designed for; please describe the label mapping and any preprocessing applied for these retrospective evaluations.","section":"§6.2, Fig. 12"}],"minor_comments":[{"comment":"The notation T_c for the set of test samples in class c conflicts with the use of T for target domain elsewhere in the paper; consider renaming to D_c or S_c.","section":"§3.2, Eq. (3)"},{"comment":"The text cites 'Erickson et al.' in the Koos grading protocol description without a year or a full reference entry; please add the complete citation.","section":"§2.4"},{"comment":"The claim that the top-ranked team maintained its position across all bootstrap samples is reported separately for 2022 and 2023, but the blob plot in Fig. 9 would be easier to verify if the percentage of bootstrap samples in which each team held its exact rank were stated numerically in the caption.","section":"§5.6"},{"comment":"The sentence describing the bootstrap sample size states that each sample retained approximately 63% of distinct cases; this is the expected fraction for sampling with replacement, but it would be clearer to state that the statement follows from the Poisson approximation.","section":"§3.3.1"}],"recommendation":"major_revision","confidential_remarks":"The benchmark and dataset releases are genuinely useful, and the descriptive parts of the paper are solid. The main risk is that the abstract and §6.2 overclaim a causal effect from an observational comparison. This is fixable by adding a controlled experiment or by reframing the conclusions as correlational; either route would make the paper publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a serious challenge wrap-up that publishes the 2022 and 2023 crossMoDA results together for the first time, and the ranking methodology is handled with care. The retrospective winner comparison in Fig. 12 is the most interesting part, but the headline claim — that heterogeneous training data reduces outliers on homogeneous test sets — is not actually established by the data as presented.\n\nWhat's new and good: the Koos classification results, the intra-/extra-meatal segmentation results, and the cross-edition winner comparison are genuinely new in this consolidated form. The datasets are public, the evaluation code is on GitHub, and the bootstrapped ranking stability analysis is well done — median Kendall's tau of 1 across 1000 samples is reassuring. The method summaries are useful for anyone working on cross-modality domain adaptation, and the paper will be a convenient reference for these benchmarks.\n\nThe soft spot is exactly what the stress-test note flags. The central interpretation in Section 6.2 and the abstract — that the 2023 winner's reduced outlier count demonstrates the effect of data heterogeneity — is an observational comparison, not a controlled one. The 2021, 2022, and 2023 winners differ simultaneously in translation architecture (2D CycleGAN vs NiceGAN vs 3D QS-Attn), segmentation backbone, augmentation strategy, and ensemble size (11 models vs 5). The 2023 winner's gains on the 2021/2022 test sets could just as easily come from its larger ensemble or site-specific style synthesis. The paper never defines what counts as an outlier, and the p-values in Fig. 12 are not tied to any stated statistical test. These are real but fixable issues: either reframe the observation as descriptive or add a fixed-architecture training-set ablation.\n\nThe cochlea decline explanation is explicitly hedged as \"likely\" in the abstract, which is appropriate; the main text states it a bit more firmly than the evidence supports, but that is a minor overstatement.\n\nBottom line: for the challenge results and benchmark numbers, this paper is the reference. For a proof that data diversity drives robustness, it does not get there on its own. It deserves a serious referee — the benchmark data are valuable, the evaluation is reproducible, and the confound can be addressed with a revised framing and a clearer statistical description.","headline":"Solid challenge report with valuable public benchmark results, but the causal claim about data heterogeneity driving robustness is confounded by winner-model differences and needs a softer framing.","tokens_in":33523,"tokens_out":2170,"would_cite":true,"duration_ms":27169,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The crossMoDA challenge reports that the 2023 winning model, trained on the most heterogeneous multi-institutional dataset, reduced outliers on earlier single-site test sets while maintaining vestibular schwannoma segmentation accuracy…","keywords":["vestibular schwannoma segmentation","cochlea segmentation","cross-modality domain adaptation","unsupervised domain adaptation","MRI domain shift","data heterogeneity","challenge benchmark","intra-extra-meatal sub-segmentation"],"falsifier":"Train a single fixed model pipeline on each edition's training data—same architecture, same augmentation, same ensemble size—and evaluate on a common held-out set of homogeneous scans. If the model trained only on the 2021 data matches the outlier count and Dice of the model trained on the full 2023 data, the data-heterogeneity explanation fails. Similarly, removing the intra-/extra-meatal split from the 2023 training objective and observing unchanged cochlea Dice would falsify the claimed cause of the cochlea decline.","tokens_in":32368,"feed_emoji":"🧠","tokens_out":7285,"duration_ms":66632,"temperature":0.7,"pith_summary":"This paper reports the 2022 and 2023 rounds of the crossMoDA challenge, a benchmark in which algorithms learn to segment vestibular schwannoma tumours and cochleas on T2 MRI using only labels from contrast-enhanced T1 scans. Across three editions the organizers progressively enlarged the training data from one controlled single-site set to a multi-site routine-clinical set, and the paper retrospectively runs each year's winning model on every year's test sets. The central claim is that the 2023 winner, trained on the most heterogeneous data, produced fewer outliers and better or equal tumour Dice on the older, homogeneous test sets than the earlier winners, suggesting that data diversity itself improves generalization even to same-site scans. The paper also reports that cochlea segmentation accuracy fell in 2023, and attributes this to the added complexity of splitting the tumour into intra- and extra-meatal components. The clinical stake is that robust T2-only segmentation could replace contrast-enhanced surveillance imaging for vestibular schwannoma, which is safer and about ten times cheaper.","feed_headline":"More diverse MRI data cut outliers on every earlier test set","feed_subtitle":"The 2023 winner generalized best to old homogeneous scans, but cochlea accuracy slipped with added tumour sub-annotation.","key_machinery":"The central object is the challenge benchmark's retrospective winner-model comparison. Each edition's winning pipeline follows the same synthesis-then-segment recipe: an unpaired image-to-image translation network (2D CycleGAN in 2021, NiceGAN in 2022, 3D QS-Attn in 2023) converts labelled ceT1 scans into synthetic T2 scans, a 3D nnUNet is trained on the synthetic T2 scans with the ceT1 labels, and self-training on unlabelled real T2 scans refines the model. The analysis then fixes each edition's test set and evaluates all three winners on it, using Dice and average symmetric surface distance, and counts outliers in box-plot distributions. This cross-evaluation is what isolates the effect of changing training data composition.","core_discovery":"The retrospective analysis shows that the 2023 winning model, trained on the combined London, Tilburg, and UK multi-centre routine data, reduced the number of outliers on the 2021 and 2022 testing data compared with the 2021 and 2022 winners, even though those test sets were acquired under controlled, homogeneous protocols. On the single-institute London test set, the 2023 winner's VS Dice was significantly higher than the 2021 winner's and slightly higher than the 2022 winner's. By contrast, the 2021 winner's performance dropped sharply when applied to the 2022 and 2023 data, and the 2022 winner degraded on the 2023 data. The cochlea Dice of the 2023 winner declined relative to the 2022 winner on the shared 2022 test sets, which the authors attribute to the new intra-/extra-meatal tumour sub-annotation task rather than to poorer translation quality. The conclusion is that heterogeneous training data can improve segmentation robustness even on homogeneous target distributions, while added task complexity can cost accuracy on small structures.","pith_inferences":["A controlled version of this comparison would randomize model design: for instance, the same team's same architecture could be trained on each edition's dataset, isolating data diversity from translation and ensembling choices; the current paper does not do this.","The data-heterogeneity benefit is probably not linear: the 2023 dataset adds both more patients and much wider scanner and protocol variation, so separating sheer volume from diversity would clarify which factor drives the outlier reduction.","If the cochlea decline is truly caused by the sub-segmentation task, lightweight task-specific decoding or a multi-task loss that shields the cochlea head may recover the lost accuracy; this is a testable design hypothesis.","The same evaluation logic could apply to other cross-modality benchmarks, suggesting a general recipe: judge methods not only on their own test set but on previous editions' test sets to measure robustness transfer."],"forward_implications":["If the central claim holds, future cross-modality benchmarks should include heterogeneous routine-acquired data, because single-site controlled data understates the generalization gap.","Models trained only on homogeneous planning data should be expected to produce outlier segmentations when deployed at other sites, so clinical deployment without multi-site training data carries a robustness risk.","Adding clinically motivated sub-tasks such as intra-/extra-meatal splitting can lower accuracy on small nearby structures, so task complexity needs to be budgeted against core segmentation targets.","Leading challenge methods are plateauing on the current task, so a harder cross-modal benchmark may be needed to continue driving methodological progress.","T2-only automated segmentation, if robust across sites, could reduce the need for gadolinium contrast in vestibular schwannoma surveillance, cutting cost and safety risk."],"supporting_citations":[{"why":"Reports the crossMoDA 2021 edition, defines the 2021 winner and the original single-site benchmark against which later winners are compared.","marker":"Dorent et al., 2023"},{"why":"Provides the London SC-GK dataset and segmentation protocol used in all three editions, including the homogeneous 2021 test set.","marker":"Shapey et al., 2021"},{"why":"Introduces the UK multi-centre routine MRI dataset whose heterogeneity is the key new ingredient in the 2023 training data.","marker":"Kujawa et al., 2024"},{"why":"Supplies the nnUNet segmentation framework used by the winning teams, the common model whose generalization is measured across editions.","marker":"Isensee et al., 2021"},{"why":"Describes the 2023 winning method based on site-specific style synthesis with 3D QS-Attn and self-training, whose outlier reduction is the central observation.","marker":"Liu et al., 2023"},{"why":"Describes the 2022 winning PAST method based on NiceGAN translation and self-training, one of the models compared in the retrospective analysis.","marker":"Dong et al., 2021"}],"fun_headline_variants":["Diverse MRI data trims outliers on clean scans","Heterogeneous training boosts old test set robustness","2023 winner generalizes best to homogeneous data","More data variety improves segmentation, but cochlea slips","Varied scans cut outliers, yet cochlea accuracy falls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the design differences between the three winning models do not explain their performance gap, so the better generalization of the 2023 winner can be credited to its more diverse training data rather than to a better model.","fun_headline_variants_meta":{"raw":{"variants":["Diverse MRI data trims outliers on clean scans","Heterogeneous training boosts old test set robustness","2023 winner generalizes best to homogeneous data","More data variety improves segmentation, but cochlea slips","Varied scans cut outliers, yet cochlea accuracy falls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1442,"prompt_tokens":1109,"completion_tokens":333,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":725,"completion_tokens_details":{"reasoning_tokens":258}},"tokens_in":725,"tokens_out":333,"duration_ms":75612,"temperature":1.0,"reasoning_tokens":258,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:57:26.032757+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a single fixed model pipeline on each edition's training data—same architecture, same augmentation, same ensemble size—and evaluate on a common held-out set of homogeneous scans. If the model trained only on the 2021 data matches the outlier count and Dice of the model trained on the full 2023 data, the data-heterogeneity explanation fails. Similarly, removing the intra-/extra-meatal split from the 2023 training objective and observing unchanged cochlea Dice would falsify the claimed cause of the cochlea decline.","supporting_citations":[],"review_version":1}