{"id":"55b2bd79-16b0-440e-9bb8-cadc8787cdf2","arxiv_id":"2502.00563","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A mutual information loss computed on complex steerable pyramid subbands improves semantic segmentation for small instances and thin boundaries in tests on four datasets.","lead":"This paper proposes a new loss function for semantic segmentation that compares predictions with ground truth using mutual information in a complex wavelet domain. It reports better accuracy and topology metrics than existing loss functions on four segmentation datasets, but one results table contains a duplicated block that needs correction.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MASS ROAD Attention U-Net baselines in Table 1 are largely identical to the GlaS Attention U-Net column, so the claim of superiority across all datasets and both architectures lacks support for that block until it is rerun.","rationale":"I read the paper as an empirical proposal: CWMI is useful if it actually improves segmentation metrics across datasets and architectures at low overhead. The method is clearly described, the code is promised, and the intact SNEMI3D U-Net/Attention U-Net and VMUNet results plus the DRIVE results provide some support. However, the central empirical generalization in Section 4.2 depends on all four datasets, and the MASS ROAD Attention U-Net block is compromised: a large majority of its baseline rows exactly duplicate the GlaS Attention U-Net block. This is an objective data-integrity issue, not a matter of interpretation. Even though RMI, CWMI-Real, and CWMI rows differ slightly from the GlaS block, the shared baselines make the block unusable as independent evidence for MASS ROAD Attention U-Net comparisons. The abstract's broad significance claim also outruns the table's asterisks: for GlaS and MASS ROAD, pixel-wise metrics are not marked as significantly better, so the claim of significant improvements in both pixel-wise accuracy and topological metrics is not supported as stated. My proposed check is deliberately concrete: rerun the affected block and see whether CWMI still wins. If it does, the central claim can be restored; if not, the paper should be narrowed to the architectures and datasets where the evidence is clean. Because the reader already assigned CONDITIONAL and flagged the table duplication, my read does not move the verdict; it sharpens the reason for the condition.","tokens_in":17320,"tokens_out":7424,"duration_ms":62847,"concrete_test":"Run the released CWMI repository to retrain the MASS ROAD Attention U-Net experiment for all eleven baseline losses plus CWMI under the paper's stated repeated three-fold protocol, and compare the resulting Table 1 block against the published GlaS Attention U-Net block. If the published MASS ROAD baseline values are not reproduced and remain identical to GlaS, the block must be replaced with the rerun results; then determine whether CWMI still outperforms the rerun baselines on a majority of metrics for MASS ROAD Attention U-Net. Only after this rerun can the Section 4.2 claim about 'all datasets using both U-Net and Attention U-Net' be evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim is that CWMI 'achieves significant improvements in both pixel-wise accuracy and topological metrics compared to state-of-the-art methods' across diverse datasets. The empirical basis is Table 1, and Section 4.2 specifically asserts superiority 'across the majority of evaluation metrics for all datasets using both U-Net and Attention U-Net.' This claim is load-bearing and rests on a table with a data-integrity problem: in the MASS ROAD Attention U-Net block, the CE, BCE, Dice, Focal, Jaccard, Tversky, WCE, ABW, Skea-topo, clDice, and Sensitive rows exactly reproduce the corresponding GlaS Attention U-Net rows, e.g., CE .640±.209/.725±.195/.906±.107/.406±.327/7.914±6.637. Only RMI, CWMI-Real, and CWMI differ slightly. Independent MASS ROAD baseline results for Attention U-Net therefore cannot be confirmed from this table; if those baselines are wrong, the starred CWMI comparisons in that block are not valid evidence. Additionally, the abstract's 'significant... pixel-wise' claim is not backed by the table's own asterisks for GlaS and MASS ROAD, where mIoU/mDice show no significance markers. The theoretical Gaussian-MI concern is secondary: even if Eqs. 7-9 are a heuristic surrogate, the intact U-Net/VMUNet results could still support the method, but the duplicated table blocks cannot.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Complex Wavelet Mutual Information (CWMI) loss, a multi-scale loss for semantic segmentation. The method decomposes prediction and ground-truth images with a complex steerable pyramid, estimates a closed-form regional mutual information term in each subband (borrowing the Gaussian estimator from RMI), and adds it to cross-entropy. Experiments compare CWMI with 11 loss functions on SNEMI3D, GlaS, DRIVE, and MASS ROAD using U-Net and Attention U-Net, plus VMUNet on SNEMI3D, reporting pixel-wise (mIoU, mDice) and topological/clustering (VI, ARI, HD) metrics. The central claim is that CWMI achieves significant improvements in pixel-wise accuracy and topological metrics with minimal computational overhead.","tokens_in":17617,"tokens_out":3674,"duration_ms":38865,"significance":"If the empirical results are reproduced, CWMI would be a useful plug-in loss: it is architecture-agnostic, has a favorable O(HW log HW) complexity analysis, and the ablations (phase vs. magnitude, decomposition levels, orientations, and regularization weight) are informative. The public code and the breadth of comparisons are strengths. However, the current manuscript contains a serious data-integrity issue in Table 1 that invalidates one of the four dataset blocks for Attention U-Net, and the abstract's significance claim is not fully supported by the table's own asterisks. The core idea remains credible, but the evidence base needs repair before the claims can be accepted.","major_comments":[{"comment":"The entire MASS ROAD Attention U-Net block for the baseline losses is numerically identical to the GlaS Attention U-Net block: for example, CE is .640±.209/.725±.195/.906±.107/.406±.327/7.914±6.637 in both, and the same holds for BCE, Dice, Focal, Jaccard, Tversky, WCE, ABW, Skea-topo, clDice, and Sensitive rows. Two different datasets cannot yield identical means and standard deviations for eleven independent baseline training runs. This invalidates the MASS ROAD Attention U-Net comparisons, including the starred CWMI/CWMI-Real rows, and removes the support for the Section 4.2 statement that CWMI outperforms other losses 'for all datasets using both U-Net and Attention U-Net architectures.' The authors must rerun these baselines (or remove the block) and verify that the correct results are reported.","section":"Table 1 (MASS ROAD, Attention U-Net block)"},{"comment":"The abstract claims 'significant improvements in both pixel-wise accuracy and topological metrics,' but the paper's own significance markers do not support the pixel-wise part for GlaS and MASS ROAD. In Table 1, CWMI has no asterisks on mIoU or mDice for GlaS (both U-Net and Attention U-Net) or for MASS ROAD (both architectures); only VI and ARI are starred for those datasets, and MASS ROAD U-Net mIoU/mDice improvements are within one standard deviation of RMI. The significance claim in Section 4.2 is more carefully worded ('majority of evaluation metrics'), but the abstract overstates the result. The authors should either provide significance tests for mIoU/mDice on all datasets or revise the abstract to match the evidence.","section":"Abstract and Section 4.2"},{"comment":"The mutual information estimator is a closed-form Gaussian surrogate borrowed from RMI, extended to complex subbands via Hermitian covariance. Ground-truth segmentation masks are binary, and their wavelet subband coefficients are highly non-Gaussian, so the interpretation of the objective as 'mutual information' is heuristic unless the surrogate is validated. The manuscript offers no such validation (e.g., comparison with a nonparametric MI estimator, or a synthetic experiment showing the surrogate tracks structural agreement). This does not necessarily invalidate the empirical results, but the theoretical motivation should be explicitly qualified or supported.","section":"Section 3.2, Eqs. (7)-(9)"}],"minor_comments":[{"comment":"The text says 'three public segmentation datasets' but then enumerates four (SNEMI3D, GlaS, DRIVE, MASS ROAD); this should read 'four.'","section":"Section 4.1, Datasets"},{"comment":"The column header 'SENMI3D' is a typo and should be 'SNEMI3D.'","section":"Table 1 header"},{"comment":"The paragraph on decomposition level N and orientation K refers to 'Table 3' for the N/K and layer ablations, but these results appear in Table 5; the cross-references should be corrected.","section":"Section 4.3, Ablation references"},{"comment":"The text says the computational overhead comparison is 'as shown in Table 4,' but the timing results are in Table 6; the cross-reference is incorrect.","section":"Section 4.3, Computational complexity"},{"comment":"The phrase '11 state-of-art loss functions' should use the standard spelling 'state-of-the-art.'","section":"Section 1, Contributions"}],"recommendation":"major_revision","confidential_remarks":"The duplicated rows in the MASS ROAD Attention U-Net block are the most serious issue; they are most plausibly a table-assembly artifact, but I cannot rule out a data-processing error. Please ask the authors to supply raw per-split results and the exact training scripts for the MASS ROAD Attention U-Net runs, and to re-examine the significance markers. The abstract should also be brought in line with the asterisks in Table 1 before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the core idea is reasonable: take RMI's regional mutual information, swap in complex steerable pyramid subbands, and use a Hermitian covariance extension. Second, one block of Table 1 is copied verbatim from another block, so some headline claims currently rest on a table error.\n\nWhat's genuinely new is the combination itself, then the careful ablation work. The paper tests N, K, lambda, and the real/magnitude/phase components, and compares against L1/L2/SSIM in the wavelet domain. On SNEMI3D and DRIVE the intact results show consistent metric gains with modest overhead (about 0.23s per epoch over CE). The VMUNet experiment also helps generalization. The complexity analysis is straightforward and the code is public.\n\nNow the soft spots, in proportion. The duplicated block is serious: in Table 1, the MASS ROAD Attention U-Net column for CE, BCE, Dice, Focal, Jaccard, Tversky, WCE, ABW, Skea-topo, clDice, and Sensitive exactly matches the GlaS Attention U-Net column. That means the MASS ROAD Attention U-Net baselines are unverified, and the starred CWMI comparisons in that block are not valid evidence until the block is rerun. The abstract also overclaims: it promises significant pixel-wise improvements on GlaS and MASS ROAD, but the table shows no asterisks on mIoU/mDice for those datasets; only VI/ARI are significant. The Gaussian MI surrogate is a real theoretical gap, but the paper is honest that it borrows the estimator from RMI, and the intact empirical results could stand even if the surrogate is loose. Minor issues: 'SENMI3D' typo, and the claim that wavelet losses are unexplored in segmentation deserves a broader citation check.\n\nMy take: this could be a useful loss, but the table must be corrected and rerun, and the abstract must match the asterisks. I would not desk-reject it. Send it to a serious referee who will verify the table and ask for a corrected version. If the MASS ROAD numbers hold after rerunning, the method is likely a solid within-subfield contribution.","headline":"The CWMI loss is a plausible and well-ablated segmentation loss, but Table 1 contains a duplicated MASS ROAD/Attention U-Net block that undermines the abstract's across-dataset claims until it is fixed.","tokens_in":18113,"tokens_out":2461,"would_cite":true,"duration_ms":24907,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a loss function computing mutual information between complex steerable pyramid subbands of prediction and ground truth improves both pixel-wise and topological segmentation metrics across four datasets with minimal…","keywords":["complex wavelet mutual information loss","semantic segmentation","steerable pyramid","class imbalance","instance imbalance","loss function","topological consistency","biomedical image segmentation"],"falsifier":"Train the same U-Net on the same SNEMI3D folds with CWMI and with Dice and RMI losses, freezing all other hyperparameters; if CWMI does not beat both on mIoU and VI across repeated runs, the reported advantage does not generalize. To test the mechanism separately, compare Eq. 9 against a nonparametric mutual-information estimate on identical subbands and check whether training outcomes change.","tokens_in":17108,"feed_emoji":"🩺","tokens_out":9513,"duration_ms":89072,"temperature":0.7,"pith_summary":"CWMI loss is a training objective that measures how much information the predicted segmentation shares with the ground truth in a multi-scale, multi-orientation wavelet representation. The paper argues that pixel-wise losses miss smaller instances and thin boundaries, while regional and topological losses are either local or expensive, and that a complex steerable pyramid exposes structure across scales and orientations at low cost. On four imbalanced datasets, CWMI is reported to outperform eleven existing loss functions on most pixel-wise, clustering, and boundary metrics with U-Net, Attention U-Net, and a Mamba-based U-Net, with per-epoch cost about 0.23 seconds above plain cross entropy. The paper limits its evidence to binary 2D segmentation and flags multi-class and 3D cases as future validation.","feed_headline":"Complex-wavelet MI loss tops 11 segmentation baselines","feed_subtitle":"By comparing predictions with labels in multiscale oriented subbands, CWMI sharpens thin structures at negligible extra cost.","key_machinery":"The engine is the complex steerable pyramid, a redundant wavelet transform with orientation-tuned band-pass filters whose analytic complex subbands encode local phase as well as magnitude, so edges and corners appear as phase structure. For each pyramid level, the K directional subband coefficients at each pixel are treated as a random vector, and mutual information between prediction and label is estimated by the closed-form Gaussian formula $I \\approx -\\frac{1}{2} \\log \\det(M_n)$, extended to complex coefficients through Hermitian covariance and cross-covariance matrices (Eq. 9). Summing $-I$ over the $N$ pyramid levels and combining with cross entropy yields the CWMI loss; because each decomposition level downsamples by a factor of four, the total cost stays near $O(HW \\log(HW))$.","core_discovery":"The central claim is that segmentation quality improves when mutual information is maximized between complex steerable pyramid subbands of the prediction and ground truth, rather than between their raw pixels. The CWMI loss sums, over pyramid levels, a closed-form Gaussian estimate of mutual information between each complex subband pair, adds a cross-entropy term, and is reported to outperform eleven baseline loss functions on the majority of metrics across SNEMI3D, GlaS, DRIVE, and Massachusetts Roads with both U-Net and Attention U-Net, and on SNEMI3D with VMUNet. Its reported gains include overlap measures such as mIoU and mDice but also structural measures such as variation of information, adjusted Rand index, and Hausdorff distance. The paper also claims the complex phase representation carries part of the benefit, since the full CWMI beats real-only, magnitude-only, and phase-only variants in the ablation study.","pith_inferences":["A testable extension beyond the paper: replace the Gaussian mutual-information estimate in Eq. 9 with a nonparametric estimator on the same subbands; if the results change materially, the Gaussian surrogate is load-bearing.","The layer ablation shows the third level is best for regional metrics while the fourth is best for Hausdorff distance, so a per-level weighting scheme could improve on the paper's uniform sum.","Because the loss penalizes structural mismatch in a phase-sensitive, translation-tolerant representation, it should transfer to other dense prediction tasks such as image-to-image translation and super-resolution, which the paper mentions only as future work.","The paper's evidence is restricted to binary 2D segmentation; extending to multi-class and volumetric data is plausible but untested, so the practical scope of the claim is currently narrower than the title suggests."],"forward_implications":["Networks trained with CWMI should segment small instances and thin boundaries better in class- and instance-imbalanced datasets, without any change to the network architecture.","The reported gains on variation of information, adjusted Rand index, and Hausdorff distance mean CWMI can serve as a cheap substitute for expensive topology-preserving losses on thin structures.","The full complex form matters: replacing it with real-only, magnitude-only, or phase-only subbands reduces at least some metrics, so phase information carries part of the signal.","With measured overhead of roughly 0.23 seconds per epoch over cross entropy and complexity near linear in pixel count, CWMI remains practical for high-resolution images.","The loss also improved a Mamba-based U-Net on SNEMI3D, indicating compatibility beyond convolutional attention architectures."],"supporting_citations":[{"why":"Supplies the Gaussian mutual-information estimator and determinant formula that Eq. 7 extends to complex subbands.","marker":"Zhao et al., 2019"},{"why":"Defines the complex steerable pyramid whose phase-sensitive subbands carry the structural features CWMI measures.","marker":"Portilla & Simoncelli, 2000"},{"why":"Introduces the steerable pyramid multiscale and multi-orientation decomposition that the complex version builds on.","marker":"Simoncelli et al., 1992"},{"why":"U-Net is the primary baseline architecture and also the source of weighted cross entropy, one of the compared losses.","marker":"Ronneberger et al., 2015"},{"why":"Attention U-Net provides the second architecture used for generalization testing.","marker":"Oktay et al., 2018"},{"why":"VMUNet provides the Mamba-based architecture used to test CWMI beyond convolutional attention models.","marker":"Ruan et al., 2024"},{"why":"SNEMI3D is one of the four evaluation datasets and the site of the ablation studies.","marker":"Arganda-Carreras et al., 2013"},{"why":"GlaS is the gland segmentation dataset used to evaluate CWMI on histological images.","marker":"Sirinukunwattana et al., 2017"},{"why":"DRIVE is the retinal vessel dataset on which CWMI shows consistent gains across metrics.","marker":"Staal et al., 2004"},{"why":"Massachusetts Roads provides the aerial road segmentation dataset with its class-imbalance setting.","marker":"Mnih, 2013"}],"fun_headline_variants":["CWMI loss: multiscale MI beats 11 segmentation losses","Mutual info on wavelet subbands sharpens segmentation","CWMI loss improves topological metrics in segmentation","Complex wavelet MI loss adds minimal overhead to segmentation","CWMI loss preserves thin boundaries in segmentation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the complex wavelet subband coefficients of prediction and label behave like a jointly Gaussian random vector, so that mutual information can be computed from covariance matrices alone; the paper borrows this approximation from regional mutual information without checking it on binary segmentation masks, whose wavelet coefficients are highly non-Gaussian.","fun_headline_variants_meta":{"raw":{"variants":["CWMI loss: multiscale MI beats 11 segmentation losses","Mutual info on wavelet subbands sharpens segmentation","CWMI loss improves topological metrics in segmentation","Complex wavelet MI loss adds minimal overhead to segmentation","CWMI loss preserves thin boundaries in segmentation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001078,"raw_usage":{"total_tokens":4504,"prompt_tokens":931,"completion_tokens":3573,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":547,"completion_tokens_details":{"reasoning_tokens":3499}},"tokens_in":547,"tokens_out":3573,"duration_ms":25732,"temperature":1.0,"reasoning_tokens":3499,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:31:47.409612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same U-Net on the same SNEMI3D folds with CWMI and with Dice and RMI losses, freezing all other hyperparameters; if CWMI does not beat both on mIoU and VI across repeated runs, the reported advantage does not generalize. To test the mechanism separately, compare Eq. 9 against a nonparametric mutual-information estimate on identical subbands and check whether training outcomes change.","supporting_citations":[{"cited_title":"Region mutual information loss for semantic segmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the Gaussian mutual-information estimator and determinant formula that Eq. 7 extends to complex subbands."},{"cited_title":"and Simoncelli, E","cited_arxiv_id":null,"evidence_quote":"Defines the complex steerable pyramid whose phase-sensitive subbands carry the structural features CWMI measures."},{"cited_title":"P., Freeman, W","cited_arxiv_id":null,"evidence_quote":"Introduces the steerable pyramid multiscale and multi-orientation decomposition that the complex version builds on."},{"cited_title":"S., Vishwanathan, A., and Berger, D","cited_arxiv_id":null,"evidence_quote":"SNEMI3D is one of the four evaluation datasets and the site of the ablation studies."},{"cited_title":"P., Chen, H., Qi, X., Heng, P.-A., Guo, Y","cited_arxiv_id":null,"evidence_quote":"GlaS is the gland segmentation dataset used to evaluate CWMI on histological images."},{"cited_title":"D., Niemeijer, M., Viergever, M","cited_arxiv_id":null,"evidence_quote":"DRIVE is the retinal vessel dataset on which CWMI shows consistent gains across metrics."},{"cited_title":"Machine Learning for Aerial Image Labeling","cited_arxiv_id":null,"evidence_quote":"Massachusetts Roads provides the aerial road segmentation dataset with its class-imbalance setting."}],"review_version":1}