{"id":"5fb54084-e5ce-41ea-b706-e7ddbb21ee89","arxiv_id":"2507.11886","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"CAA-Seg combines selective slice alignment with hierarchical feature fusion to improve infarction, edema, and myocardium segmentation on a 397-patient multi-sequence CMR dataset.","lead":"CAA-Seg, a two-stage alignment-aware deep learning framework, improves myocardial lesion segmentation from multi-sequence cardiac MRI, especially infarction, reaching 51.11% Dice, 5.54% higher than the next best method. It targets a real clinical problem: heart MRI sequences are often acquired at mismatched slice positions, so its selective slice alignment could make automated cardiac assessment more reliable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 5.54-point MI gain is confounded by an input-channel asymmetry: general baselines get only LGE while CAA-Seg gets LGE+T1m+T2m, so the margin may reflect extra sequences rather than the alignment framework.","rationale":"The reader identified the input-channel asymmetry as the weakest assumption, and this is also the most load-bearing concern I find. The paper's central quantitative claim is a 5.54% MI Dice improvement over the second-best method, but the comparison gives CAA-Seg three sequences and the strongest general baselines only one. The text explicitly states that general frameworks were restricted to LGE because of performance degradation with multi-sequence input, but it does not show those results or demonstrate that SSA preprocessing would not eliminate the degradation. Since the proposed framework is itself a fusion of SSA plus HA-Net, the experiment that would isolate the alignment contribution is a general segmentation network trained on the same SSA-aligned multi-sequence data. The ablation table does not include this cell, so the superiority claim is conditional on an untested comparison protocol. I do not see a stronger internal inconsistency: the method description, including the mutual-information formulation, is coherent enough to reproduce, and the reported improvements are plausible for a well-tuned alignment pipeline. The correct next step is therefore to keep the CONDITIONAL verdict and require the specific control experiment plus confidence intervals before accepting the superiority claim. This is not an objection to the engineering contribution itself; it is a constraint on how the empirical evidence can be interpreted.","tokens_in":7779,"tokens_out":8015,"duration_ms":91356,"concrete_test":"Run nnU-Net and UMamba on the same three-sequence input used by CAA-Seg, applying the proposed SSA slice alignment identically as preprocessing (or, minimally, feeding the same resampled multi-channel volumes), with the same train/validation/test split, loss, and 200-epoch training schedule. If either baseline closes the 5.54-point MI Dice gap or surpasses 51.11, the reported advantage is explained by additional input information rather than by the alignment modules. Also report bootstrap 95% confidence intervals over the MI test subset to check stability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that CAA-Seg outperforms the best prior method by 5.54% Dice on myocardial infarction (51.11 vs 45.57, Table 1) and that this improvement is produced by the proposed composite alignment framework. The protocol in Section 3 does not support that attribution. General frameworks (nnU-Net, UNet++, TransUNet, Swin UNETR, UTNet, UMamba) were evaluated using only LGE images, citing performance degradation with multi-sequence input, while CAA-Seg and cardiac-specific baselines received LGE plus T1m/T2m. The comparison therefore varies the method and the available input information at the same time. The paper does not report the degraded multi-input results it used to justify restricting the baselines, nor does it test the critical control: a general framework fed with the same SSA-aligned multi-sequence volumes. The ablation in Table 2 only varies registration, network, and input within the proposed pipeline; it does not include a general network trained on SSA-aligned multi-sequence input. Without that control, the 5.54-point MI margin cannot be uniquely credited to alignment-aware fusion. The MI result is also based on only 208 of 397 patients, with no confidence intervals, so the margin may not be stable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript proposes CAA-Seg, a two-stage framework for myocardial lesion segmentation from multi-sequence CMR images. The first stage, selective slice alignment (SSA), matches LGE slices to corresponding T1-mapping and T2-mapping slices via mutual-information optimization with a sliding-window constraint. The second stage, a hierarchical alignment network (HA-Net), fuses features with local deformable-convolution corrections and global cross-attention, plus a task-aware controller at the bottleneck. The authors evaluate on a 397-patient in-house dataset and report Dice and HD95 for myocardium, edema, and infarction, claiming a 5.54-point improvement over the second-best method for infarction segmentation (Table 1). Ablations vary registration, network, and input modality within their pipeline (Table 2). Code is provided.","tokens_in":8070,"tokens_out":5335,"duration_ms":58501,"significance":"If the results are reproducible, CAA-Seg addresses a real clinical problem—heterogeneous slice sampling across CMR sequences—and the two-stage alignment approach is a sensible design. The paper's strengths include a relatively large dataset (397 patients), release of code, an ablation study, and statistical significance testing with Wilcoxon tests. However, the central quantitative claim is currently undermined by a comparison protocol that varies the method and the input modality simultaneously, and by the absence of variance estimates and external validation. The contribution is potentially significant for the multi-sequence CMR segmentation community, but the empirical evidence as presented does not yet support the stated superiority.","major_comments":[{"comment":"The comparison is confounded: general frameworks (nnU-Net, UNet++, TransUNet, Swin UNETR, UTNet, UMamba) were evaluated using only LGE images, while CAA-Seg and cardiac-specific baselines received LGE plus T1m/T2m. Thus the reported 5.54-point MI improvement (51.11 vs 45.57) may reflect the additional input sequences rather than the proposed alignment and fusion modules. The authors should either report the multi-sequence-input results for the general baselines (they state these degraded but do not show the numbers) or add a control condition in which a general network is trained on SSA-aligned multi-sequence volumes. Without that control, the central superiority claim is not supported.","section":"Section 3, Experimental Setup; Table 1"},{"comment":"The ablation rows vary registration (MvMM vs SSA), network (nnU-Net vs HA-Net), and input (LGE vs Multi), but all conditions operate on the authors' own pipeline. There is no row combining SSA-aligned multi-sequence input with a general baseline network such as nnU-Net. Consequently, the ablation cannot isolate the effect of the alignment mechanism from the effect of the additional sequences, and it does not provide the control needed to interpret the comparison in Table 1.","section":"Table 2, Ablation"},{"comment":"Only 208 of 397 patients have MI annotations, and the paper reports no confidence intervals, standard deviations, or repeated-run variability for the Dice/HD95 values. The reported p-values from the Wilcoxon signed rank test indicate significance but not effect-size stability. Given the small MI target and the single in-house dataset, the 'substantial 5.54% improvement' may not be robust. The authors should report bootstrap or repeated-run intervals and, ideally, validate on an external or public dataset such as MyoPS.","section":"Section 3, Dataset paragraph; Table 1"},{"comment":"The selective slice alignment method is not fully specified. The values of λ, the search-range parameters N and M, and the window size are not given. Furthermore, the text and Fig. 2 indicate both T1m and T2m are aligned to LGE, but the manuscript does not explain how the two moving sequences (each with only 2-3 slices, as stated in the Introduction) are jointly handled by the single optimization in Eq. (4), nor how the 'LGE x2' duplicated input is generated and used. These details are needed to reproduce the method and to assess whether the alignment is anatomically meaningful.","section":"Section 2.1, Eqs. (1)-(7)"}],"minor_comments":[{"comment":"The method name 'A WSNet' should be 'AWSNet' to match reference [14]; please unify the notation across the table and text.","section":"Table 1"},{"comment":"The table header 'Settings Registration Network Input Overall' is not aligned with the row contents; the row labels (e.g., 'Baseline', 'MvMM', 'SSA') are not separated into clear columns, making it difficult to read which configuration corresponds to each Dice/HD95 value. Please reformat the table.","section":"Table 2"},{"comment":"In Eq. (9), γ and β are described as 'pixel-wise' modulation parameters but are produced from a GlobalPool operation, which suggests they are channel-wise; please specify the exact dimensions and operation.","section":"Section 2.2, Eq. (9)"},{"comment":"The task prompt Ptask is not defined; specify how it is initialized, whether it is learned, and how the task ID (0 for LGE+T1m+T2m, 1 for LGE x3 in Fig. 2) is encoded.","section":"Section 2.2, Eq. (11)"},{"comment":"The Wilcoxon signed rank test is used, but the pairing unit (slice or patient) is not stated; please clarify.","section":"Section 3, Statistical analysis"},{"comment":"The claim that general frameworks show 'performance degradation with multi-sequence input' is not supported by any reported results; at minimum, include the corresponding numbers in the supplementary material.","section":"Section 3, Experimental Setup"}],"recommendation":"major_revision","confidential_remarks":"The comparison asymmetry is the main risk to the paper's central claim. The authors should be asked to supply the multi-input baseline numbers or a nnU-Net control condition; without this, the headline improvement is not defensible. There is no indication of misconduct; the issue is methodological presentation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: this is a competent engineering paper, but treat the headline MI margin as unproven until the authors run the missing control.\n\nWhat's new: a selective slice alignment (SSA) stage that matches LGE to T1/T2 slices with mutual information and a windowed search, plus a hierarchical alignment network (HA-Net) that uses deformable convolutions at low levels and cascaded cross-attention at high levels. The idea of explicitly rejecting mismatched slices and aligning only reliable correspondences is a sensible response to a real clinical problem, and the ablation in Table 2 does show each component earns its keep. Credit also for a 397-patient dataset and for releasing code.\n\nThe soft spot is the evaluation protocol. General baselines (nnU-Net, UNet++, TransUNet, etc.) are run on LGE only, with the paper claiming multi-sequence input caused performance degradation, while CAA-Seg and the cardiac-specific methods get all three sequences. That varies the method and the available input information at the same time. The 5.54-point gain on MI over UMamba could come from having T1/T2, not from the alignment mechanism. The paper doesn't report the degraded multi-sequence results for the general baselines, and the ablation doesn't include the obvious control: a general U-Net or nnU-Net trained on the same SSA-aligned multi-sequence volumes. Without that, the central attribution doesn't hold.\n\nOther weaknesses are less severe but real: single in-house dataset, no external validation, no confidence intervals or standard deviations, and the MI metric is based on only 208 of 397 patients. The Wilcoxon test is fine, but p<0.05 alone is thin.\n\nIf the authors rerun the general baselines on aligned multi-sequence input and get the same margin, that's a genuinely useful result. As it stands, the paper is worth refereeing, but the reviewer should push for the control experiment and public-data validation. My sense is the fair verdict is major revision, not accept-as-is.","headline":"The architecture is thoughtful and the ablation is helpful, but the headline 5.54-point MI gain is confounded by an input-channel asymmetry the paper never controls for.","tokens_in":8606,"tokens_out":2230,"would_cite":false,"duration_ms":25774,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CAA-Seg, a two-stage alignment-aware framework, reports infarction Dice of 51.11%, beating the previous best by 5.54 points.","keywords":["multi-sequence CMR","myocardial lesion segmentation","selective slice alignment","hierarchical feature fusion","deformable convolution","cross-attention","late gadolinium enhancement","T1/T2 mapping"],"falsifier":"Run nnU-Net or UMamba on the identical three-sequence, SSA-aligned input that CAA-Seg receives, keeping every other training setting the same; if infarction Dice moves from 45.30% toward 51.11%, the headline margin is an input advantage rather than an alignment effect. A second check: replace the MMI-optimized slice matching with order-preserving random pairing inside the full pipeline; if the overall 64.89% Dice is unchanged, the selective slice alignment carries none of the gain.","tokens_in":7595,"feed_emoji":"🫀","tokens_out":12878,"duration_ms":134051,"temperature":0.7,"pith_summary":"Multi-sequence cardiac MRI carries complementary markers for myocardial infarction and edema, but the sequences are acquired with different slice counts and positions, so fusing them naively forces anatomically false correspondences or interpolation artifacts. The paper's claim is that the fix is composite alignment: first match only the slice pairs that truly correspond anatomically between LGE and T1/T2 mapping sequences, then align the merged features at two semantic levels. Validated on 397 patients, the framework reports 78.09% Dice for myocardium, 65.49% for edema, and 51.11% for infarction, beating the second-best method by 5.54 points on infarction. A sympathetic reading is that selective slice correspondence, rather than simply more fusion capacity, is what makes multi-sequence CMR segmentation work on routine, imperfectly aligned acquisitions.","feed_headline":"Infarction Dice tops 51%, 5.54 points beyond prior best","feed_subtitle":"Heart MRI slices rarely line up between sequences; selective matching lifts infarction Dice to 51%.","key_machinery":"The load-bearing object is the selective slice alignment (SSA) scheme, an optimization over slice correspondences: for each LGE slice the method chooses a moving T1m/T2m slice by minimizing $L_{MMI}(I_f^k, I_m^j \\circ \\phi_k) + \\lambda R(\\phi_k)$ over a sliding window $j \\in [j_{k-1}, N-M+k]$, where Mattes mutual information scores multi-sequence similarity and the window enforces sequential order while reserving enough slices for later matches. This yields a registered volume built only from reliably matched slice pairs. The second mechanism is the hierarchical alignment network (HA-Net), which splits the residual alignment problem by feature level: low-level features pass through deformable convolution offset fields with learned $\\gamma,\\beta$ modulation, high-level features pass through two cascaded cross-attention blocks that query the LGE features against T1m and then T2m, and a task-aware controller injects a learned prompt at the bottleneck. The division of labor—slice selection first, then feature-level correction—is what the paper identifies as the source of its infarction-segmentation advantage.","core_discovery":"The paper argues that the two barriers to multi-sequence CMR lesion segmentation—anatomical slice mismatch caused by different acquisition protocols and intensity variation across sequences—should be handled as separate problems before segmentation. Stage one, selective slice alignment, searches paired LGE and T1/T2 mapping slices for the anatomically most plausible correspondences, scoring candidate pairs with Mattes mutual information and constraining the search so that slice order is preserved and enough slices remain for the rest of the sequence; mismatched pairs are excluded rather than force-registered. Stage two, the hierarchical alignment network, corrects residual local deformation in low-level features with deformable convolutions and pixel-wise modulation, fuses high-level semantic features through cascaded cross-attention, and conditions the bottleneck on a task prompt. The authors report that the full system reaches an overall Dice of 64.89% and an infarction Dice of 51.11%, a 5.54-point gain over the second-best method, and they attribute the gain to the alignment-aware design.","pith_inferences":["Editorially, the 5.54-point infarction margin is not fully isolated from the input advantage: because the general baselines were evaluated with LGE only, a same-input comparison would be needed to attribute the gain specifically to the alignment machinery.","Editorially, the sliding-window slice matching could be made differentiable and folded into the network as a soft correspondence layer, letting registration and segmentation objectives train jointly rather than as separate stages.","Editorially, the same two-stage selection-then-alignment recipe should apply to other cross-modality imaging problems with asymmetric slice sampling, such as MRI-to-CT or echocardiography-to-CMR, where dense registration is currently assumed rather than selectively chosen."],"forward_implications":["If the reported numbers hold, clinical pipelines could ingest unaligned LGE and T1/T2 acquisitions directly, without manual slice screening or aggressive resampling of 3-slice maps to 8-slice volumes.","The selective slice matching principle transfers to other multi-sequence settings with unequal slice sampling, such as combining cine with LGE or adding parametric mapping sequences.","The largest measured gain is on infarction, the smallest and most clinically consequential target, which suggests the alignment machinery matters most where lesions are subtle and easy to corrupt.","The accuracy gain costs modest compute: 0.71 s and 1.8 GB per case versus 0.52 s and 1.1 GB for nnU-Net, which the paper presents as an acceptable trade."],"supporting_citations":[{"why":"MFU-Net, the max-fusion multi-sequence baseline that CAA-Seg is directly compared against in Table 1.","marker":"[13]"},{"why":"AWSNet, the auto-weighted supervision attention baseline whose multi-sequence fusion is outscored in Table 1.","marker":"[14]"},{"why":"MyoPS-Net, the flexible multi-sequence combination baseline and the reference method for cardiac-specific fusion.","marker":"[15]"},{"why":"The MyoPS benchmark paper that documents the 25-patient, pre-aligned dataset on which earlier fusion methods were tested, framing the misalignment problem this paper targets.","marker":"[17]"},{"why":"nnU-Net, the strongest general baseline and the source of the identical preprocessing, training, and postprocessing protocol applied to every compared method.","marker":"[19]"},{"why":"UMamba, the closest overall competitor in Table 1, which trails on infarction detection despite competitive myocardium and edema scores.","marker":"[24]"},{"why":"MvMM, the multivariate mixture model registration used as the alternative to selective slice alignment in the ablation study.","marker":"[25]"}],"fun_headline_variants":["Alignment-aware CAA-Seg boosts infarction Dice by 5.54 points","Infarction Dice hits 51.11% after alignment-aware CMR fusion","Selective slice alignment drives 5.54-point infarction Dice gain","CAA-Seg: hierarchical alignment lifts cardiac lesion Dice to 51.11%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison protocol gives the general baselines only LGE images while CAA-Seg receives all three sequences, so the 5.54-point infarction margin is credited to the alignment framework even though richer input alone could explain part of it.","fun_headline_variants_meta":{"raw":{"variants":["Alignment-aware CAA-Seg boosts infarction Dice by 5.54 points","Infarction Dice hits 51.11% after alignment-aware CMR fusion","Selective slice alignment drives 5.54-point infarction Dice gain","CAA-Seg: hierarchical alignment lifts cardiac lesion Dice to 51.11%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001188,"raw_usage":{"total_tokens":4906,"prompt_tokens":950,"completion_tokens":3956,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":3872}},"tokens_in":566,"tokens_out":3956,"duration_ms":30741,"temperature":1.0,"reasoning_tokens":3872,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:58:52.282940+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run nnU-Net or UMamba on the identical three-sequence, SSA-aligned input that CAA-Seg receives, keeping every other training setting the same; if infarction Dice moves from 45.30% toward 51.11%, the headline margin is an input advantage rather than an alignment effect. A second check: replace the MMI-optimized slice matching with order-preserving random pairing inside the full pipeline; if the overall 64.89% Dice is unchanged, the selective slice alignment carries none of the gain.","supporting_citations":[{"cited_title":"Max-fusion u-net for multi-modal pathology segmentation with attention and dynamic resampling","cited_arxiv_id":null,"evidence_quote":"MFU-Net, the max-fusion multi-sequence baseline that CAA-Seg is directly compared against in Table 1."},{"cited_title":"Awsnet: An auto-weighted su- pervision attention network for myocardial scar and edema segmentation in multi- sequence cardiac magnetic resonance images","cited_arxiv_id":null,"evidence_quote":"AWSNet, the auto-weighted supervision attention baseline whose multi-sequence fusion is outscored in Table 1."},{"cited_title":"Myops-net: Myocardial pathology segmentation with flexible combination of multi-sequence cmr images","cited_arxiv_id":null,"evidence_quote":"MyoPS-Net, the flexible multi-sequence combination baseline and the reference method for cardiac-specific fusion."},{"cited_title":"Myops: A benchmark of myocardial pathology segmentation combining three-sequence car- diac magnetic resonance images","cited_arxiv_id":null,"evidence_quote":"The MyoPS benchmark paper that documents the 25-patient, pre-aligned dataset on which earlier fusion methods were tested, framing the misalignment problem this paper targets."},{"cited_title":"nnu-net: a self-configuring method for deep learning-based biomedical image segmentation","cited_arxiv_id":null,"evidence_quote":"nnU-Net, the strongest general baseline and the source of the identical preprocessing, training, and postprocessing protocol applied to every compared method."},{"cited_title":"Multivariate mixture model for myocardial segmentation com- bining multi-source images","cited_arxiv_id":null,"evidence_quote":"MvMM, the multivariate mixture model registration used as the alternative to selective slice alignment in the ablation study."}],"review_version":1}