{"id":"1d7812c3-e402-47be-8d71-144a0624e491","arxiv_id":"2607.07581","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":7,"one_line_summary":"STRMSR combines coarse-to-fine cross-view matching with a memory-based propagation mechanism to super-resolve through-plane cardiac MRI slices, improving PSNR at 4x and 8x upsampling.","lead":"The paper presents STRMSR, a deep learning method that uses high-resolution reference views and a memory bank to super-resolve the thick slices of cardiac MRI volumes. A smart generalist might read it because thick-slice acquisition is a standard clinical compromise, and better through-plane resolution directly improves 3D cardiac analysis.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Self-referencing synthetic evaluation and modest LAX gains limit evidence strength for the central claim of consistent improvement.","rationale":"The reader correctly identified clinical applicability as the key concern and arrived at a CONDITIONAL verdict, which remains appropriate. However, the reader's specific framing — that the pre-alignment assumption is the weakest point — targets a limitation the paper explicitly acknowledges and scopes out (Section 1: 'we focus on the second step and assume the input low-resolution stacks are already spatially aligned'). This is a standard decomposition in the SVR literature. The more load-bearing concern is that even within the paper's stated scope, the evaluation methodology limits the strength of evidence: (1) self-referencing views from the same volume mean the reference contains exact ground-truth anatomy, (2) synthetic degradation does not capture real MRI artifacts, (3) improvements on the clinically realistic WHS-LAX protocol are small (0.17-0.19 dB), and (4) the CFCM robustness-to-misalignment claim is never directly tested with induced misalignment. The ablation results (Table 2) show the memory contribution is only +0.12 dB, which is marginal for a claimed main contribution, though temporal profiles suggest qualitative consistency benefits. The paper is technically sound within its scope, the method is well-motivated, and the ablation study is informative. The CONDITIONAL verdict is correct — the method shows promise but needs validation on independent clinical data with realistic acquisition conditions to confirm that the improvements are not artifacts of the self-referencing synthetic setup. No code release further limits reproducibility. No internal inconsistency or correctness issue was identified in the method description.","tokens_in":9965,"tokens_out":4319,"duration_ms":337439,"concrete_test":"Introduce controlled rigid misalignment (random translations of 2-5mm and rotations of 2-5°) between the target and reference views in the WHS test set, then re-evaluate STRMSR vs. McMRSR and MASA. If STRMSR's improvement over baselines shrinks to <0.1 dB PSNR under induced misalignment, the CFCM robustness claim weakens. Additionally, train and test on a second independent cardiac MRI dataset (e.g., ACDC or MESA) with real multi-breath-hold acquisitions to verify that the LAX improvements survive beyond WHS.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that STRMSR achieves 'consistent improvements over baselines at 4x and 8x.' This is supported by experiments on a single dataset (WHS, 30 test subjects) where both the target and reference views are derived from the same underlying 3D volume with synthetic degradation (Gaussian σ=1.0 + block averaging). In WHS-Ortho, the sagittal reference is a degraded orientation of the same volume; in WHS-LAX, the 2/3/4-CH views are resampled from the same 256³ volume. This means the reference contains exact anatomical ground truth from the target volume — no inter-breath-hold motion, no contrast mismatch, no coil sensitivity differences. All baselines use the same references, so the comparison is internally fair, but the absolute performance may not transfer to clinical acquisitions where references are independently acquired. The concern is sharpened by the magnitude of improvements on the more clinically realistic WHS-LAX protocol: +0.17 dB PSNR at ×4 and +0.19 dB at ×8 over the best baseline (McMRSR). These are small effects that could be dataset-specific. Additionally, the paper claims CFCM provides robustness to 'spatial misalignment,' but the only misalignment present is the exact geometric transformation between views — no residual motion is introduced to test this claim. The reader's concern about pre-aligned inputs is related but focuses on a limitation the paper explicitly scopes out; the deeper issue is that even within the stated scope, the evaluation does not test the conditions (realistic degradation, independent references, induced misalignment) under which the method's design choices would matter most.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper proposes STRMSR, a reference- and memory-guided through-plane super-resolution framework for cardiac MRI. The method introduces three components: (1) coarse-to-fine contextual matching (CFCM) for establishing cross-view correspondence, (2) patch-wise dynamic feature aggregation (PDFA) for fusing multiple warped reference/memory features, and (3) a FIFO memory bank that propagates intermediate SR results along the through-plane axis to enforce slice-to-slice consistency. The method is evaluated on the WHS dataset under two protocols (WHS-Ortho and WHS-LAX) at ×4 and ×8 upsampling, with an ablation study isolating each component's contribution.","tokens_in":10239,"tokens_out":2181,"duration_ms":138236,"significance":"The problem of through-plane SR for cardiac MRI is clinically relevant, and the memory-based propagation mechanism is a reasonable adaptation of video object segmentation ideas to volumetric SR. The ablation study (Table 2) properly isolates component contributions, and the use of two reference protocols with different coverage densities provides useful diversity within the single dataset. The ×8 results on WHS-Ortho show a meaningful improvement (+0.97 dB PSNR over McMRSR), which is the most practically relevant regime. Statistical significance is reported via Wilcoxon signed-rank tests.","major_comments":[{"comment":"Section 1, Contribution (1): The paper claims that CFCM enables 'robust detail transfer between target and reference/memory images under cross-view geometric misalignment.' However, no experiment introduces residual misalignment to test this claim. In both WHS-Ortho and WHS-LAX protocols (Section 3.1), the reference views are resampled from the same underlying 3D volume with known geometric transformations and no inter-acquisition motion. The only 'misalignment' present is the exact geometric relationship between orthogonal views, which is trivially known. This is load-bearing because robustness to misalignment is stated as a primary contribution and motivates the coarse-to-fine design. The authors should either (a) add an experiment with controlled synthetic misalignment (e.g., random rigid perturbations applied to reference views) to demonstrate that CFCM degrades gracefully, or (b) re","section":null},{"comment":"Table 1, WHS-LAX rows: The improvements on the more clinically realistic protocol are modest — +0.17 dB PSNR at ×4 and +0.19 dB at ×8 over McMRSR. While statistical significance is reported for ×8, the practical significance of sub-0.2 dB gains is questionable, especially given that the evaluation uses synthetic degradation on a single dataset where references are derived from the same volume. The paper should discuss this limitation explicitly and acknowledge that the stronger WHS-Ortho gains may benefit from the dense reference coverage that is less representative of clinical LAX acquisitions. Without this qualification, the abstract's claim of 'consistent improvements' overstates the WHS-LAX evidence.","section":null}],"minor_comments":[{"comment":"Section 3.2: The memory bank size differs between training and inference (T=2 vs T=10 for WHS-LAX). No justification or ablation is provided for this discrepancy or for the specific values chosen.","section":null},{"comment":"Section 2.5 and Figure 3: The term 'temporal profile' is used for the through-plane spatial axis. While the video analogy is explained, this terminology may confuse readers expecting temporal (cardiac phase) information. Consider clarifying.","section":null},{"comment":"Table 1: STRMSR does not achieve the best SSIM at ×4 on either protocol (WHS-Ortho: 0.9721 vs MsFF-Net's 0.9747; WHS-LAX: 0.8892 vs MsFF-Net's 0.8891). The text mentions 'five out of six metrics at ×4' but this could be stated more prominently to avoid overstating.","section":null},{"comment":"Section 2.2, Eq. (3): The notation switches between superscript s (scale level) and subscript u/v (patch index) without explicit definition of all subscripts. A brief clarification would improve readability.","section":null},{"comment":"Section 3.1: The degradation model (Gaussian σ=1.0 + block averaging) is a standard synthetic approach but does not model realistic MRI acquisition artifacts (coil sensitivity variation, motion, contrast differences). A brief discussion of this limitation would strengthen the paper.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid contribution to the reference-based MRI SR literature, but the gap between the claimed robustness to misalignment and the actual evaluation is the main concern. If the authors can either add a misalignment robustness experiment or substantially qualify their claims, the paper would be appropriate for the journal. The single-dataset evaluation is a limitation but is common in this subfield; I would not insist on a second dataset for acceptance, but the authors should be transparent about it."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive review. The two major comments are well-taken, and we address each below. In short: (1) we agree that the robustness-to-misalignment claim is not directly validated by the current experiments and will revise the wording and add a controlled perturbation experiment; (2) we agree that the WHS-LAX gains are modest and will add an explicit discussion of this limitation, including qualifying the abstract's language.","responses":[{"response":"We agree with this comment. The referee is correct that the current experiments do not directly test robustness to residual registration error. We want to clarify the intended meaning: the 'misalignment' we refer to is the inherent cross-view geometric discrepancy between orthogonal or LAX views and the target SAX stack — i.e., the reference and target slices are not parallel and their anatomical content is related by a non-trivial geometric transformation, not a simple translation. The coarse-to-fine design is motivated by this cross-view discrepancy, which makes dense correspondence non-trivial even when the inter-acquisition transformation is known. However, we acknowledge that the wording in Contribution (1) — 'robust detail transfer under cross-view geometric misalignment' — can be read as a claim about robustness to residual registration error, which we do not validate. We will take the following steps in the revision: (a) We will add an experiment applying random rigid perturbations (translations and rotations) to the reference views at controlled magnitudes to demonstrate that CFCM degrades gracefully relative to the single-stage matching baseline (McMRSR-style). (b) We will revise the contribution statement to distinguish clearly between (i) the inherent cross-view geometric discrepancy that CFCM is designed to handle and (ii) residual registration error, which is outside the scope of the current method given our stated assumption of pre-aligned inputs (Section 1). We believe this addresses the referee's concern without overstating our claims.","revision_made":"yes","referee_comment":"Section 1, Contribution (1): The paper claims CFCM enables robust detail transfer under cross-view geometric misalignment, but no experiment introduces residual misalignment. References are resampled from the same volume with known transformations and no inter-acquisition motion. The authors should either add a misalignment experiment or revise the claim."},{"response":"We agree that the WHS-LAX gains are modest and that this should be discussed more transparently. We will make the following changes: (a) Add an explicit paragraph in Section 3.4 acknowledging that the WHS-LAX improvements are smaller in absolute terms and discussing the likely cause — the LAX protocol provides only three 2-chamber, 3-chamber, and 4-chamber views as references, which cover a small fraction of the target SAX slices, whereas WHS-Ortho uses a dense sagittal reference volume that intersects every axial slice. The denser reference coverage in WHS-Ortho provides more matching opportunities and thus a larger performance gap. (b) Acknowledge that the WHS-Ortho results, while demonstrating the method's potential under favorable reference coverage, may not generalize directly to typical clinical LAX acquisitions where reference views are sparse. (c) Revise the abstract to replace 'consistent improvements' with more precise language, e.g., 'improvements over baselines at ×4 and ×8 upsampling factors, with larger gains under dense reference coverage.' We note that the ×8 WHS-LAX improvement, while small in absolute terms (0.19 dB), is statistically significant (p < 0.001 by Wilcoxon signed-rank test) and consistent across all three metrics (PSNR, SSIM, MSE). We will report effect sizes more carefully and let the reader judge practical significance rather than characterizing the gains as large. We respectfully note that even modest gains at ×8 — the most clinically common through-plane spacing — can be meaningful when no existing method provides a significant improvement, but we agree the paper should not overstate this.","revision_made":"yes","referee_comment":"Table 1, WHS-LAX rows: Improvements are modest (+0.17 dB PSNR at ×4, +0.19 dB at ×8). Practical significance of sub-0.2 dB gains is questionable. The paper should discuss this limitation explicitly and acknowledge that stronger WHS-Ortho gains may benefit from dense reference coverage less representative of clinical LAX acquisitions. The abstract's 'consistent improvements' overstates the WHS-LAX evidence."}],"tokens_in":9739,"tokens_out":1243,"duration_ms":110781,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Bottom line: STRMSR combines coarse-to-fine matching with a memory-based inter-slice propagation mechanism for cardiac MRI through-plane SR, and the idea works. The gains at ×8 on WHS-Ortho are meaningful (+0.97 dB PSNR over McMRSR), and the ablation cleanly isolates each component. The main weakness is that all evaluation uses synthetic degradation on a single dataset, which limits how much we can trust the absolute numbers for clinical transfer.","headline":"Solid technical contribution with a real limitation in evaluation scope — synthetic degradation on a single dataset","tokens_in":10757,"tokens_out":807,"would_cite":false,"duration_ms":42288,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["87.61.-c","87.57.N-","87.57.nm"],"model":"glm-5.2","headline":"Memory and reference views sharpen cardiac MRI through-plane","keywords":[],"falsifier":"Apply STRMSR to cardiac MRI data with realistic residual inter-slice misalignment (without prior registration correction) and measure whether the coarse-to-fine contextual matching produces correct correspondences or amplifies artifacts. If the matching degrades under misalignment, the assumption of pre-aligned input is load-bearing rather than incidental.","tokens_in":10310,"feed_emoji":"🫀","tokens_out":828,"duration_ms":163266,"temperature":0.7,"pith_summary":"Clinical cardiac MRI scans capture fine detail within each slice but leave thick gaps between slices, because patients can only hold their breath briefly and the heart keeps moving. This paper proposes STRMSR, a framework that reconstructs high-resolution 3D cardiac volumes from these anisotropic scans by borrowing sharp structural detail from high-resolution reference views of the same subject and from the method's own intermediate super-resolution results stored in a memory bank. The central mechanism is a three-part pipeline: coarse-to-fine contextual matching (CFCM) progressively refines correspondences between the low-resolution target and the high-resolution reference across multiple scales, handling spatial misalignment between views; patch-wise dynamic feature aggregation (PDFA) computes content-adaptive mixture weights for each local patch so that reliable detail transfers are amplified and inconsistent matches are suppressed; and a memory-based propagation step treats the volume's through-plane axis like a temporal sequence, storing recently super-resolved slices to guide the next slice, enforcing continuity across the 3D volume. The paper validates STRMSR on the WHS cardiac MRI dataset under two reference protocols (orthogonal-plane and long-axis chamber views) at 4x and 8x upsampling, reporting consistent improvements over five baselines, with the largest gains at 8x where reference information is sparsest.","feed_headline":"Memory and reference views sharpen cardiac MRI through-plane","feed_subtitle":"STRMSR borrows detail from high-resolution reference slices and its own prior outputs to reconstruct thick-gap cardiac volumes, with largest","key_machinery":"Coarse-to-Fine Contextual Matching (CFCM), Patch-wise Dynamic Feature Aggregation (PDFA), and Memory-based SR Propagation","core_discovery":"The paper demonstrates that combining multi-scale correspondence refinement, patch-level adaptive fusion, and inter-slice memory propagation yields measurable improvements in through-plane cardiac MRI super-resolution, particularly at 8x upsampling where existing methods degrade most. The ablation study shows that replacing the coarse-to-fine matching with a single-scale matching strategy causes the largest performance drop (-0.60 dB PSNR), followed by removing the dynamic aggregation (-0.24 dB) and the memory bank (-0.12 dB), indicating that all three components contribute complementarily but correspondence refinement is the most critical. The performance advantage over baselines grows as 8","pith_inferences":[],"forward_implications":["If the memory-propagation approach generalizes, the same mechanism could be applied to other anisotropic medical imaging modalities where slice-to-slice consistency is clinically important, such as fetal brain MRI or abdominal imaging.","The finding that gains are largest at 8x upsampling—where clinical breath-hold constraints are most binding—suggests the method is most useful in exactly the regime that dominates real clinical cardiac MRI acquisition.","The patch-wise dynamic aggregation strategy, interpreted as a content-adaptive mixture-of-experts over warped reference features, could be adopted more broadly in any multi-reference image fusion task where some references are more informative than others for a given spatial region.","The memory bank's expansion from 2-3 frames during training to 10 frames during inference for the long-axis protocol suggests the method benefits from longer temporal context at test time, raising the question of whether even larger memory banks would yield further gains."],"fun_headline_variants":["STRMSR sharpens cardiac MRI through-plane via memory and references","Reference and memory guidance boost cardiac MRI through-plane SR","Coarse-to-fine matching and memory bank refine cardiac MRI through-plane SR","STRMSR uses reference views and prior outputs to reconstruct thick-gap cardiac MRI","Memory-guided super-resolution restores slice consistency in cardiac MRI"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The method assumes that the input low-resolution image stacks are already spatially aligned. If residual breath-hold or cardiac motion misalignment persists in clinical data, the coarse-to-fine matching could establish incorrect correspondences and degrade the output.","fun_headline_variants_meta":{"raw":{"variants":["STRMSR sharpens cardiac MRI through-plane via memory and references","Reference and memory guidance boost cardiac MRI through-plane SR","Coarse-to-fine matching and memory bank refine cardiac MRI through-plane SR","STRMSR uses reference views and prior outputs to reconstruct thick-gap cardiac MRI","Memory-guided super-resolution restores slice consistency in cardiac MRI"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1169,"prompt_tokens":496,"completion_tokens":673,"prompt_tokens_details":null},"tokens_in":496,"tokens_out":673,"duration_ms":29625,"temperature":1.0,"reasoning_tokens":660,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T06:05:07.975747+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Apply STRMSR to cardiac MRI data with realistic residual inter-slice misalignment (without prior registration correction) and measure whether the coarse-to-fine contextual matching produces correct correspondences or amplifies artifacts. If the matching degrades under misalignment, the assumption of pre-aligned input is load-bearing rather than incidental.","supporting_citations":[],"review_version":1}