{"id":"6b8f0a32-66d7-490e-bba5-ade08d9bcf1f","arxiv_id":"2505.00643","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A deep-learning ghost-removal and outer-volume-subtraction pipeline enables 8x-accelerated real-time cardiac MRI with image quality close to slower clinical baselines.","lead":"This paper proposes a post-processing method that estimates and subtracts signals from tissues outside the heart so that real-time cardiac MRI can run at higher acceleration without buying new hardware. In tests, the method produced images at 8x acceleration that looked close to slower clinical scans.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sensitivity maps estimated from OVR-altered k-space data could bias the encoding operator and undermine the claimed quantitative gains.","rationale":"The reader's weakest_assumption centers on the Eq. 1-2 decomposition and the TGRAPPA R=4 ghost labels. That is a legitimate concern, but the paper partially addresses it in Sec. 4.2.1 by noting the R=4 TGRAPPA labels are noisy and that the network benefits from higher-SNR composite inputs, which mitigates label-noise bias. The more load-bearing gap is the sensitivity-map estimation, because the entire PD-DL encoding operator and the SSDU loss in Eqs. 6-8 depend on it, and the paper never states how the maps were computed for the OVR data. If the maps are derived from the same composite R=8 data that contain the ghosting artifacts, the subtraction in Eq. 5 cannot cleanly separate outer volume from ROI signal, and the comparison against both non-OVR PD-DL and the BH reference becomes biased in an unquantified way. This is a single, concrete, checkable assumption that gates the central claim of diagnostic equivalence at R=8. The verdict should remain CONDITIONAL: the fix requires either a stated and validated sensitivity-map protocol or a sensitivity analysis demonstrating that the result is invariant to the map estimation choice.","tokens_in":18361,"tokens_out":1519,"duration_ms":15738,"concrete_test":"Re-run both the retrospective and prospective pipelines twice: once with sensitivity maps estimated from fully sampled or BH reference k-space (or from calibration data acquired before OVR), and once with maps estimated exactly as in the paper from composite R=8 k-space. If the proposed method's PSNR/SSIM and LV-function agreement relative to the baseline change by more than a small margin (e.g., EF change > 2 percentage points or PSNR shift > 0.5 dB), the headline claim depends critically on the sensitivity-map estimation choice and must be re-scoped.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that the proposed OVR pipeline at R=8 matches clinical baselines depends on the encoding operator E in Eqs. 6-8 being a faithful forward model. The paper never states how sensitivity maps are estimated for the OVR data. If they are computed from the same composite R=8 k-space data that contain the ghosting artifacts and outer-volume signal the method is designed to remove, then the subtraction in Eq. 5 does not cleanly separate outer volume from ROI signal: the sensitivity maps already include motion-blurred outer-volume contributions inside the ROI, so the data-consistency term can force the network to explain residual outer-volume signal within the heart. This would bias the reconstruction toward the artifacts the method claims to remove and would contaminate the comparison against both non-OVR PD-DL and the BH reference. The paper reports no sensitivity analysis on this choice, so the claimed quantitative superiority and LV-function agreement could reflect this bias rather than the OVR decomposition itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a post-processing outer volume removal (OVR) framework for highly accelerated real-time cine cardiac MRI. The authors characterize the pseudo-periodic ghosting that appears in low-temporal-resolution composite images formed from time-interleaved shifted undersampling (Sec. 2, Eqs. 1-2), train a ResNet to predict these ghosts from four adjacent composite images (Sec. 3.1, Eq. 3), subtract the resulting clean outer-volume background from each timeframe's k-space (Eq. 5), and reconstruct the residual ROI data with a 35-unroll physics-driven network trained with multi-mask SSDU and a new dual-sensitivity consistency loss (Sec. 3.3, Eqs. 7-8). Evaluations include retrospectively R=8-undersampled bSSFP data (3 test subjects; Table 1) and prospectively R=8-acquired GRE data with breath-hold segmented cine as the clinical baseline (6 test subjects; Table 2), with TGRAPPA and a matched non-OVR PD-DL network as controls. The paper claims visual parity with clinical baselines and quantitative superiority over conventional reconstructions at R=8.","tokens_in":18529,"tokens_out":11586,"duration_ms":110781,"significance":"The core idea is attractive and timely: it removes aliasing from extra-cardiac tissue in post-processing rather than through acquisition-level suppression, so it is applicable to existing protocols, and it is evaluated with a genuine prospective R=8 acquisition plus a matched non-OVR deep-learning control. The analytical decomposition of composite images into an averaged moving component, a pseudo-periodic ghost, and a stationary background (Sec. 2) is a genuine contribution, and the masked/full-sensitivity consistency loss (Eq. 8) is a clever mechanism for avoiding both residual-background artifacts and ROI signal loss. The manuscript promises code release upon acceptance, which would be valuable for reproducibility. However, the significance of the claimed quantitative gains is currently limited by thin and partly self-referential evidence: the retrospective LV analysis uses only three subjects and shows sizable differences; the reported PSNR/SSIM values are never disclosed; and the prospective agreement claim rests on a low-power n=6 paired t-test. These gaps, not the plausibility of the method, are the main barrier to accepting the abstract's claims.","major_comments":[{"comment":"The text states that 'Quantitative evaluations of PSNR and SSIM values ... further support the improvements,' but no PSNR or SSIM numbers appear anywhere in the manuscript. Since the abstract's claim of 'outperforming ... quantitatively' for the retrospective study rests on this sentence, the authors should report numeric PSNR and SSIM (mean and SD over the 3 test subjects) for TGRAPPA R=8, PD-DL without OVR, and the proposed method, and state the reference used (TGRAPPA R=4 or BH) and whether the metric is evaluated over the full FOV or restricted to the ROI.","section":"5.1.2"},{"comment":"The retrospective LV analysis on n=3 subjects reports EF 55.0 (6) versus 65.4 (7), ESV 56.0 (13) versus 42.2 (15), and ESM 137.7 (30) versus 105.7 (24) against the TGRAPPA R=4 baseline. These are not 'close agreement' in any conventional sense, and with n=3 and no statistical test, the claim in Sec. 5.1.3 is unsupported. The authors should report subject-level values with limits of agreement, add a proper statistical treatment, or substantially temper the retrospective agreement claim.","section":"Table 1 / Sec. 5.1.3"},{"comment":"The ghost-detection network is trained to reproduce composite-minus-TGRAPPA(R=4) reference labels, and the retrospective quantitative comparison is computed against the same TGRAPPA(R=4) reconstructions. Moreover, the final images recombine the masked background estimate (Fig. 3e), which is trained to resemble TGRAPPA(R=4) in the outer volume. The claimed quantitative advantage over the baseline is therefore entangled with the training target, and whole-FOV PSNR/SSIM against TGRAPPA(R=4) does not provide an independent evaluation. Metrics restricted to the ROI, and a comparison against the BH reference for the prospective data, would decouple the evaluation from the training labels.","section":"4.2.1 vs 5.1.2 and Fig. 3e"},{"comment":"The encoding operator E in Eqs. (6)-(8) requires coil sensitivity maps, but the manuscript does not state how these are estimated for the OVR pipeline (e.g., from raw R=8 k-space, from composite data, from OVR-subtracted data, or from a separate calibration scan). If sensitivities are computed from data containing the ghosting and outer-volume signal that the method is designed to remove, the forward model becomes inconsistent with the OVR target, and the improvements attributed to the Eq. (8) consistency loss could reflect that bias. Specify the estimation procedure and include a sensitivity analysis comparing reconstructions obtained with sensitivities from unmodified versus OVR-processed data.","section":"3.3 / 4.2.3"},{"comment":"For the prospective study, 'strong agreement ... with no statistical differences (P > .05 in all cases)' is inferred from a paired t-test on n=6 subjects. Failure to reject the null at n=6 has low power and cannot establish agreement or equivalence. Report Bland-Altman limits of agreement or an equivalence test, together with subject-level EDV/ESV/EF values, and discuss the fact that the BH comparator was acquired in a separate scan.","section":"5.2.3 / Sec. 4.3"}],"minor_comments":[{"comment":"Learning rates are printed as '10-3' and '10 -4' with inconsistent spacing; they should be formatted as 10^-3 and 10^-4.","section":"4.2.1 / 4.2.3"},{"comment":"The Shapiro-Wilk test is stated as performed, but no test statistic or outcome is reported; with n=3 the test has little power, so the authors should either report the result or remove the sentence.","section":"4.3"},{"comment":"The notation in Eq. (3) and the surrounding text (e.g., xconcatcom(t0) and the indexing over tau = t0-2 ... t0+1) is hard to parse; it should be written out more explicitly.","section":"Eq. (3)"},{"comment":"The table captions should state the number of subjects and clarify that values are per-subject means with SDs, and whether EDV/ESV are body-surface-area indexed.","section":"Tables 1 and 2"},{"comment":"The 'clinical baseline' for the prospective study is a breath-hold ECG-gated segmented cine from a separate acquisition with no in-plane acceleration, and the RT sequence is GRE rather than bSSFP; a sentence stating that this cross-sequence comparison is inherent to the study design would improve clarity.","section":"Abstract / 5.2.2"},{"comment":"The stationarity assumption over R consecutive frames is asserted as approximate, but at R=8 with 61 ms temporal resolution the window is about 0.5 s, over which respiratory chest-wall motion is non-negligible; a brief quantitative discussion of this timescale would strengthen the analytical derivation.","section":"Sec. 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an arXiv v1 preprint, and my assessment is based on that version. The method is plausible and the prospective R=8 evaluation is a genuine strength, but the abstract overstates the quantitative evidence: the only reported quantitative image metrics are LV-function tables with n=3 and n=6, and the retrospective table actually shows substantial differences (EF 55.0 vs 65.4). I would ask the editor to require the PSNR/SSIM numbers, a description of sensitivity-map estimation, and a more appropriate statistical analysis (Bland-Altman or equivalence testing) before considering the paper further. If those requirements are met, the paper could become a solid contribution to the real-time MRI literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Things you should know: this is a genuinely new idea in real-time cardiac MRI—explicit DL-based removal of pseudo-periodic ghosting before outer-volume subtraction—and the prospective validation against an independent breath-hold reference is the right experiment. The method is plausible and the paper is mostly clearly argued.\n\nThe core contribution is real: the Sec. 2 decomposition of composite images into an averaged moving component, ghosting, and stationary background, with ghost-aware background estimation, is not in the prior composite-image work (Blaimer, Gómez-Talavera). The OVR-specific training loss in Eq. 8, using masked sensitivity maps during training and full maps at inference, is a sensible fix for the residual-signal problem, and Fig. 4 shows the artifacts it addresses. The authors also deserve credit for using SSDU self-supervision and for explicitly comparing against a matched PD-DL network without OVR.\n\nSoft spots, in order of severity. First, the paper claims quantitative improvement but never reports PSNR/SSIM numbers—only a sentence in Sec. 5.1.2 saying the metrics “support” the improvements. That needs to be fixed. Second, the retrospective LV analysis has only 3 subjects and shows EF 55.0 vs 65.4 and ESV 56.0 vs 42.2 against the R=4 baseline. Those are clinically meaningful differences, and with n=3 no t-test was possible. The abstract’s “quantitatively” wording overreaches. Third, there is some circularity: the ghost-detection network is trained on TGRAPPA(R=4) difference maps, and the retrospective comparison is against TGRAPPA(R=4). The authors acknowledge the baseline is not a true reference, but the connection is still tighter than the text admits. Fourth, the paper never states how sensitivity maps are computed for the OVR-altered k-space data. If they are estimated from the same composite data that contain the ghosts and outer-volume signal, the stress-tester’s concern about a biased encoding operator is legitimate; if they come from a separate calibration or temporal average, that should be stated. This is a missing detail rather than an observed flaw, but it is load-bearing for Eq. 6.\n\nWho is this for: anyone working on accelerated real-time cardiac MRI or PD-DL reconstruction. It deserves a serious referee. I would send it out and ask for a revision. The core method is sound, but the quantitative claims need to match the evidence: add numbers with error bars, increase the retrospective cohort or soften the claim, clarify the sensitivity-map estimation, and discuss the label circularity. With those changes it would be a solid contribution.","headline":"A genuinely new ghost-aware outer-volume removal method for RT cardiac MRI, with a prospective validation design that is better than most in the field, but the quantitative evidence is thinner than the abstract claims.","tokens_in":19086,"tokens_out":2393,"would_cite":true,"duration_ms":23696,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A post-processing pipeline that estimates and subtracts the signal from tissues surrounding the heart lets free-breathing real-time cine MRI be accelerated eightfold with image quality comparable to clinical breath-hold references.","keywords":["real-time MRI","dynamic MRI","cardiac cine","outer volume removal","ghosting artifacts","deep learning reconstruction","parallel imaging","self-supervised learning"],"falsifier":"Acquire a prospective R=8 free-breathing cine dataset from subjects with pronounced respiratory or diaphragmatic motion and compare the pipeline's systolic-frame reconstruction to a breath-hold reference at the same cardiac phase; visible ghost energy from other cardiac phases inside the myocardium, or an LV ejection fraction that drifts outside the agreement range reported in the prospective cohort, would falsify the central decomposition.","tokens_in":18114,"feed_emoji":"🫀","tokens_out":8335,"duration_ms":82012,"temperature":0.7,"pith_summary":"This paper aims to show that a purely reconstruction-side step—estimating and subtracting the signal from tissues outside the heart—removes the main obstacle to high acceleration in real-time cine cardiac MRI. It analyzes low-temporal-resolution composite images formed from time-interleaved shifted undersampling as a sum of a motion-averaged heart, a pseudo-periodic ghost from cardiac motion, and a stationary background, then trains a neural network to detect the ghost and a second network to locate the heart. After subtracting the cleaned background estimate from k-space, a physics-driven unrolled network trained with a self-supervised, OVR-specific loss restores the image. The claimed payoff is that free-breathing real-time cine at R=8 acceleration matches the visual quality and left-ventricular function measurements of R=4 or breath-hold clinical references, with no change to the acquisition protocol.","feed_headline":"Outer-volume subtraction makes 8x real-time heart MRI viable","feed_subtitle":"Free-breathing cardiac scans at 8x undersampling match breath-hold reference quality without acquisition changes.","key_machinery":"The governing identity is the composite-image decomposition $x_{com}(t_0) = x_{moving}(t_0) + x_{ghost}(t_0) + x_{background}(t_0)$, which turns the outer-volume-removal problem into three tractable subproblems. The load-bearing mechanism is a ResNet ghost detector that consumes four adjacent composite images and outputs the ghost maps for those frames; after ghost subtraction the stationary background is masked by a U-Net-predicted heart mask and subtracted from each frame's k-space. The final component is the OVR-specific SSDU loss for the 35-unroll PD-DL network, which mixes a standard self-supervised data-consistency loss with a consistency term between reconstructions using full and ROI-masked coil sensitivities, so the network learns to keep ROI signal while pushing residual outer-volume signal outside the mask.","core_discovery":"On the paper's own terms, the central claim is that outer-volume aliasing—not temporal resolution—is the limiting factor at high acceleration and that removing the outer volume in k-space unlocks R=8 real-time cine. The authors characterize composite cine images as $x_{com} = x_{moving} + x_{ghost} + x_{background}$, where time-interleaved shifted sampling gives each foldover a distinct modulation phase; the moving components add constructively at the true heart location, stationary components cancel at side foldovers, and the moving components form a pseudo-periodic ghost. A ResNet estimates the ghost, the background is obtained by subtracting it, an OVR mask isolates the heart, and k-space subtraction $y_{OVR} = y - F(m_{OVR} \\cdot x_{background})$ removes the outer-volume signal before an unrolled physics-driven network performs frame-by-frame reconstruction. The PD-DL network is trained self-supervised with a loss that adds a consistency term between full-sensitivity and ROI-masked-sensitivity reconstructions, preventing both signal loss in the ROI and artifacts from residual outer-volume signal. Across retrospective bSSFP data and prospective GRE data at R=8, the method is reported to produce images visually comparable to clinical references and LV function values with no statistically significant difference from breath-hold cine in the prospective cohort.","pith_inferences":["If the decomposition holds, the same outer-volume-removal recipe should transfer to other dynamic MRI settings in which a large stationary FOV aliases into the ROI; for non-Cartesian trajectories the ghost structure would need its own derivation, a point the paper leaves open.","The two-stage pipeline (ghost detection then reconstruction) is likely not the end of the road: end-to-end training that backpropagates the reconstruction loss through the ghost detector could refine the background estimate, which the paper notes as future work.","A plausible next test is pushing to R=10 or higher; if outer-volume aliasing is genuinely eliminated, the remaining limit should be g-factor noise inside the ROI, and combining OVR with virtual coils or spatiotemporal regularization might extend the regime further.","The method's generalization boundaries (different scanner vendors, field strengths beyond 3T, arrhythmic patients, pediatric sizes) are not established in the paper, so those are natural stress tests before clinical adoption."],"forward_implications":["At R=8, the OVR pipeline is claimed to yield images visually comparable to the clinical reference, whereas TGRAPPA at R=8 is severely artifact-degraded and PD-DL without OVR blurs the myocardium and papillary muscles.","Left-ventricular volumes, ejection fraction, stroke volume, and mass measured from OVR-based R=8 reconstructions agree with breath-hold segmented cine to the point of no statistically significant difference in the prospective study.","Because OVR acts in post-processing, it can be added to existing time-interleaved acquisitions without outer-volume-suppression pulses, avoiding SAR, steady-state disruption, and signal regrowth.","Ghost detection trained on R=4 data transfers to prospectively accelerated R=8 data, meaning no fully sampled reference at the target acceleration is needed to build the pipeline.","The comparisons are deliberately run without temporal regularization, so the reported improvements are attributed to outer-volume removal rather than temporal blurring."],"supporting_citations":[{"why":"Supplies the TGRAPPA time-interleaved shifted undersampling pattern, the composite-image formation principle, the R=4 baseline in retrospective experiments, and the g-factor-noisy labels used to train the ghost detector.","marker":"Breuer et al., 2005"},{"why":"Establishes the TSENSE time-interleaved acquisition that the analytical ghosting characterization builds on.","marker":"Kellman et al., 2001"},{"why":"Provides the SSDU self-supervised training paradigm that lets the PD-DL network train without fully sampled reference data.","marker":"Yaman et al., 2020a"},{"why":"Introduces the multi-mask SSDU loss used to train both the masked-sensitivity network and the final OVR PD-DL network.","marker":"Yaman et al., 2022"},{"why":"Supplies the MoDL-style unrolled network architecture whose proximal-operator ResNets and data-fidelity blocks form the PD-DL backbone.","marker":"Aggarwal et al., 2018"},{"why":"Provides the U-Net architecture trained to predict the heart-boundary mask used in the k-space subtraction.","marker":"Ronneberger et al., 2015"},{"why":"Supplies the SENSE formulation in which masked coil sensitivities are interpreted as inverting a lower-rank system, motivating the full-versus-masked sensitivity consistency loss.","marker":"Pruessmann et al., 1999"},{"why":"Supplies the Segment software and its DL segmentation tool used to quantify left-ventricular function in both retrospective and prospective analyses.","marker":"Heiberg et al., 2010"}],"fun_headline_variants":["Subtract outer volume for 8x real-time cardiac MRI","Outer-volume removal enables 8x free-breathing cardiac cine","Deep learning subtracts outer volume for 8x real-time heart MRI","OVR via DL yields 8x accelerated real-time cardiac cine","Remove extra-cardiac signal for 8x real-time heart imaging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands or falls on the Section 2 assumption that, over the R consecutive frames used to form a composite image, every frame separates into a moving cardiac component and a stationary background, with respiration and diaphragm motion slow enough to be negligible; if that separation fails, the ghost estimate is wrong and subtracting it from k-space will corrupt the heart region rather than clean it.","fun_headline_variants_meta":{"raw":{"variants":["Subtract outer volume for 8x real-time cardiac MRI","Outer-volume removal enables 8x free-breathing cardiac cine","Deep learning subtracts outer volume for 8x real-time heart MRI","OVR via DL yields 8x accelerated real-time cardiac cine","Remove extra-cardiac signal for 8x real-time heart imaging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000787,"raw_usage":{"total_tokens":3551,"prompt_tokens":1106,"completion_tokens":2445,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":722,"completion_tokens_details":{"reasoning_tokens":2354}},"tokens_in":722,"tokens_out":2445,"duration_ms":16692,"temperature":1.0,"reasoning_tokens":2354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:37:03.298487+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire a prospective R=8 free-breathing cine dataset from subjects with pronounced respiratory or diaphragmatic motion and compare the pipeline's systolic-frame reconstruction to a breath-hold reference at the same cardiac phase; visible ghost energy from other cardiac phases inside the myocardium, or an LV ejection fraction that drifts outside the agreement range reported in the prospective cohort, would falsify the central decomposition.","supporting_citations":[{"cited_title":"A., Kellman, P., Griswold, M","cited_arxiv_id":null,"evidence_quote":"Supplies the TGRAPPA time-interleaved shifted undersampling pattern, the composite-image formation principle, the R=4 baseline in retrospective experiments, and the g-factor-noisy labels used to train the ghost detector."},{"cited_title":"K., Mani, M","cited_arxiv_id":null,"evidence_quote":"Supplies the MoDL-style unrolled network architecture whose proximal-operator ResNets and data-fidelity blocks form the PD-DL backbone."}],"review_version":1}