{"id":"910ecd07-0fb5-4e44-a0bf-15375c2d7a35","arxiv_id":"2501.09305","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A diffusion reconstruction method that conditions on time-resolved k-space and adds x-t and k-t priors reports improved dynamic MRI reconstruction metrics on Cartesian cardiac and radial lung data.","lead":"This paper introduces dDiMo, a diffusion-model method that adds temporal information to MRI reconstruction and tests it on fast cardiac and lung scans. It reports higher image-quality metrics than several comparison methods, though the gains are small at the most aggressive undersampling.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-Cartesian lung results may not measure true fidelity: the 283-spoke GROG reference is generated by the same pipeline that forms dDiMo's training pairs, so its PSNR/SSIM edge over XD-GRASP could reflect matching that pipeline; XD-GRASP's consistently higher Tenengrad suggests the reference may…","rationale":"The central claim has two independent legs: Cartesian cardiac cine and non-Cartesian radial lung. The cardiac leg appears solid: CMRxRecon is a public dataset, the comparisons include L+S, CRNN, DiMo, and the improvements are consistent across 4x/8x/10x for SAX and LAX, with modest standard deviations. The load-bearing uncertainty is entirely in the lung leg, where the ground truth is not an independent acquisition but a reconstruction produced by the same motion-binning and GROG gridding steps that define the network's training target. The reader's weakest assumption identifies this, and I agree. I sharpened it by pointing to a specific observable consequence: XD-GRASP's Tenengrad is higher at every spoke count, and the authors attribute this to noise without validation. In the absence of a gold standard, higher similarity to a possibly over-smoothed GROG reference is not evidence of structural recovery. The 17-spoke row of Table II already breaks the stated claim: dDiMo does not beat XD-GRASP on NMSE or Tenengrad, and its PSNR advantage is 0.01 dB with overlapping standard deviations. The paper's own Discussion (Section V) concedes that k-t priors are unreliable in radial lung data, and the inference time of 6 minutes/volume is a practical limitation but not a correctness issue. The proposed concrete test, recomputing metrics against an independent non-Cartesian reference, would settle whether the lung numbers are an artifact. If it passes, the claim stands; if not, the non-Cartesian superiority should be withdrawn. The verdict remains CONDITIONAL because the method is plausible and the Cartesian results are strong, but the lung claim must be re-validated against an independent reference before acceptance.","tokens_in":23039,"tokens_out":7656,"duration_ms":71765,"concrete_test":"Reconstruct the six lung test motion-phase references without using the dDiMo training pipeline: e.g., apply a non-Cartesian motion-resolved compressed-sensing reconstruction (or XD-GRASP itself with all 1700 spokes) on the native radial trajectory to obtain reference images, then recompute PSNR/SSIM/NMSE/Tenengrad for dDiMo and XD-GRASP against this independent reference on the same 70/35/17-spoke test cases. If dDiMo retains its advantage across all spoke counts, the measurement-substrate concern is resolved; if the ranking changes or the 17-spoke gap persists, the original lung metrics are artifacts of the GROG/motion-binning reference. A secondary check: measure motion-bin purity by registering spokes within each bin to a common respiratory position; if intra-bin displacement exceeds approximately one voxel, the reference is temporally blurred.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"For the non-Cartesian lung experiments (Section II-C, III-B, Table II), the reference 'ground truth' is reconstructed by binning 283 golden-angle radial spokes with a projection-based respiratory signal and gridding to Cartesian with self-calibrating GROG. dDiMo's training pairs are generated by randomly selecting subsets of those same spokes and gridding them with the same GROG operator. The network therefore learns to invert the specific undersampling-plus-GROG mapping that produces the reference. PSNR/SSIM/NMSE are computed in this GROG-gridded Cartesian space, so dDiMo can achieve high scores by reproducing the pipeline's own output, including any smearing introduced by GROG or by motion-bin errors. The claim that dDiMo provides 'improved temporal alignment and structural recovery' in lung imaging is thus only as strong as the validity of this reference. Table II shows XD-GRASP has higher Tenengrad at every acceleration level (0.0309 vs 0.0286 at 70 spokes; 0.0263 vs 0.0189 at 35; 0.0197 vs 0.0113 at 17), which the authors dismiss as noise. But it is equally consistent with the reference being over-smoothed, so dDiMo's higher similarity to a blurred reference is not evidence of better structural recovery. At 17 spokes, dDiMo's PSNR (30.63±1.55) is statistically indistinguishable from XD-GRASP (30.62±1.04), and its NMSE is worse (0.1623 vs 0.1575), directly contradicting the abstract's claim of consistent superiority. No statistical tests, a test set of only 2 lung datasets, and the unresolved reference-validity question mean the non-Cartesian half of the central claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes dDiMo, a diffusion-based reconstruction method for accelerated dynamic MRI. The method extends the authors' prior DiMo framework by processing multi-frame multi-coil k-space sequences jointly, using a 3D U-Net for noise estimation, explicit x-t and k-t prior networks, data consistency, and a nonlinear CG refinement inside each reverse diffusion step. The method is evaluated on CMRxRecon cardiac cine at 4x, 8x, and 10x acceleration and on an in-house free-breathing stack-of-stars lung MRI dataset at 70, 35, and 17 spokes per respiratory motion bin, with comparisons to L+S, CRNN, DiMo, XD-GRASP, and zero-filled reconstruction. The main claim is that dDiMo provides improved temporal alignment and structural recovery over these methods for both Cartesian and non-Cartesian dynamic MRI.","tokens_in":23387,"tokens_out":8655,"duration_ms":81644,"significance":"The methodological integration is a reasonable and potentially useful extension of diffusion models to dynamic MRI, and the use of the public CMRxRecon benchmark and several independent baselines is a strength. If the results hold, the method could be valuable for high-acceleration cardiac and free-breathing lung imaging. The manuscript also includes helpful ablations for the weighting factors, diffusion steps, and the CG module. However, the lung evaluation is currently insufficient to support the central claim: the reference is generated from the same binning/gridding pipeline used to create training pairs, the test set contains only two subjects, and at 17 spokes the method does not beat XD-GRASP on NMSE. The paper would be strengthened by a disentangled evaluation of the non-Cartesian pipeline and by proper statistical testing.","major_comments":[{"comment":"The non-Cartesian lung evaluation has a reference/training coupling problem. The reference for each motion state is formed by binning 283 golden-angle radial spokes with a projection-based respiratory signal and gridding them to Cartesian with self-calibrating GROG; the training pairs are then generated by randomly selecting subsets of those same binned spokes and gridding them with the same GROG operator (Fig. 3). All PSNR/SSIM/NMSE numbers are computed in this GROG-gridded Cartesian space. dDiMo is therefore trained to invert exactly the undersampling-plus-GROG mapping that produces the reference, and its scores can reflect fidelity to that pipeline rather than to true anatomy. XD-GRASP, by contrast, reconstructs directly from radial k-space and is not matched to the GROG operator. The authors' explanation that XD-GRASP's consistently higher Tenengrad is due to noise is plausible but not the only possible reading; the reference itself may be over-smoothed. A concrete test is needed, e.g., an independent reference reconstructed from all 1700 spokes or a retrospective simulation with a known ground-truth sequence, and XD-GRASP results should also be reported after the same GROG gridding so that the comparison is not biased by the training/reference pipeline.","section":"Section II-C, Section III-B, Table II"},{"comment":"At 17 spokes, dDiMo does not outperform XD-GRASP on key quantitative metrics: NMSE is worse (0.1623 ± 0.0492 vs 0.1575 ± 0.0226), PSNR is equal within uncertainty (30.63 ± 1.55 vs 30.62 ± 1.04), and Tenengrad is lower at every spoke count (e.g., 0.0113 vs 0.0197 at 17 spokes). This contradicts the abstract's sweeping claim of 'improved temporal alignment and structural recovery' and the text's statement that dDiMo 'consistently achieves the highest PSNR and overall image similarity' across all undersampling levels. The claim needs to be revised or supported with paired statistical tests; no significance testing is reported anywhere in the manuscript.","section":"Table II, Section IV-B"},{"comment":"The lung test set consists of only two subjects, and no per-subject results are reported. With n = 2, the means and standard deviations in Table II and the violin plots in Figure 8 cannot support a claim of consistent, generalizable superiority over XD-GRASP; the results should be presented as a pilot or the test set should be expanded. At minimum, per-subject metrics and confidence intervals are needed.","section":"Section III-B, Section IV-B"},{"comment":"The data-consistency update in Algorithm 1 (line 7) and Algorithm 2 (line 2) uses a coefficient λ_t that is never defined in the text; the forward and reverse diffusion formulas in Section II-B use only β_t, α_t, and σ_t. The authors should specify the schedule or value of λ_t; without this, the method cannot be reimplemented exactly.","section":"Algorithm 1, Algorithm 2"}],"minor_comments":[{"comment":"The text says 'F represents a Fourier transform applied to the estimated clean k-space data to convert it into x-t space, and F^H denotes the inverse Fourier transform operation.' This is opposite to the convention in Eqs. (1)-(2), where F maps to k-space and F^H maps from k-space. The equation itself is consistent, but the prose should be corrected.","section":"Section II-B.2, Eq. (16)"},{"comment":"The caption states that results are shown for 4x, 8x, and 16x undersampling, but the text and Table I consistently report 10x as the highest acceleration. This should be corrected to 10x.","section":"Figure 4 caption"},{"comment":"The number of motion bins, the sliding window overlap, and the number of respiratory phases used for training are not specified; the evaluation uses 6 motion states, but the training details are needed for reproducibility.","section":"Section II-C, Section III-B"},{"comment":"The statement that the golden-angle radial lung results rely more on the x-t component than the k-t component is not supported by a lung-specific ablation; the only lambda ablation shown in Figure S3 is for cardiac cine. A lung ablation or a qualifying statement is needed.","section":"Section V"},{"comment":"No code or data availability statement is included; providing one would improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper relies heavily on the authors' own DiMo and prior k-space work, and the novelty over DiMo is incremental. The main weakness is the lung evaluation's reference/training coupling and the very small test set; if the authors can address these with additional experiments and statistical tests, the work could be publishable in a medical imaging journal. The paper also lacks a defined λ_t in the algorithms, which is a reproducibility issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper extends the authors' own DiMo diffusion framework to dynamic MRI by adding temporal processing (3D U-Nets), x-t and k-t priors, and a nonlinear CG refinement layer. On the Cartesian cardiac side, the results look believable: dDiMo beats L+S, DiMo, and CRNN on PSNR/SSIM/NMSE/Tenengrad at 4x, 8x, and 10x, the gains are consistent across SAX and LAX views, and the public CMRxRecon dataset gives the evaluation a solid footing. The ablation study is reasonable and the authors are upfront about the 6-minute inference time. They also deserve credit for reporting Tenengrad even though it goes against them in the lung experiments.\n\nThe soft spots are concentrated in the non-Cartesian lung half, and they are substantive. The reference for training and evaluation is not a fully sampled ground truth; it is built by binning 283 radial spokes with a respiratory signal and gridding them with self-calibrating GROG. dDiMo's training pairs are generated by randomly subsetting those same spokes and gridding with the same operator, so the network learns to invert the specific undersampling-plus-GROG mapping that produced the reference. PSNR/SSIM/NMSE are then computed in that same GROG-gridded space. XD-GRASP, by contrast, reconstructs directly from radial k-space and is evaluated against that GROG reference, which biases the comparison. The consistently lower Tenengrad of dDiMo at every lung acceleration level is consistent with over-smoothing rather than with the authors' 'noise inflates XD-GRASP' explanation. At 17 spokes the numbers flatly contradict the abstract: dDiMo's PSNR is a tie (30.63 vs 30.62) and its NMSE is worse (0.1623 vs 0.1575). With only 2 test subjects and no significance testing, the lung half of the central claim is not established.\n\nThe cardiac results alone are enough to justify a serious look, but the manuscript overstates the lung claim. The authors should release code, add statistical testing, compare against recent dynamic diffusion baselines they currently omit, and either fix the 17-spoke narrative or qualify it. The GROG-reference issue needs a direct response; ideally the lung reference should be rebuilt from more spokes or evaluated on a separate fully sampled acquisition. This is an incremental but honest piece of work, and I would not desk reject it. Send it to reviewers who know radial MRI and diffusion, and expect major revisions.","headline":"The cardiac experiments are credible and the temporal-diffusion idea is a real, if incremental, step; the lung results are not established because the reference is generated by the same GROG pipeline the network is trained to invert.","tokens_in":24031,"tokens_out":1911,"would_cite":true,"duration_ms":22723,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Dynamic MRI reconstruction improves when diffusion denoises the whole frame sequence with temporal guidance, not each frame alone.","keywords":["dynamic MRI reconstruction","diffusion model","temporal priors","k-t space","cardiac cine","radial lung MRI","non-Cartesian reconstruction","conjugate gradient"],"falsifier":"Simulate a radial acquisition with known ground truth and exact spoke-to-state assignments, run dDiMo and the compressed-sensing baseline through the same GROG-binned pipeline, and check whether dDiMo's reported margins persist; if the margin shrinks or reverses when the reference is exact rather than bin-derived, the claimed superiority depends on the approximation in gridding and motion binning, not on the temporal priors themselves.","tokens_in":22760,"feed_emoji":"🧲","tokens_out":13467,"duration_ms":130528,"temperature":0.7,"pith_summary":"Diffusion models have been used to reconstruct static MRI by denoising k-space, but a beating heart or breathing lung moves between frames. The paper tries to show that this temporal structure should be part of the diffusion process itself: instead of reconstructing each time frame independently, dDiMo denoises several adjacent frames together and guides every reverse step with learned spatiotemporal and frequency-temporal priors plus a nonlinear conjugate gradient refinement. If the claim holds, accelerated dynamic scans could recover both spatial detail and motion fidelity at higher undersampling than frame-wise methods can manage, which matters because cardiac and respiratory motion are exactly where fast MRI currently blurs or misaligns. The method is tested on Cartesian cardiac cine at 4x, 8x, and 10x acceleration and on non-Cartesian radial lung data at 70, 35, and 17 spokes per motion state.","feed_headline":"Temporal-aware diffusion sharpens accelerated heart and lung MRI","feed_subtitle":"Joining image-domain and frequency-domain temporal priors keeps motion aligned at up to 10x acceleration and 17 radial spokes.","key_machinery":"The load-bearing object is the temporally guided reverse diffusion step: the update that mixes the current noisy state with a temporally refined clean estimate to produce the next state. Three learned modules feed that step. A 3D noise-estimation U-Net sees several time frames at once, so predicted noise carries inter-frame context. A 3D spatiotemporal network acts on the estimated clean sequence in image space to sharpen temporal dynamics, and a 3D frequency-temporal network, trained in a self-consistency fashion on the auto-calibration region of k-space, enforces consistency with measured data; a nonlinear conjugate gradient layer with a temporal finite-difference penalty closes the step. The role of this machinery is to make temporal coherence a per-step constraint of the diffusion process rather than a post-processing afterthought.","core_discovery":"The central claim is that for dynamic MRI the object of diffusion should be the whole frame sequence, not a single image. The reverse process starts from noise, and at each step a 3D noise-estimation network predicts noise using both spatial and temporal context; a clean estimate is pulled out of the noisy state, refined by an image-domain temporal network and by a frequency-domain self-consistency network trained on the auto-calibration region, and then polished by a nonlinear conjugate gradient layer that enforces temporal sparsity. The update then carries this temporally refined estimate into the next reverse step, so temporal coherence is enforced inside the diffusion loop rather than applied afterward. On Cartesian cardiac cine the paper reports consistent gains over low-rank-plus-sparse, frame-wise diffusion, and recurrent-network baselines at all tested accelerations, and on gridded radial lung data it reports gains over the motion-resolved compressed-sensing baseline at all tested spoke counts, with better temporal alignment and lower residual error.","pith_inferences":["A testable extension the paper does not run is a native non-Cartesian diffusion baseline; comparing dDiMo with such a method would separate the benefit of the temporal priors from the benefit of converting radial data into a Cartesian grid.","The authors down-weight the frequency-temporal prior for radial lung data because the densely sampled k-space center makes auto-calibration learning unreliable; a trajectory-aware version that accounts for radial sampling density is a natural next step for very few spokes.","If temporal coherence is the active ingredient, the same sequence-denoiser design should transfer to other time-resolved reconstructions such as perfusion or contrast-dynamics imaging, where adjacent frames share structure and motion; the paper does not test these.","The reported performance saturates beyond optimal prior weights, so the practical gains depend on per-dataset tuning; automatic selection of those weights would be needed for reliable black-box use in the clinic."],"forward_implications":["Temporal guidance converts diffusion-based dynamic MRI reconstruction from a frame-by-frame denoiser into a sequence denoiser, so motion coherence is carried across adjacent frames by the same k-space conditioning machinery.","At every tested cardiac acceleration factor and radial spoke count, dDiMo is reported to have the best PSNR, SSIM, and NMSE among the compared methods, indicating the benefit persists as undersampling becomes more aggressive.","The ablations show that the spatiotemporal and frequency-temporal prior weights and the conjugate gradient temporal penalty each affect output quality, with quality rising up to an optimum and degrading beyond it, so each component is load-bearing rather than decorative.","Because radial data is gridded onto a Cartesian grid before entering the diffusion model, the same trained-in-k-space architecture transfers between Cartesian and non-Cartesian acquisitions without designing a new denoiser.","At 1000 diffusion steps, inference takes about six minutes per lung volume, so practical clinical use would require fewer steps or a latent-space diffusion formulation."],"supporting_citations":[{"why":"Supplies the base domain-conditioned diffusion method that dDiMo extends from static imaging to dynamic MRI, and serves as one of the comparison baselines.","marker":"[47]"},{"why":"Provides the approach for extracting a clean estimate from the noisy diffusion state, which the spatiotemporal prior branch operates on at each reverse step.","marker":"[48]"},{"why":"Contributes the nonlinear conjugate gradient optimization and compressed-sensing sparsity ideas used in the CG refinement layer.","marker":"[4]"},{"why":"L+S is the classical low-rank-plus-sparse baseline against which dDiMo is compared on cardiac cine.","marker":"[10]"},{"why":"CRNN is the recurrent-network baseline that already uses temporal information, making it a strong comparison for the temporal-guidance claim.","marker":"[25]"},{"why":"XD-GRASP is the motion-resolved radial reconstruction baseline for the lung experiments and supplies the projection-based motion-state framing.","marker":"[52]"},{"why":"Self-calibrating GROG gridding maps the non-Cartesian radial lung data onto a Cartesian grid so the same diffusion machinery can be used.","marker":"[56]"},{"why":"Supplies the multi-coil cardiac cine k-space dataset used for the Cartesian training and evaluation experiments.","marker":"[58]"},{"why":"Coil compression reduces the acquired coil arrays to ten virtual coils, making multi-coil training and inference computationally tractable.","marker":"[59]"}],"fun_headline_variants":["Whole-sequence diffusion sharpens accelerated cardiac and lung MRI","Temporal-guided diffusion keeps motion aligned in fast MRI","Diffusion over space-time sharpens accelerated MRI scans","Sequence-aware diffusion improves dynamic MRI reconstruction","Temporal diffusion priors boost acceleration in cardiac and lung MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"For the non-Cartesian lung experiments, the training pairs and the reported metrics rest on the assumption that gridding radial k-space onto a Cartesian grid with self-calibrating GROG (a calibration-based regridding) and binning spokes into respiratory states by the projection-based motion signal preserve the true image content; if gridding is lossy or spokes go to the wrong motion state, the quantitative gains are measured against a distorted reference rather than the true anatomy.","fun_headline_variants_meta":{"raw":{"variants":["Whole-sequence diffusion sharpens accelerated cardiac and lung MRI","Temporal-guided diffusion keeps motion aligned in fast MRI","Diffusion over space-time sharpens accelerated MRI scans","Sequence-aware diffusion improves dynamic MRI reconstruction","Temporal diffusion priors boost acceleration in cardiac and lung MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000612,"raw_usage":{"total_tokens":2872,"prompt_tokens":995,"completion_tokens":1877,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":611,"completion_tokens_details":{"reasoning_tokens":1802}},"tokens_in":611,"tokens_out":1877,"duration_ms":14994,"temperature":1.0,"reasoning_tokens":1802,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:05:58.137985+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a radial acquisition with known ground truth and exact spoke-to-state assignments, run dDiMo and the compressed-sensing baseline through the same GROG-binned pipeline, and check whether dDiMo's reported margins persist; if the margin shrinks or reverses when the reference is exact rather than bin-derived, the claimed superiority depends on the approximation in gridding and motion binning, not on the temporal priors themselves.","supporting_citations":[{"cited_title":"Diffusion modeling with domain-conditioned prior guidance for accelerated mri and qmri reconstruction,","cited_arxiv_id":null,"evidence_quote":"Supplies the base domain-conditioned diffusion method that dDiMo extends from static imaging to dynamic MRI, and serves as one of the comparison baselines."},{"cited_title":"Diffbir: Toward blind image restoration with generative diffusion prior,","cited_arxiv_id":null,"evidence_quote":"Provides the approach for extracting a clean estimate from the noisy diffusion state, which the spatiotemporal prior branch operates on at each reverse step."},{"cited_title":"Sparse mri: The application of compressed sensing for rapid mr imaging,","cited_arxiv_id":null,"evidence_quote":"Contributes the nonlinear conjugate gradient optimization and compressed-sensing sparsity ideas used in the CG refinement layer."},{"cited_title":"Low-rank plus sparse matrix decomposition for accelerated dynamic mri with separation of background and dynamic components,","cited_arxiv_id":null,"evidence_quote":"L+S is the classical low-rank-plus-sparse baseline against which dDiMo is compared on cardiac cine."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"XD-GRASP is the motion-resolved radial reconstruction baseline for the lung experiments and supplies the projection-based motion-state framing."},{"cited_title":"Self- calibrating grappa operator gridding for radial and spiral trajectories,","cited_arxiv_id":null,"evidence_quote":"Self-calibrating GROG gridding maps the non-Cartesian radial lung data onto a Cartesian grid so the same diffusion machinery can be used."},{"cited_title":"Cmrxrecon: A publicly available k-space dataset and benchmark to advance deep learning for cardiac mri,","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-coil cardiac cine k-space dataset used for the Cartesian training and evaluation experiments."}],"review_version":1}