{"id":"77898824-d664-49b8-8c97-c69c96f4e737","arxiv_id":"2501.03293","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A diffusion model trained on single-slice MRI can reconstruct simultaneous multislice k-space data by enforcing a Slice GRAPPA consistency constraint during sampling, avoiding SMS training data.","lead":"Researchers combined a k-space diffusion model, trained only on ordinary single-slice MRI brain scans, with a standard SMS separation technique called Slice GRAPPA. This lets the model reconstruct simultaneous multislice (SMS) images from accelerated scans without needing SMS data for training, and the authors report improved quality and higher acceleration factors than traditional SMS methods.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never states a held-out test split; if the same fastMRI T2 volumes used to synthesize SMS test data also appeared in diffusion training, Table 1 gains could be memorization artifacts.","rationale":"The reader's weakest assumption was that simulated SMS data faithfully represents real acquisitions. That is a serious transferability concern, but the more immediately load-bearing evaluation flaw is the apparent absence of a documented train/test split. Both training and test construction draw on the same FastMRI T2 brain dataset, and the paper never states that the test SMS volumes were excluded from training. Without that exclusion, the diffusion prior could have memorized or partially memorized the exact test images, making the comparison against non-learning baselines unfair regardless of how realistic the SMS simulation is. The paper's own Discussion also concedes incomplete high-frequency recovery and promises future comparisons, which further underscores that the current evidence is conditional. I therefore keep the reader's CONDITIONAL verdict unchanged, while adding the train/test split as an explicit condition for trusting the quantitative claims.","tokens_in":4288,"tokens_out":6350,"duration_ms":134654,"concrete_test":"Request or reconstruct the train/test split: obtain the fastMRI volume or slice identifiers used for training and for SMS simulation; retrain the heat-diffusion model on only the training subset and recompute Table 1 exclusively on held-out SMS simulations that share no patient or volume with training. If the PSNR gaps over Slice GRAPPA+SENSE at 3x and 4x shrink substantially, for example by more than 3-5 dB, the central claim of no-SMS-training superiority is not established. Also report whether any same-subject or same-volume slices overlap between training and test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the method requires no SMS training data yet outperforms Slice GRAPPA+SENSE and SMS-COOKIE depends on generalization to unseen anatomy. Section 3.1 says only that the FastMRI T2 brain dataset was used to train the model and that test SMS data were simulated from the same FastMRI T2 brain dataset, cropped to 320x320; no train/test split, volume identifiers, or patient-level separation is reported. Because the diffusion model h_theta is trained to output the clean k-space z(0) from degraded input, any test image that appeared during training gives the proposed method access to high-frequency content that the linear baselines cannot use. The magnitude of the reported advantage, for example 12.5 dB over Slice GRAPPA+SENSE at 4x, is exactly what training-set memorization or near-duplicate leakage would look like. This is not an accusation of misconduct; it is a missing verification step. If a held-out split was used, it should be reported; if not, the headline comparison is biased and the numerical claims are not reliable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an SMS (simultaneous multislice) MRI reconstruction method that combines a k-space heat-diffusion model (trained only on single-slice FastMRI T2 images) with Slice GRAPPA applied during the reverse-diffusion sampling process. The authors simulate SMS k-space data from cropped 320x320 FastMRI T2 images with MB=3, a 2/3 FOV CAIPIRINHA shift, and 32 ACS lines, and compare against Slice GRAPPA+SENSE and SMS-COOKIE at in-plane acceleration factors 3x and 4x. They report consistently higher PSNR/SSIM and lower NMSE (e.g., PSNR 36.91 dB vs. 24.43 dB at 4x), and further show that their method maintains relatively high quality up to 8x acceleration. The central claim is that the method does not require SMS data for training yet outperforms traditional SMS reconstruction methods and supports higher in-plane acceleration factors.","tokens_in":4488,"tokens_out":2948,"duration_ms":30475,"significance":"If the reported gains hold on held-out data and on real SMS acquisitions, the paper offers a practically important result: a deep-learning SMS reconstruction method that avoids the need to curate diverse SMS training datasets (with variable MB factors and CAIPIRINHA shifts) by training only on single-slice k-space data and injecting the SMS physics through Slice GRAPPA during sampling. The method builds on the authors' prior heat-diffusion framework (Ref. [7]) and their ISMRM abstract (Ref. [6]), and the paper provides quantitative comparisons at multiple acceleration factors. However, the evaluation is currently too weak to establish the central claim: the training and test data appear to come from the same FastMRI T2 dataset without any reported split, the test data are simulated rather than real SMS acquisitions, no error bars or statistical significance are given, and only two traditional baselines are used. The paper also does not release code, trained models, or test-set identifiers, which limits reproducibility. With a proper held-out evaluation and additional baselines, the contribution could be significant for the SMS-MRI community.","major_comments":[{"comment":"The paper never reports a train/test split. Section 3.1 states that the FastMRI T2 brain dataset was used to train the Heat Diffusion model and that SMS test data were simulated from the same FastMRI T2 brain dataset cropped to 320x320, but no volume identifiers, patient-level separation, or number of test slices is given. Since the diffusion model is trained to reconstruct high-frequency k-space content, any test volume that appeared in training (or came from the same patient) could inflate the reported gains, for example the 12.5 dB improvement over Slice GRAPPA+SENSE at 4x. The authors must specify and implement a held-out patient-level split, or retrain on a disjoint set, before the headline quantitative claims can be considered reliable.","section":"Section 3.1, Table 1"},{"comment":"The evaluation uses only simulated SMS data with a single parameter set (MB=3, 2/3 FOV CAIPIRINHA shift, 32 ACS lines, uniform undersampling). No real SMS acquisitions, variable MB/CAIPI patterns, or realistic slice-leakage and noise-correlation effects are considered. The claim that the method outperforms Slice GRAPPA+SENSE and SMS-COOKIE therefore rests on an untested assumption that simulation faithfully captures the SMS forward model. The authors should validate on at least one real SMS dataset or realistic phantom scan, or clearly scope the claim to simulated data.","section":"Section 3.1, Section 4.1"},{"comment":"The method description omits several parameters needed for reproducibility: the number of reverse-diffusion steps, the noise schedule σ(t), the Gaussian kernel parameters G_t, the predictor-corrector settings, the number of Monte Carlo samples, and the details of the SPIRiT initialization step. Without these, experiments cannot be reproduced, and it is unclear whether the comparison to the baselines is performed under matched computational budgets. The authors should provide a complete algorithmic specification or release code.","section":"Section 2.2, Section 2.3"},{"comment":"The quantitative comparison reports single-point PSNR/NMSE/SSIM values without error bars, confidence intervals, or the number of test volumes/slices. The extremely large differences may be statistically meaningful, but the magnitude of noise is unknown. The authors should report mean and standard deviation over multiple test scans, and ideally per-volume results, to support the claimed superiority.","section":"Table 1, Section 4.1"},{"comment":"The paper compares only two traditional non-deep-learning baselines (Slice GRAPPA+SENSE and SMS-COOKIE). Because the proposed method is a generative deep-learning approach, the absence of any SMS-aware deep-learning baseline (e.g., SMS-RAKI, VCC-RAKI, or a diffusion-based SMS method) makes the claim of significant advantage over existing state-of-the-art overly broad. Adding at least one such baseline is important for positioning the contribution.","section":"Section 4.3"}],"minor_comments":[{"comment":"There are typographical errors: \"SNESE\" in Section 4.1, \"k-sapce\" in the Introduction, and inconsistent spacing after periods (e.g., in the abstract). These should be corrected.","section":"Throughout"},{"comment":"The score-matching loss in Eq. (5) uses S^*G(t)⊙(h_θ(ẑ(t),t) - ẑ(0)), but the relationship between the network output and the score ∇ẑ log p_t(ẑ) is not stated explicitly. A brief derivation or reference would help the reader connect Eq. (5) to the reverse SDE in Eq. (4).","section":"Section 2.2, Eq. (5)"},{"comment":"Table 2 reports results up to 8x, but the text in Section 4.2 says the method \"does not exhibit undersampling artifacts even at 8x,\" while the NMSE worsens to 0.0305 and PSNR drops to 30.5 dB. The wording overstates the visual quality; consider a more nuanced claim consistent with the admitted high-frequency loss.","section":"Section 4.2, Table 2"},{"comment":"Figure 2 and Figure 3 are referenced in the text, but the caption and figure content are not fully described (e.g., which slices, which coil configuration, and the display window). Including slice numbers and error-map color scales would improve interpretability.","section":"Figures"}],"recommendation":"major_revision","confidential_remarks":"The paper reads like a 4-page conference abstract rather than a full journal article. The central idea is reasonable, but the evaluation is not at the standard expected for a journal: no data split, simulated-only data, no statistical uncertainty, and no deep-learning baselines. The main risk is that the headline numbers are inflated by training/test leakage; this must be addressed before the claims can be trusted. I would not reject outright because the problem is fixable, but the revision requires substantial new experiments. Also, the paper relies heavily on the authors' own prior heat-diffusion work with limited comparison to independent methods; a clearer statement of novelty relative to Ref. [7] and Ref. [6] would help."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What's actually new here: the authors take a k-space heat diffusion model, train it on single-slice fastMRI data, and then inject the Slice GRAPPA physics as a constraint during reverse diffusion sampling. That lets them avoid SMS training data entirely, which is a real advantage given how variable SMS acquisitions are. The results over Slice GRAPPA+SENSE and SMS-COOKIE are large — 12.5 dB at 4x — and the qualitative claim that aliasing disappears is plausible if the prior is doing what they say.\n\nCredit where due: the combination is not in the prior literature, it is a sensible extension of their own heat diffusion work, and the sampling strategy handles different MB and CAIPIRINHA shifts at inference without retraining. They also honestly note that high-frequency detail degrades at 8x.\n\nBut there is a hole in the evaluation that a referee must probe first. The paper says the model was trained on the fastMRI T2 brain dataset and that test SMS data were simulated from the same fastMRI T2 brain dataset, cropped to 320x320. No train/test split, volume IDs, or patient-level separation is reported. If the same anatomy appeared during training, the diffusion model has an unfair advantage over the linear baselines, and the reported gains would be memorization artifacts. This is not an accusation of misconduct; it is a missing verification step. The authors need to state the split, ideally at the volume level, or the headline comparison is biased.\n\nBeyond that, the evaluation is thin in ways that matter. It is simulated SMS data only, with no real acquisitions, so slice leakage and coil geometry effects are untested. No error bars or repeated runs, so the dB differences between baselines could be noise. The comparison set is two conventional methods; the discussion of H-DSLR and MoDL is a hand-wave and should be backed by numbers or dropped. Some hyperparameters (diffusion steps, noise schedule) are missing, though that is typical for a short paper.\n\nThe central idea holds up — I do not see circularity or a fitted-parameter problem. The method is not a performance; it is a genuine attempt to combine physics and a diffusion prior. The main issue is evaluation hygiene, not the underlying approach.\n\nWho this is for: people working on SMS reconstruction, especially those who want to avoid collecting SMS training data. It deserves a serious referee, but the referee should demand the train/test split and more rigorous comparison before accepting. If the split is clean, this is a nice contribution; if not, the numbers cannot be trusted.","headline":"The paper's combination of a single-slice-trained heat diffusion prior with Slice GRAPPA at sampling is new and plausible, but the evaluation never states a train/test split, so the big numbers are not yet trustworthy.","tokens_in":5045,"tokens_out":2879,"would_cite":false,"duration_ms":30484,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A k-space diffusion model trained only on single-slice images reconstructs simultaneous multislice MRI by folding Slice GRAPPA into the sampling loop, beating conventional SMS methods at 3x and 4x in-plane acceleration.","keywords":["simultaneous multi-slice MRI","k-space diffusion model","heat diffusion","Slice GRAPPA","CAIPIRINHA","MRI reconstruction","in-plane acceleration","generative diffusion prior"],"falsifier":"Acquire genuine SMS k-space data on a scanner with MB=3 and a 2/3 FOV CAIPIRINHA shift, run the single-slice-trained diffusion model with Slice GRAPPA consistency, and compare against Slice GRAPPA+SENSE at 4x in-plane acceleration: if the PSNR advantage is much smaller than the reported 36.91 dB versus 24.43 dB gap, or if visible slice leakage appears in the separated slices, the simulated-data assumption is the likely cause.","tokens_in":4078,"feed_emoji":"🧠","tokens_out":10715,"duration_ms":92256,"temperature":0.7,"pith_summary":"Simultaneous multislice (SMS) MRI excites several slices at once to cut scan time, but its reconstruction has resisted deep learning because SMS training data vary with multiband factor and CAIPIRINHA shift. The paper argues that an SMS dataset is unnecessary: a k-space diffusion model trained on ordinary single-slice images can serve as the prior, with the SMS physics imposed at sampling time. At every reverse-diffusion step the method recombines per-slice estimates into SMS k-space, enforces data consistency against the measured SMS data, and separates the slices again with Slice GRAPPA. On simulated T2-weighted brain data with MB=3, the method reports PSNR 38.08 dB at 3x and 36.91 dB at 4x in-plane acceleration, outperforming Slice GRAPPA+SENSE and SMS-COOKIE, and it shows no in-plane aliasing up to 8x.","feed_headline":"Diffusion model reconstructs multislice MRI with no SMS training data","feed_subtitle":"It beats Slice GRAPPA+SENSE and SMS-COOKIE at 4x while training only on single-slice k-space.","key_machinery":"The central object is the SMS-constrained reverse diffusion sampler built on a heat-diffusion k-space model, in which the forward process $d\\hat z = \\hat G_t \\odot \\hat z(0) dt + \\sqrt{d\\sigma^2(t)/dt}\\,SS^* dw$ attenuates k-space with a 2-D Gaussian $\\hat G_t$ and coil-sensitivity-weighted noise, and a network $h_\\theta$ trained by score matching learns to invert that attenuation. During sampling, the method takes each slice estimate, composes the multi-slice k-space $\\hat x^{\\mathrm{sms}}_{0|t}$, performs data consistency with the measured $\\hat x^{\\mathrm{sms}}_{00}$, applies Slice GRAPPA kernels estimated from 32 ACS lines to separate the slices again, and uses SPIRiT kernels in initialization, repeating predictor-corrector steps through the reverse SDE.","core_discovery":"The central claim is that SMS-specific training data are not needed for high-quality SMS reconstruction by a generative diffusion model: a prior learned on single-slice k-space is sufficient, provided the SMS forward model is injected during inference. The method's sampling loop alternates heat-diffusion reverse steps with an SMS constraint: single-slice estimates are combined into simulated SMS k-space, corrected against the acquired k-space in a data-consistency step, and decomposed again by Slice GRAPPA. This combined loop removes the inter-slice aliasing that Slice GRAPPA alone leaves behind and resolves the in-plane aliasing that Slice GRAPPA+SENSE and SMS-COOKIE show at 4x, while staying stable through 8x in-plane acceleration with only gradual high-frequency blurring.","pith_inferences":["If the approach transfers to real scanner data, the practical bottleneck for SMS deep learning shifts from building large SMS training sets to obtaining accurate Slice GRAPPA kernels and coil sensitivity maps at scan time.","The same train-on-single-slices / constrain-at-sampling recipe could apply to other k-space undersampling tasks with known forward models but scarce paired training data, such as multi-echo or diffusion-weighted acquisitions.","The rising NMSE at higher acceleration factors identifies high-frequency information as the limiting component, so a frequency-weighted or high-frequency-preserving loss is a natural next experiment."],"forward_implications":["A single diffusion model trained on single-slice k-space can reconstruct SMS data across different multiband and CAIPIRINHA configurations by swapping the Slice GRAPPA and data-consistency constraint, avoiding SMS data collection for training.","SMS reconstruction is no longer limited by the in-plane acceleration ceiling of traditional slice-separation methods; the paper's experiments show usable reconstructions up to 8x without in-plane aliasing.","At 3x and 4x in-plane acceleration, the method's reported PSNR and SSIM exceed both Slice GRAPPA+SENSE and SMS-COOKIE, e.g., PSNR 36.91 dB versus 24.43 dB and 27.19 dB at 4x.","Because the SMS physical model is imposed during sampling rather than learned, adapting the method to a new SMS acquisition mode should require changing the constraint, not retraining the network."],"supporting_citations":[{"why":"Supplies the Slice GRAPPA kernel method used to separate slices in the sampling loop and the CAIPIRINHA shift protocol used in the simulated data.","marker":"[2]"},{"why":"Supplies the SENSE-GRAPPA combination that forms the Slice GRAPPA+SENSE baseline and the SENSE/GRAPPA physics the method builds on.","marker":"[3]"},{"why":"The authors' preceding slice-diffusion method that this work extends by adding a k-space heat-diffusion sampler and Slice GRAPPA constraint.","marker":"[6]"},{"why":"Provides the heat-diffusion k-space forward/reverse SDE, the score-matching loss, and the reverse sampling equation the method trains and samples with.","marker":"[7]"},{"why":"SMS-COOKIE, one of the two traditional SMS reconstruction methods the proposed method is compared against in the experiments.","marker":"[8]"}],"fun_headline_variants":["K-space diffusion model reconstructs SMS MRI without SMS training data","Diffusion model for SMS MRI with no SMS training outperforms standard methods","SMS MRI reconstruction without SMS data: diffusion model wins","Diffusion model beats Slice GRAPPA+SENSE at 4x with no SMS training","No SMS data? Diffusion model still wins at multislice MRI"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that k-space data simulated from single-slice T2-weighted brain images with MB=3, a 2/3 FOV CAIPIRINHA shift, and 32 ACS lines faithfully reproduces real simultaneous multislice acquisitions, including slice leakage, noise correlation, and coil geometry.","fun_headline_variants_meta":{"raw":{"variants":["K-space diffusion model reconstructs SMS MRI without SMS training data","Diffusion model for SMS MRI with no SMS training outperforms standard methods","SMS MRI reconstruction without SMS data: diffusion model wins","Diffusion model beats Slice GRAPPA+SENSE at 4x with no SMS training","No SMS data? Diffusion model still wins at multislice MRI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001365,"raw_usage":{"total_tokens":5468,"prompt_tokens":813,"completion_tokens":4655,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":429,"completion_tokens_details":{"reasoning_tokens":4560}},"tokens_in":429,"tokens_out":4655,"duration_ms":31604,"temperature":1.0,"reasoning_tokens":4560,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:00:37.054323+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire genuine SMS k-space data on a scanner with MB=3 and a 2/3 FOV CAIPIRINHA shift, run the single-slice-trained diffusion model with Slice GRAPPA consistency, and compare against Slice GRAPPA+SENSE at 4x in-plane acceleration: if the PSNR advantage is much smaller than the reported 36.91 dB versus 24.43 dB gap, or if visible slice leakage appears in the separated slices, the simulated-data assumption is the likely cause.","supporting_citations":[{"cited_title":"K-space Diffusion Model Based MR Reconstruction Method for Simultaneous Multislice Imaging","cited_arxiv_id":"2501.03293","evidence_quote":"Supplies the Slice GRAPPA kernel method used to separate slices in the sampling loop and the CAIPIRINHA shift protocol used in the simulated data."},{"cited_title":"Dataset FastMRI Dataset was used to train the Heat Diffusion Model","cited_arxiv_id":null,"evidence_quote":"Supplies the SENSE-GRAPPA combination that forms the Slice GRAPPA+SENSE baseline and the SENSE/GRAPPA physics the method builds on."},{"cited_title":"Simultaneous multislice (sms) imaging techniques,","cited_arxiv_id":null,"evidence_quote":"The authors' preceding slice-diffusion method that this work extends by adding a k-space heat-diffusion sampler and Slice GRAPPA constraint."},{"cited_title":"Blipped-controlled aliasing in parallel imaging for simultaneous multislice echo planar imaging with reduced g-factor penalty,","cited_arxiv_id":null,"evidence_quote":"Provides the heat-diffusion k-space forward/reverse SDE, the score-matching loss, and the reverse sampling equation the method trains and samples with."},{"cited_title":"Accelerated volumetric mri with a sense/grappa combination,","cited_arxiv_id":null,"evidence_quote":"SMS-COOKIE, one of the two traditional SMS reconstruction methods the proposed method is compared against in the experiments."}],"review_version":1}