{"id":"e193d35e-dbd5-46f5-90e8-36f7cd4156e3","arxiv_id":"2506.11771","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An AI denoising network embedded in iterative reconstruction reduces 3D IR-UTE knee MRI scan time from 30 minutes to 2.5-10 minutes with preserved image quality.","lead":"Researchers trained a denoising neural network to speed up a specialized MRI scan that images bone in the knee, cutting scan time from 30 minutes to about 5 minutes. The method was tested on healthy volunteers and shows promise for making radiation-free bone imaging practical in clinics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Retrospective-to-prospective transfer for 3D IR-UTE acceleration is assumed, not demonstrated; intra-TR spoke grouping and sampling-pattern differences could invalidate the 5-min prospective validation.","rationale":"I read the paper in good faith and agree with the reader that the central claim is plausible but conditional. The proposed S3MOB method is technically sensible: a plug-and-play Landweber iteration with a custom DnCNN is a reasonable way to regularize undersampled IR-UTE reconstruction, and the paper provides qualitative and quantitative comparisons against CG-SENSE, CG-SENSE-Dn, and LW-dIf on a held-out set. The most load-bearing concern is indeed the unvalidated equivalence between retrospectively undersampled simulations and true prospective acceleration. However, I do not fully endorse the reader's specific mechanism: because the temporal offset pattern repeats every 7 spokes and the acceleration factors 3, 6, and 12 are coprime with 7, the retrospective subsets actually contain a balanced distribution of the seven spoke positions over each 84-spoke cycle, so the average long-T2 cancellation is approximately reproduced. The more precise issue is that the prospective scan acquires the seven spokes within a single inversion pulse, creating a correlated noise structure and a data-consistency coupling that the every-Nth retrospective subset does not replicate, and no quantitative check is provided to show this difference is negligible. The paper's own 'mutual averaging' argument is hand-waving, not a derivation or experiment. Because the prospective validation is a single subject and the 5-min prospective result is central to the abstract and conclusion, this assumption is load-bearing. A concrete computational test using existing data can settle it: construct a prospective-like retrospective subset by selecting whole TR groups and compare reconstructions. If the metrics match the true prospective scan, the concern is resolved; if not, the retrospective evidence cannot support the prospective claim. I therefore keep the reader's CONDITIONAL verdict unchanged and partially agree with the identified weakest assumption.","tokens_in":12953,"tokens_out":11082,"duration_ms":110970,"concrete_test":"Using the V12 reference scan, construct a 'prospective-matching' retrospective dataset by taking all seven spokes from every 6th inversion recovery pulse, yielding 2000 TRs x 7 = 14k spokes and preserving the intra-TR null-crossing cancellation and noise correlation. Reconstruct this dataset with S3MOB and compare SSIM, PSNR, NRMSE, and PSI against (a) the standard every-6th-spoke retrospective subset and (b) the actual prospective 14k scan from V12. If the prospective-matching metrics are substantially closer to the true prospective metrics than the standard retrospective metrics are, then the paper's equivalence assumption in the Discussion is unsupported and the retrospective evaluation cannot stand in for prospective acceleration. This check uses only data already described in Section 2 and would settle whether the grouping effect matters for the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that S3MOB produces good reconstructions at clinically feasible scan times rests on the transferability of retrospectively undersampled evaluations to prospectively accelerated acquisitions. The paper's only prospective evidence is a single subject (V12), and its Discussion (third paragraph) argues that differences between prospective and retrospective signal cancellation are negligible because of averaging over 84k spokes. This argument is not quantitatively demonstrated and is load-bearing: if wrong, the prospective metrics in Table 2 may overstate real-world performance. Concretely, in the prospective 14k-spoke (5-min) acquisition, seven spokes are acquired per IR pulse at tau = 3.8 ms around the null point, so long-T2 signal components change sign across the null and interfere in the reconstruction. The retrospective subsets are formed by selecting every 3rd, 6th, or 12th projection from the 84k reference (Section 2.3). Because 7 is coprime to these factors, the selected subset does contain a balanced distribution of the seven temporal positions over each LCM cycle, so the average cancellation may be roughly preserved. However, the grouping of spokes from the same inversion pulse, the noise correlation structure, and the exact k-space sampling pattern of the prospective scan are different, and no experiment isolates these effects. The paper does not state that the prospective trajectory exactly matches the every-Nth retrospective pattern; if it does not, the DnCNN may operate out of distribution. Thus the equivalence of retrospective and prospective acceleration is an unverified modeling assumption that directly supports the main conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes S3MOB, a Landweber iterative reconstruction with a plug-and-play DnCNN, for accelerating 3D IR-UTE knee MRI. The DnCNN is trained on CG-SENSE reconstructions of retrospectively undersampled data from eight volunteers, paired with CG-SENSE references. The method is evaluated on four held-out volunteers (V9-V12), including one volunteer (V12) with two prospectively accelerated scans, using SSIM, PSNR, NRMSE, and PSI. The authors report improved similarity metrics compared to CG-SENSE and LW-dIf, and improved sharpness compared to CG-SENSE-Dn, especially for a 5-minute acquisition.","tokens_in":13292,"tokens_out":5644,"duration_ms":50211,"significance":"If the reported results hold, the method could make 3D IR-UTE bone imaging clinically feasible within about 5 minutes without ionizing radiation. The use of an open-source Pulseq sequence, the detailed reconstruction pseudo-code, and the held-out evaluation design are strengths. However, the evaluation relies on a single prospective subject and on an untested equivalence between retrospective undersampling and true prospective acceleration; the quantitative tables are also not populated in the submitted text. The claimed improvement therefore remains plausible but not yet fully demonstrated.","major_comments":[{"comment":"The manuscript refers to Table 1 and Table 2 for the quantitative image quality metrics, but the actual numerical values are absent from the submitted text; only the captions are present. Since the central claims of improved SSIM/PSNR/NRMSE relative to CG-SENSE and LW-dIf rest on these numbers, the evaluation cannot be verified as submitted. The tables must be populated and the corresponding averages and standard deviations reported.","section":"Section 2.6, Tables 1 and 2"},{"comment":"The equivalence between retrospectively undersampled and prospectively accelerated acquisitions is a load-bearing assumption. As the authors acknowledge, prospective scans acquire seven spokes per inversion pulse around the null point, producing signal cancellation between spokes before and after the null, whereas retrospective subsets are formed by selecting every Nth projection from the combined trajectory and do not reproduce this intra-TR grouping. The claim that 'mutual averaging can be assumed' over 84k spokes is not quantitatively supported, and no experiment isolates the effect of spoke grouping, noise correlation, or the exact k-space sampling pattern. Because the only prospective evidence comes from a single volunteer (V12), this assumption must be tested or explicitly bounded before the 5-minute clinical feasibility claim can be accepted.","section":"Section 2.3 and Discussion, third paragraph"},{"comment":"The prospective evaluation is based on a single volunteer (V12), and the prospectively acquired images are rigidly registered to the reference before computing metrics, while the retrospectively undersampled images are not. Registration improves the similarity metrics but lowers PSI, which the authors attribute to blurring introduced by the registration optimizer. This asymmetric handling makes the prospective-versus-retrospective comparison in Table 2 difficult to interpret, and the single-subject design does not support generalization of the conclusion. At minimum, the unregistered prospective metrics should be reported alongside the registered ones, and the analysis should discuss how registration affects the comparison.","section":"Section 2.6, Table 2"}],"minor_comments":[{"comment":"The text states that 'the network was trained using only magnitude data', but the Figure 2 caption says 'Input and target slices were normalized and phase-shifted prior to training'; please clarify whether phase information is used in training or not.","section":"Section 2.4 and Figure 2"},{"comment":"The phrase 'S3MOB also exhibits significantly higher PSI values' uses the term 'significantly' without any statistical test; either provide significance testing or rephrase to 'higher' to avoid implying formal inference.","section":"Section 4, first paragraph"},{"comment":"The training duration for R=12 is listed as '45 (R = 12) hours' with the unit 'hours' missing; this should read '45 hours (R = 12)'.","section":"Section 3, first paragraph"},{"comment":"The text refers to a '2D rectangular region-of-interest (ROI)' and then to a 'volume-of-interest (VOI)' without clearly defining how the VOI is derived from the 2D ROIs; please specify the relationship.","section":"Section 2.6"},{"comment":"The abstract says 'one prospectively accelerated scan' but Section 2.2 describes two prospective acquisitions (7k and 14k projections) for volunteer V12; please reconcile this inconsistency.","section":"Abstract and Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The missing numeric tables are a serious barrier to verification, and the single-subject prospective validation together with the untested retrospective-to-prospective transfer assumption are, in my view, load-bearing for the paper's central claim. The work is otherwise technically sound and within the journal's scope; the authors should be given the opportunity to address these points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a straightforward engineering feasibility study: a custom DnCNN embedded in Landweber iterations to accelerate 3D IR-UTE knee MRI from 30 to 5 minutes. The application to this specific radial sequence with a prospective scan is new, and the authors are unusually transparent about the limits. My main caveat is that we cannot actually see the numbers — Tables 1 and 2 are referenced but not present in the text I reviewed, so the claimed SSIM/PSNR gains over CG-SENSE and LW-dIf are unverifiable from the manuscript.\n\nWhat it does well: the pulse sequence is built on Pulseq (open source), the network is trained on eight subjects with held-out evaluation on four, and the reconstruction details (sigma, iterations, phase handling) are given clearly. The authors also describe a real limitation others often ignore: prospective acquisitions acquire seven spokes per inversion pulse around the null point, so long-T2 signal components change sign and cancel; retrospective undersampling by every Nth projection does not reproduce that exact cancellation. Their counterargument — that averaging over 84k spokes makes the difference negligible — is plausible, and the every-3rd/6th/12th selection does balance the seven temporal positions because 7 is coprime to those factors. Still, they do not quantify the residual mismatch in noise correlation or spoke grouping, and it is a load-bearing assumption for the headline 5-minute claim.\n\nSoft spots: single prospective subject (V12), no code or data availability, no compressed sensing baseline, and the \"good agreement\" language runs ahead of the visible metrics. The registration of prospective images improves similarity but lowers sharpness, which they discuss. These are proportionate concerns for a feasibility study; none of them sink the central idea.\n\nWho should read it: anyone working on accelerated UTE or bone MRI. It is a useful data point that plug-and-play denoising can cut IR-UTE scan time, and the honest limitations section makes it a good model for reporting.\n\nRecommendation: send it to peer review, but the referees should see the actual tables, ask for a comparison or explicit justification against compressed sensing, and request a supplementary analysis of the retrospective-prospective equivalence. It is not a home run, but it is a legitimate contribution and deserves a fair look.","headline":"A solid, honest feasibility study of plug-and-play DnCNN acceleration for 3D IR-UTE knee MRI, but the quantitative tables are missing from the text and the retrospective-to-prospective transfer is plausible rather than proven.","tokens_in":13820,"tokens_out":2835,"would_cite":false,"duration_ms":27727,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an AI-driven iterative reconstruction of undersampled IR-UTE knee MRI can produce bone images that agree closely with a 30-minute reference, with the strongest result at a 5-minute scan time.","keywords":["knee MRI","adiabatic inversion recovery","ultra-short echo time","DnCNN","Landweber iteration","plug-and-play reconstruction","undersampled radial MRI","deep learning reconstruction"],"falsifier":"Run S3MOB on true accelerated 5-minute IR-UTE knee scans from several volunteers and compare SSIM and PSNR against their matching retrospectively undersampled reconstructions: a systematic, artifact-linked drop in the prospective images would falsify the central claim.","tokens_in":12809,"feed_emoji":"🦴","tokens_out":11500,"duration_ms":105595,"temperature":0.7,"pith_summary":"The authors are trying to establish that MR-based bone imaging of the knee does not require a 30-minute scan. They propose a reconstruction method, S3MOB, in which a denoising convolutional neural network trained on this specific sequence regularizes the image inside each Landweber iteration, and they report that it reconstructs 2.5-, 5-, and 10-minute acquisitions in good agreement with the full reference, with the 5-minute case especially favorable. Quantitative evaluation on four volunteers excluded from training and on one prospectively accelerated scan showed higher structural similarity and peak signal-to-noise ratio than CG-SENSE and than the same iterative loop with a generic denoiser, and sharper edges than applying the network only once. If the claim carries to patients, radiation-free bone assessment of the knee becomes practical at scan times compatible with clinical workflow.","feed_headline":"AI reconstruction brings bone MRI of the knee down to 5 minutes","feed_subtitle":"A task-trained denoiser inside a Landweber loop keeps bone contrast, hinting at radiation-free fracture imaging.","key_machinery":"The load-bearing mechanism is a plug-and-play Landweber loop. Starting from the gridding reconstruction $x_0=A^*y$, each of six iterations applies the MRI forward operator $A$ (non-Cartesian gridding plus coil sensitivities), subtracts the measured data, and feeds the residual back through the adjoint $A^*$ as a data-consistency correction; before the network step, the complex phase is stored and the magnitude slice is normalized, denoised by the DnCNN, weighted by $\\sigma=0.15$ against the previous iterate, and rescaled. The custom DnCNN is a residual convolutional network with 482.9k parameters, trained on pairs of CG-SENSE reconstructions of retrospectively undersampled subsets and their full 30-minute references from eight volunteers, so it learns IR-UTE-specific noise suppression. This integration reduces noise without the blur and striation artifacts seen when the same network is applied once to a CG-SENSE image.","core_discovery":"The paper's central claim is that S3MOB, six Landweber iterations with a plug-and-play DnCNN as regularizer, reconstructs undersampled 3D IR-UTE knee data in good agreement with the 30-minute reference dataset. On four held-out volunteers, the method achieved higher SSIM and PSNR than CG-SENSE and LW-dIf, with a slight reduction in sharpness relative to those baselines, and higher perceptual sharpness than CG-SENSE-Dn. For volunteer V12, a true 5-minute accelerated scan reconstructed with S3MOB compared favorably with its retrospectively undersampled counterpart and showed no visible streaking. The authors conclude that the method preserves contrast and structural detail while suppressing noise, and that it is poised to make MR-based bone assessment possible in clinically feasible scan times.","pith_inferences":["Beyond the paper, the equivalence between prospective and retrospective acceleration could be tested directly by varying the number of spokes acquired per inversion pulse, the parameter the authors identify as the main source of discrepancy.","If the 5-minute protocol holds in patients, the same reconstruction pipeline offers a path toward replacing CT for selected knee indications, since IR-UTE-derived bone measures already correlate with bone density.","A variational-network version, with the denoiser unrolled through the iterations and the weighting learned per step, is the natural follow-up and would plausibly recover some of the sharpness that the fixed weight of 0.15 sacrifices."],"forward_implications":["A 5-minute knee acquisition reconstructed with S3MOB is the operating point the authors recommend, since it gives high similarity to the 30-minute reference while keeping sharpness above the one-shot denoiser baseline.","The same fixed settings, six iterations and a denoising weight of 0.15, performed consistently across 10-, 5-, and 2.5-minute protocols, so the pipeline does not require per-scan parameter tuning.","The comparison with LW-dIf shows that a denoiser trained on IR-UTE data is a necessary ingredient: swapping in a generic pretrained denoiser lowers SSIM and PSNR, so the learned denoiser, not the iterative loop alone, carries much of the quality gain.","In the one prospectively accelerated scan, the 5-minute image showed no visible streaking, suggesting the method can survive real acceleration rather than only simulated undersampling; the 2.5-minute image stayed usable but showed slight blurring and streaking."],"supporting_citations":[{"why":"supplies the 3D IR-UTE-Cones pulse sequence design, including the seven spokes per inversion pulse and the inversion time, on which all acquisitions are based.","marker":"[8]"},{"why":"justifies the multi-spoke inversion-recovery timing around the null point and the sign-cancellation model used to discuss prospective acceleration.","marker":"[35]"},{"why":"provides CG-SENSE, which generates both the paired training inputs and the baseline reconstructions.","marker":"[40]"},{"why":"provides the coil-sensitivity estimation used inside the MRI forward operator.","marker":"[41]"},{"why":"provides the DnCNN architecture that the method embeds as a learned denoiser.","marker":"[42]"},{"why":"is the Landweber update formula at the core of the iterative reconstruction.","marker":"[46]"},{"why":"supplies the plug-and-play methodology for combining physical data consistency with a learned regularizer.","marker":"[47]"}],"fun_headline_variants":["AI reconstruction speeds knee bone MRI to 5 minutes","5-minute knee bone MRI via AI denoising","Fast bone MRI: AI recovers knee detail in 5 min","AI-driven MRI method cuts bone scan to 5 minutes","Knee bone MRI in 5 minutes with AI-based reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that retrospectively thinning a fully sampled 30-minute acquisition reproduces a genuinely accelerated scan, even though the accelerated sequence samples seven spokes around the inversion null while the retrospective subset does not.","fun_headline_variants_meta":{"raw":{"variants":["AI reconstruction speeds knee bone MRI to 5 minutes","5-minute knee bone MRI via AI denoising","Fast bone MRI: AI recovers knee detail in 5 min","AI-driven MRI method cuts bone scan to 5 minutes","Knee bone MRI in 5 minutes with AI-based reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1242,"prompt_tokens":1046,"completion_tokens":196,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":662,"completion_tokens_details":{"reasoning_tokens":114}},"tokens_in":662,"tokens_out":196,"duration_ms":2645,"temperature":1.0,"reasoning_tokens":114,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:03:29.475466+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run S3MOB on true accelerated 5-minute IR-UTE knee scans from several volunteers and compare SSIM and PSNR against their matching retrospectively undersampled reconstructions: a systematic, artifact-linked drop in the prospective images would falsify the central claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the 3D IR-UTE-Cones pulse sequence design, including the seven spokes per inversion pulse and the inversion time, on which all acquisitions are based."},{"cited_title":"Carl, G.M","cited_arxiv_id":null,"evidence_quote":"justifies the multi-spoke inversion-recovery timing around the null point and the sign-cancellation model used to discuss prospective acceleration."},{"cited_title":"Landweber, An Iteration Formula for Fredholm Integral Equations of the First Kind, American Journal of Mathematics 73 (1951) 615","cited_arxiv_id":null,"evidence_quote":"is the Landweber update formula at the core of the iterative reconstruction."}],"review_version":1}