{"id":"15688edd-7836-4f78-8c3f-27bc009980f0","arxiv_id":"2506.05975","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Real-world paired MRI data plus a new feature-space metric give more reliable evaluation of 3D motion correction than simulated or reference-free approaches, which are shown to mislead.","lead":"This paper releases PMoC3D, a dataset of real-world motion-corrupted 3D brain MRI scans paired with motion-free references, and introduces MoMRISim, a learned metric for scoring motion artifacts. It compares three evaluation approaches and concludes that real-world evaluation with MoMRISim is the most reliable, while simulated motion overestimates algorithm performance.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Validation of MoMRISim relies on a single-annotator PMAS ground truth; without inter-rater agreement or uncertainty quantification, the claimed reliability of real-world evaluation is not established.","rationale":"Agree with the reader's weakest assumption. The paper's central claim is that MoMRISim, combined with PMoC3D, gives reliable evaluation. The evidence for this is the correlation between MoMRISim and PMAS. But PMAS is not a well-established ground truth; in Appendix B, the reconstruction-level PMAS used in Section 4.2 is explicitly from a single evaluator. No inter-rater study is provided, so we cannot tell whether this score reflects a stable human consensus or one person's preference. This is load-bearing because all three comparative findings—real-world vs. simulated, metric quality, and reference-free limitations—are anchored to this PMAS. If the PMAS is unreliable, the correlations (especially MoMRISim rho=0.95) may be inflated by the rater's idiosyncrasies. The dataset and code are publicly available, making a multi-rater validation feasible. The paper has strong points: MoMRISim is trained on simulated severity rather than on PMAS, reducing circularity; the visual examples are compelling; and the dataset is a genuine contribution. But the central reliability claim currently rests on a single-annotator benchmark, so the verdict should remain conditional pending an inter-rater reliability study. No change to the reader's CONDITIONAL verdict is needed.","tokens_in":16112,"tokens_out":6097,"duration_ms":58451,"concrete_test":"Have at least two independent raters (ideally radiologists or trained MRI technicians) repeat the pairwise comparison protocol from Appendix B on the reconstruction-level PMAS set used in Section 4.2, or on a stratified random subset of at least 50 pairs. Compute inter-rater reliability (e.g., Krippendorff's alpha or Fleiss' kappa) and a consensus PMAS. Then recompute the Spearman correlation between MoMRISim and the consensus PMAS, with a bootstrap confidence interval. If the correlation drops materially below 0.8 or inter-rater agreement is poor, the reliability claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that real-world evaluation with PMoC3D and MoMRISim is 'relatively reliable' is validated in Section 4.2 by correlating each metric with the perceived motion artifact score (PMAS). Appendix B states that for the reconstruction-level PMAS used in Section 4.2, 'annotations were performed by a single evaluator.' This is the ground truth against which MoMRISim's Spearman rho=0.95 is computed, yet no inter-rater agreement, rater blinding, or confidence interval is reported. If that single evaluator's notion of artifact severity is idiosyncratic, the high correlation does not demonstrate alignment with human perception generally. The same dependence recurs in the Section 4.3 comparison of simulated vs. real-world reconstructions, where a single 'human annotator' judges residual artifacts. The mild-motion caveat does not fix this: the reliability claim specifically covers moderate-to-severe motion, and the PMAS-based validation is the only quantitative support for it. The small number of subjects (8) and the absence of confidence intervals on the correlations amplify the risk that the headline result is not robust.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the evaluation of 3D MRI motion correction methods. It introduces PMoC3D, a dataset of paired raw k-space motion-corrupted and motion-free scans from 8 subjects, and MoMRISim, a learned feature-space similarity metric trained on triplets labeled by simulated motion severity. The authors compare three evaluation strategies: real-world reference-based evaluation, simulated-motion evaluation, and reference-free evaluation. Their main findings are that (i) real-world evaluation combined with MoMRISim correlates well with human-perceived artifact scores for moderate-to-severe motion, (ii) simulated-motion evaluation systematically overestimates algorithm performance, and (iii) reference-free metrics, including a proposed VLM score, align poorly with human judgment and can favor oversmoothed outputs. The dataset and code are released.","tokens_in":16341,"tokens_out":7367,"duration_ms":69563,"significance":"The PMoC3D dataset release is a valuable contribution: raw paired k-space data with motion corruption is scarce, and the paper provides a thoughtful discussion of the mild-motion failure mode of reference-based evaluation. If the central claims hold, the paper would establish a preferred benchmark and metric for 3D MRI motion correction evaluation. The validation of MoMRISim against an independent human-based score is a strength, as is the honest reporting of limitations. However, the empirical support is weakened by the small sample size (8 subjects), the use of a single-annotator ground truth for the key correlation, and the absence of confidence intervals or inter-rater agreement measures. The claims should therefore be treated as preliminary until the authors provide uncertainty quantification and ideally multi-rater evidence.","major_comments":[{"comment":"The central claim that MoMRISim provides a reliable real-world evaluation metric rests on the Spearman correlation rho=0.95 with the perceived motion artifact score (PMAS). Appendix B states that for the PMAS used in Section 4.2, 'annotations were performed by a single evaluator.' No inter-rater agreement, rater blinding, or confidence interval for the correlation is reported. Because the central claim defines reliability as agreement with human perception, this single-annotator ground truth is load-bearing. I recommend adding a second (or more) rater on at least a subset of the comparisons, reporting inter-rater agreement, and providing bootstrap confidence intervals on the reported correlations. Alternatively, the claims should be explicitly limited to agreement with this one annotator.","section":"Section 4.2 and Appendix B"},{"comment":"The conclusion that simulated-motion evaluation systematically exaggerates algorithm performance is primarily supported by a 64-pair comparison performed by a single 'human annotator,' with no reported expertise, blinding procedure, or inter-rater reliability. The matching of artifact severity between real-world and simulated volumes is described only qualitatively ('closely resembling'), which could bias the comparison if the simulated volumes were in fact more severely corrupted. Please provide the exact number of simulated volumes, the severity-matching procedure (including any quantitative criteria), and ideally a second annotator or a sensitivity analysis.","section":"Section 4.3"},{"comment":"The correlations in Figure 2 and Appendix F.1 are reported without confidence intervals or significance tests on the differences between metrics (e.g., MoMRISim rho=0.95 versus DreamSim rho=0.93). The sample arises from 8 subjects, and the reconstruction-level PMAS involves only 48 or 72 volumes; with such a small sample, the ordering of metrics may not be stable. Please provide bootstrap confidence intervals on all reported Spearman correlations and test whether the difference between MoMRISim and DreamSim is statistically significant.","section":"Section 4.2 and Appendix F.1"}],"minor_comments":[{"comment":"The x-axis extends to 6 while the PMAS values in Table 2 lie in [-0.44, 2.42]; the caption also says 'medium and severe' but negative PMAS values (mild) appear. Please clarify the axis range and the subset of scans actually shown.","section":"Figure 2"},{"comment":"An acceleration factor of 4.9 is reported, whereas Section 2.1 and Table 1 report 4.94; please make these values consistent.","section":"Section 3.2"},{"comment":"The learning rates and regularization weights contain missing superscripts ('learning rate of 10 8' should be 10^{-8}; 'lambda = 3-8' should be 3 x 10^{-8}).","section":"Appendix C.1 and C.2"},{"comment":"The caption reads 'MoMRIScore' instead of 'MoMRISim'; please correct the typo.","section":"Figure 2 caption"},{"comment":"The sentence 'We provide access to the corrupted scans (S3 4, S5 4, S8 4) with involuntary motion resulting in a total of 27 motion-corrupted scans' is unclear: the notation and the total count (24 voluntary plus 3 involuntary?) should be explained.","section":"Section 2.1"},{"comment":"The reported triplet-ranking accuracy of 0.933 is computed on the training set; please also report validation accuracy to indicate generalization.","section":"Appendix D"},{"comment":"The VLM score uses GPT-4o at temperature 0.5, but no specific model version or access date is given, which limits reproducibility; please add this information.","section":"Appendix E"}],"recommendation":"major_revision","confidential_remarks":"The paper's main claims are plausible and the dataset release is a real contribution, but the evidence base is thin. The key issue is the single-annotator PMAS used to validate MoMRISim; without inter-rater agreement or uncertainty quantification, the headline correlation of 0.95 is not yet convincing as a general result. I would support publication after the authors add uncertainty quantification and multi-rater evidence, or substantially soften the central claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: PMoC3D is the real deliverable, and it's a good one. The paper gives the community the first public raw k-space dataset of paired motion-corrupted and motion-free 3D brain scans, with sensitivity maps, trajectory, and motion instructions. That alone justifies engagement. MoMRISim is a reasonable proposal, and the three-way comparison of real, simulated, and reference-free evaluation is new and useful.\n\nThe paper is also honest. It reports the mild-motion breakdown where reference-based evaluation breaks down, and it concedes that the reference scan is not true ground truth. The finding that simulated motion exaggerates performance is plausible and I believe it; the qualitative examples are strong.\n\nThe soft spots are real but mostly in the validation of the central reliability claim. The PMAS ground truth for the reconstruction-level correlations (Section 4.2) comes from a single annotator, per Appendix B. That means the Spearman rho = 0.95 for MoMRISim is agreement with one person's notion of severity, not necessarily with human perception generally. No inter-rater agreement, no confidence intervals, and only 8 subjects. The simulated-vs-real comparison in Section 4.3 uses one annotator and 64 pairs, with the same caveat. I don't think these flaws are fatal — the paper is upfront about many of its own limitations and the dataset is genuinely valuable — but the headline claim 'real-world evaluation with MoMRISim is most reliable' is not yet nailed down. It needs a proper observer study with multiple raters and uncertainty quantification.\n\nFor peer review: yes, this deserves a serious referee. A desk rejection would be wrong. The dataset and code are public, the metric is trainable, and the claims are important. I'd recommend the editor send it out and ask the authors to strengthen the human-rater evidence, but even as is, the work is a useful contribution.","headline":"The raw k-space dataset is the real contribution; the reliability claim leans on a single human rater and needs stronger validation.","tokens_in":16868,"tokens_out":2937,"would_cite":true,"duration_ms":29138,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Real-world paired evaluation with the new PMoC3D dataset and MoMRISim metric reliably tracks human judgment of MRI motion correction, while simulated-motion and reference-free evaluation mislead.","keywords":["MRI motion correction","PMoC3D dataset","MoMRISim","evaluation methodology","perceptual image quality metrics","simulated motion artifacts","reference-free evaluation","3D brain MRI"],"falsifier":"Have several independent radiologists rank the same set of PMoC3D reconstructions pairwise, average their preferences into a consensus PMAS, and recompute correlations: if MoMRISim's Spearman correlation with the consensus drops well below 0.95, or if another metric matches or exceeds it, the paper's claim that MoMRISim is the most reliable reference-based metric would be falsified. A second check: retrain MoMRISim on a different brain dataset with the same triplet scheme and see whether its correlation with human judgment on PMoC3D is preserved, which would test whether the result depends on the specific training data.","tokens_in":15933,"feed_emoji":"🧠","tokens_out":9001,"duration_ms":78176,"temperature":0.7,"pith_summary":"Correcting motion artifacts in 3D brain MRI is clinically valuable, but the field has no agreed way to tell whether a correction method actually works, because true ground-truth scans of moving patients can never be obtained. This paper compares the three evaluation strategies in use—scoring against a motion-free reference scan, scoring against simulated motion, and scoring without any reference—and finds that two of them mislead. Simulated motion makes correction methods look dramatically better than they perform on real scans, even when artifact severity is matched; reference-free metrics, including a GPT-4o-based score, systematically overrate smooth deep-learning outputs that have lost anatomical detail. The paper's positive contribution is a real-world paired dataset, PMoC3D, together with a learned feature-space metric, MoMRISim, whose rankings agree with human perception (Spearman's $\\rho = 0.95$) for moderate-to-severe motion. If this holds, the field finally has a benchmark and metric that measure real correction performance rather than simulation or smoothness.","feed_headline":"Simulated motion flatters MRI correction methods","feed_subtitle":"A new paired benchmark and perceptual metric track human judgment (ρ=0.95); no-reference scores reward oversmoothing.","key_machinery":"The argument is carried by three objects. PMoC3D is a dataset of unprocessed, paired 3D brain MRI acquisitions—one motion-free reference and three motion-corrupted scans per subject, with k-space (raw spatial-frequency) data, sensitivity maps, and per-shot motion instructions—so that motion-estimation methods that need raw data can be tested on real artifacts. MoMRISim is a learned perceptual metric: a DINO-vitb16 visual encoder fine-tuned with LoRA (a lightweight low-rank adapter) on triplets (reference, reconstruction A, reconstruction B) labeled by simulated motion severity, trained so that it ranks the less-corrupted reconstruction as closer to the reference; it thereby learns motion-artifact features without human annotation. PMAS, the perceived motion artifact score, is the yardstick of reliability: pairwise human comparisons are fit with a Bradley-Terry model, and every evaluation approach and metric is judged by how well its rankings agree with PMAS. The third ingredient is the three-way comparison protocol itself—real-world paired, simulated, and reference-free evaluation run on the same three baseline methods (alternating optimization, MotionTTT, and stacked U-nets)—which is what lets the paper attribute over- and under-rating to the evaluation approach rather than to any single algorithm.","core_discovery":"The central claim is that evaluation on real-world paired data—the PMoC3D benchmark the paper releases, scored with the MoMRISim metric it also introduces—gives a relatively reliable and meaningful measure of 3D MRI motion-correction performance under moderate to severe motion, and that the two popular alternatives do not. Using perceived motion artifact scores (PMAS), obtained from pairwise human comparisons fitted with a Bradley-Terry model, as the yardstick, the paper reports that MoMRISim's ranking of reconstructions correlates with human judgment at Spearman $\\rho = 0.95$, above PSNR, SSIM, DISTS, DreamSim, and the proposed VLM score. On simulated motion, a blind comparison showed that in 75% of matched-severity pairs the real-world reconstruction had clearly more residual artifacts than its simulated counterpart, so simulation-based evaluation systematically exaggerates algorithm performance. Reference-free metrics, including the vision-language-model score proposed here, weakly align with human judgment and hand high scores to oversmoothed stacked U-net outputs that visibly lose anatomy. The paper also documents a boundary of its own recommendation: under mild motion, corrected images can look cleaner than the motion-free reference, so reference-based evaluation loses validity exactly where artifacts are already subtle.","pith_inferences":["A multi-rater study of the same reconstructions would test how stable the reported 0.95 correlation is; if agreement between raters is low, the metric ranking itself may change.","MoMRISim is trained on simulated motion yet validated against real motion—if that transfer holds, the same triplet scheme could be extended to other artifact families (non-rigid motion, pulsation, spin-history effects) that current simulation cannot capture.","The VLM-score failure is a caution for the wider practice of using vision-language models as automatic image-quality judges: without explicit anti-oversmoothing constraints, such scores will systematically favor plausible-looking but detail-poor reconstructions.","Because PMoC3D records per-shot motion instructions and time stamps, it could double as a testbed for predicting motion severity from raw data, not just for evaluating correction algorithms."],"forward_implications":["Papers evaluating 3D MRI motion correction should report results on real paired data like PMoC3D with a feature-based metric, since simulated-motion numbers will systematically overstate progress.","Simulated-motion evaluation remains useful only for relative comparisons or the mild-motion regime; matched-severity real artifacts are consistently harder to correct.","Reference-free metrics, including VLM-based scores, are not reliable for ranking deep-learning motion-correction methods because they reward smoothness at the cost of anatomy.","Motion-correction methods that estimate motion parameters from raw k-space can finally be evaluated on real motion, because PMoC3D provides unprocessed measurements rather than only magnitude images.","For mild motion, reference-based scoring is unsafe: a corrected image can legitimately beat the motion-free reference, so benchmark design must separate mild from moderate and severe cases."],"supporting_citations":[{"why":"Supplies the Bradley-Terry paired-comparison model used to convert pairwise human judgments into the PMAS yardstick that every metric is validated against.","marker":"[BT52]"},{"why":"Provides the triplet-based training approach and DINO-based architecture that MoMRISim adapts with LoRA fine-tuning and simulated-motion labels.","marker":"[Fu+23]"},{"why":"The Calgary Campinas brain MRI dataset is the training source for MoMRISim, MotionTTT, and the stacked U-net, and the base for the simulated-motion evaluation.","marker":"[Sou+18]"},{"why":"Defines the motion-simulation protocol and the MotionTTT baseline; MoMRISim's triplet severity labels and the paper's simulated artifacts follow this protocol.","marker":"[Klu+24]"},{"why":"Supplies the classical alternating-optimization baseline and the sensitivity-encoding motion-model framework used for real-data reconstruction.","marker":"[Cor+16]"},{"why":"Supplies the stacked U-net baseline whose oversmoothed outputs expose the failure of reference-free metrics.","marker":"[Al-+22]"},{"why":"Motivates the preprocessing pipeline for paired scans and documents that classical gradient-based metrics correlate poorly with human judgment.","marker":"[Mar+24]"}],"fun_headline_variants":["Real-world data beats simulation for MRI motion evaluation","Simulation exaggerates MRI motion correction success","New MRI benchmark and metric target motion artifact quality","No-reference metrics misjudge MRI motion fixes","MoMRISim ranks MRI corrections like human experts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Reliability is defined as agreement with the perceived motion artifact score, but the pairwise human comparisons behind that score were made by a single evaluator (Appendix B), so the reported 0.95 correlation shows agreement with one person's judgment rather than with a consensus of radiologists.","fun_headline_variants_meta":{"raw":{"variants":["Real-world data beats simulation for MRI motion evaluation","Simulation exaggerates MRI motion correction success","New MRI benchmark and metric target motion artifact quality","No-reference metrics misjudge MRI motion fixes","MoMRISim ranks MRI corrections like human experts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000375,"raw_usage":{"total_tokens":2006,"prompt_tokens":954,"completion_tokens":1052,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":570,"completion_tokens_details":{"reasoning_tokens":981}},"tokens_in":570,"tokens_out":1052,"duration_ms":10834,"temperature":1.0,"reasoning_tokens":981,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:12:45.354905+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several independent radiologists rank the same set of PMoC3D reconstructions pairwise, average their preferences into a consensus PMAS, and recompute correlations: if MoMRISim's Spearman correlation with the consensus drops well below 0.95, or if another metric matches or exceeds it, the paper's claim that MoMRISim is the most reliable reference-based metric would be falsified. A second check: retrain MoMRISim on a different brain dataset with the same triplet scheme and see whether its correlation with human judgment on PMoC3D is preserved, which would test whether the result depends on the specific training data.","supporting_citations":[],"review_version":1}