{"id":"95e2fcf6-6b38-4c13-a436-339286ee25ff","arxiv_id":"2412.18389","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Reference-based image quality metrics, especially with percentile normalization and brain masking, best match radiologists' ratings of motion-affected MRI scans.","lead":"This study compares ten image quality metrics against radiologist scores on brain MRI scans with real motion artifacts. It finds that reference-based metrics match human ratings well, while preprocessing choices such as normalization and brain masking strongly change the results.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strong-correlation claim rests on treating repeated volumes per participant as independent; without subject-level clustering, reported Spearman magnitudes and p-values may be inflated.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the statistical analysis treats each image volume as an independent observation despite clear nesting within participants. I agree this is the most important threat to the central claim. My read reinforces the reader's point rather than introducing a different objection. In fact, the concern is slightly stronger than the reader states: non-independence can inflate not only p-values but also the point estimates of Spearman rho, because shared participant-level anatomy and scoring tendencies can create a spurious subject-level correlation between IQM values and observer scores. A concrete subject-level re-analysis would settle whether the 'strong correlation' claim survives. The reader's CONDITIONAL verdict already reflects that this concern is specific and addressable, so I do not propose changing the verdict. I also note the paper's acknowledged limitations and the missing code link as secondary issues, but neither is as load-bearing as the clustered-data problem.","tokens_in":12708,"tokens_out":5556,"duration_ms":59745,"concrete_test":"Recompute the main Spearman correlations separately for each sequence after (i) removing reference volumes and (ii) collapsing the remaining repeated acquisitions to one observation per participant, e.g. the median metric value and median observer score across that participant's motion/correction variants. Report the resulting rho values with cluster-bootstrap 95% confidence intervals. If a reference-based metric's |rho| drops below 0.6 in any sequence, or if the confidence intervals are wide enough to include weak correlations, the claim of uniformly strong correlation is not robust to the repeated-measures structure and the manuscript should instead report within-subject or mixed-effects associations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('All reference-based IQMs show a strong correlation with radiological assessment, with small variations in their relative performance for different MR sequences', Section 3.2) is supported by Spearman correlations computed across individual image volumes. The NRU dataset contains 22 participants and the CUBRIC dataset 9, with each participant contributing multiple motion/correction variants per sequence. Volumes from the same participant are not independent: they share anatomy, the same reference image, likely similar motion behavior, and are scored by the same raters. The analysis in Section 2.4 treats each volume as an observation with no random effects, clustering, or subject-level bootstrap. This has two consequences. First, p-values are anticonservative: the effective sample size is much smaller than the number of volumes, so the significance flags in Figures 3 and 6 are weaker than presented. Second, the point estimates themselves can be biased upward: if participants differ systematically in how scoreable or atypical their anatomy is, both IQM values and observer scores inherit a subject-level component that is mistaken for agreement about motion artifacts. The reported 'strong correlation' could therefore partly reflect subject identity rather than the ability of metrics to track radiological quality within subjects. The manuscript acknowledges small samples in Section 4.1 but does not quantify the clustering effect. The missing code link and commit hash also prevent independent re-analysis, but the statistical issue is the more load-bearing one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares ten image quality metrics (five reference-based, five reference-free) against radiologist/radiographer Likert scores on two small MRI datasets with real motion artifacts. Metrics are computed under several preprocessing choices (masking, normalization, slice reduction), and Spearman correlations are used to rank metrics and to identify which preprocessing choices best agree with human assessment. The central claim is that all reference-based metrics correlate strongly with radiological evaluation across sequences and datasets, that preprocessing choices (especially normalization and brain masking) substantially affect correlation values, and that among reference-free metrics only AES and TG are reasonably consistent. The manuscript concludes with practical recommendations to use percentile normalization and brain masking and to prefer reference-based metrics when a reference is available.","tokens_in":12870,"tokens_out":4211,"duration_ms":43221,"significance":"The study addresses a practical question in MR motion-correction evaluation: which IQMs actually track radiological quality on real motion-corrupted data. Its strengths are the use of real rather than simulated motion, the inclusion of two datasets and multiple sequences, the comparison of five preprocessing conditions, and the public availability of one dataset and (claimed) analysis code. If the correlation claims survive reanalysis, the paper would provide useful empirical guidance for metric choice and preprocessing standardization. The conclusions are, however, currently vulnerable to the statistical treatment of repeated volumes per participant, and no code link or confidence intervals are provided, so the quantitative strength of the claims is not yet established.","major_comments":[{"comment":"The Spearman correlations are computed over individual image volumes, but the NRU dataset contains 22 participants and the CUBRIC dataset 9, with multiple volumes per participant (different sequences, motion conditions, and correction states). Volumes from the same participant share anatomy, the same reference image, and the same raters, and therefore are not independent. Treating them as independent makes p-values anticonservative and can inflate the correlation estimates by mixing a subject-level component into the association between IQMs and observer scores. This directly affects the central claim in Section 3.2 and the significance flags shown in Figs. 3 and 6. Please reanalyze with subject-level clustering (e.g., cluster bootstrap by participant, mixed-effects models, or within-subject correlations) and report effective sample sizes.","section":"Section 2.4, Figs. 3 and 6"},{"comment":"No confidence intervals are reported for any Spearman coefficient, and no correction is applied for the many comparisons across ten metrics and multiple preprocessing settings. Given the small and dependent samples, a binary display of 'significant' versus 'non-significant' (as in Figs. 3 and 6) is not sufficient support for the numeric strength of the correlations. Add confidence intervals for the coefficients and either apply multiplicity control or explicitly label the exploratory preprocessing comparisons as hypothesis-generating.","section":"Sections 2.4 and 3.2"},{"comment":"The statement 'We did not observe a significant difference in the correlation coefficients for different slice reduction methods' is presented as a finding, but no formal comparison, confidence interval, or equivalence test is given. The visual comparison in Fig. 6 cannot establish that the two reduction methods are equivalent, particularly when the underlying correlations already lack uncertainty measures. Please either provide an appropriate statistical comparison or soften the claim to a descriptive observation.","section":"Section 3.3"}],"minor_comments":[{"comment":"The phrase 'intra-variability between evaluators' should be 'inter-rater variability' or 'inter-rater agreement', since Krippendorff's alpha measures agreement between raters, not variability within a single rater.","section":"Section 2.4"},{"comment":"The manuscript states that the code is 'publicly available on GitHub' but provides no repository URL, username, or version/commit identifier. This makes the reproducibility claim incomplete and should be fixed in the revision.","section":"Code Availability"},{"comment":"The funding paragraph contains an incomplete sentence: 'Additionally we would like to acknowledge the following funding sources:' is immediately followed by the Declarations section without any funding text. Please complete this sentence or remove the dangling phrase.","section":"Funding statement"},{"comment":"The abstract says the metrics were 'recalculated seven times' with varying preprocessing steps, but the preprocessing grid described in Fig. 1 appears to contain more than seven combinations (three mask options, several normalization options, and two reduction options). Clarify how the 'seven times' is defined, or state the exact number of settings used.","section":"Abstract and Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The central empirical claim is plausible and the dataset is valuable, but the statistical analysis as presented does not yet support the strength of the conclusions. The missing GitHub link is a simple fix, while the independence issue requires a substantive reanalysis, which is why I recommend major revision rather than rejection. I do not see any circularity or misconduct concern; the authors' prior ISMRM abstracts are appropriately extended here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper has the best data of its kind—real motion artifacts with radiologist scores—and it sweeps preprocessing systematically, which is genuinely new. The main qualitative finding, that reference-based IQMs beat reference-free ones and that AES is the best reference-free option, is plausible and consistent with earlier work. But the quantitative claim of “strong correlation” is built on Spearman correlations that treat each image volume as an independent observation, with only 22 and 9 subjects underneath. That inflates significance and can overstate the correlation magnitude.\n\nWhat’s actually new: the real-motion datasets (NRU public, CUBRIC private), the systematic variation of masking, normalization, and slice reduction, and the inclusion of VIF and LPIPS. The authors are also reasonable about limitations: they note the scans are research-grade, 3T only, and that IQMs are not proper metrics. Their recommendation to use a set of metrics rather than a single one is sound.\n\nThe soft spots are real. The main one is the independence assumption: volumes from the same subject share anatomy, reference image, and motion behavior, so the effective sample size is far smaller than the number of volumes. No clustering, mixed effects, or subject-level bootstrap is used. P-values are anticonservative, and the point estimates can pick up subject identity rather than pure motion-related quality. The authors acknowledge small samples in Section 4.1 but do not quantify this. Also, no confidence intervals or multiple-comparison correction across the many metrics and preprocessing settings are provided. The missing GitHub URL and commit hash in the Code Availability statement is a minor reproducibility gap. None of this kills the direction, but the specific rho values and p<0.05 flags should be treated as optimistic.\n\nThis paper is for researchers benchmarking motion correction methods or choosing IQMs for their pipeline. It deserves a serious referee because the dataset and question matter, though the statistical treatment needs revision before I’d trust the numbers. I’d send it to peer review, but I’d ask for a subject-level analysis—cluster bootstrap or mixed model—plus confidence intervals, and the actual code link.","headline":"Real-motion IQM comparison with a genuinely useful dataset, but the correlation analysis needs subject-level clustering before the strong quantitative claims can be trusted.","tokens_in":13456,"tokens_out":2225,"would_cite":true,"duration_ms":21259,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper finds that reference-based image quality metrics correlate strongly with radiologist ratings of real motion-corrupted MRI, and that percentile normalization plus brain masking gives the strongest agreement.","keywords":["MRI","image quality metrics","motion artifacts","radiological evaluation","Spearman correlation","brain masking","normalization","reference-based metrics"],"falsifier":"Recompute the Spearman correlations with one averaged data point per participant instead of one per volume; if the reference-based metrics no longer exceed the strong-correlation threshold of |ρ| > 0.7, the paper's central claim is not robust to the non-independence of repeated volumes from the same subject.","tokens_in":12467,"feed_emoji":"🧠","tokens_out":8657,"duration_ms":66468,"temperature":0.7,"pith_summary":"This paper asks whether quantitative image quality metrics, the kind used to rank motion correction algorithms, actually agree with what radiologists see when MRI scans are corrupted by real head motion. On two datasets of brain scans acquired with and without intentional motion, the authors compare five reference-based metrics (SSIM, PSNR, FSIM, VIF, LPIPS) and five reference-free metrics (Tenengrad, Average Edge Strength, Normalized Gradient Square, Image Entropy, Gradient Entropy) against 1–5 Likert ratings from radiologists and radiographers, recalculating each metric under seven preprocessing variants. The central result is that all reference-based metrics correlate strongly with observer scores across sequences and datasets, that among reference-free metrics only Average Edge Strength and Tenengrad correlate consistently, and that preprocessing choices—especially percentile normalization and brain masking—strongly modulate the correlation. If this is right, then reference-based metrics can be trusted as proxies for expert image quality in motion-correction research, provided the preprocessing is reported and standardized.","feed_headline":"Reference-based MRI metrics track radiologist scores under motion","feed_subtitle":"Percentile normalization and brain masking give the strongest match to expert scores on motion-corrupted scans.","key_machinery":"The engine of the analysis is the Spearman rank correlation between each image quality metric and the averaged observer score, computed over many image volumes per sequence and dataset, together with a preprocessing grid that varies normalization (none, min-max, mean-std, percentile), brain masking (none, direct masking, multiplication), and slice reduction (mean or worst). The paper also uses the observer score as the gold standard, with Krippendorff's alpha to check inter-rater reliability. The preprocessing grid is what allows the paper to isolate which implementation choices drive agreement with radiological evaluation.","core_discovery":"On its own terms, the paper establishes that reference-based image quality metrics agree with radiological assessment in real-motion MRI: Spearman correlations between metric values and expert scores were strong for SSIM, PSNR, FSIM, VIF, and LPIPS across MP-RAGE, T2 FLAIR, T1 STIR, and T2 TSE acquisitions, with small differences among metrics. The paper also shows that the preprocessing recipe matters as much as the metric family: percentile normalization and restricting computation to a skull-stripped brain mask produced the strongest correlations, while omitting the mask or using min-max/no normalization degraded them; the choice of mean versus worst-slice reduction had little effect. Among reference-free metrics, Average Edge Strength and Tenengrad correlated with observer scores most consistently but more weakly than the reference-based group. The authors conclude that reference-based metrics are preferable when a reference image exists and that preprocessing choices must be documented for results to be reproducible.","pith_inferences":["Because multiple volumes per participant enter the Spearman correlation as independent points, the reported p-values and confidence levels likely overstate how precisely the correlations are known; a subject-clustered re-analysis might shrink the correlations, though the ranking between metrics could survive.","The brain-masking result is probably specific to anatomies with large background fractions; in cardiac or abdominal imaging, where the background is smaller, masking may matter less, as the paper itself notes.","A natural next step, which the paper's outlook gestures toward, is to train a reference-free model on the observer scores directly; the strong preprocessing sensitivity found here suggests such a model should mimic the percentile-normalized, brain-masked view of the image.","The 'hidden noise' in reference images that some prior work has identified could mean that even reference-based metrics are limited by reference quality; pairing this approach with reference-quality assessment would test that boundary."],"forward_implications":["Reference-based IQMs (SSIM, PSNR, FSIM, VIF, LPIPS) can be used as stand-ins for radiological scoring when benchmarking motion correction on brain MRI with a reference scan.","Percentile normalization and brain masking should be adopted as default preprocessing for these metrics, since other choices materially weaken correlation with expert judgment.","Reference-free metrics, especially Average Edge Strength and Tenengrad, offer a weaker but usable fallback when no reference image exists.","The slice reduction choice (mean vs worst) does not change conclusions, so simpler mean-based implementations are acceptable.","IQM studies that skip brain masking or use min-max normalization risk drawing different conclusions about which reconstruction is best."],"supporting_citations":[{"why":"Supplies the publicly available dataset of real instructed-motion scans with reference images across four sequences.","marker":"[18]"},{"why":"Supplies the private dataset of varied real motion patterns with retrospective correction for MP-RAGE scans.","marker":"[19]"},{"why":"Prior comparison of objective IQMs to radiologist scoring, which this study extends to real-motion data.","marker":"[6]"},{"why":"Larger-scale prior evaluation of MR IQMs against expert ratings, providing the baseline that the present metric selection builds on.","marker":"[7]"},{"why":"Defines SSIM, one of the five reference-based metrics whose correlations carry the central claim.","marker":"[21]"},{"why":"Defines PSNR, another reference-based metric in the comparison.","marker":"[22]"},{"why":"Defines FSIM, a reference-based metric shown here to correlate strongly with observers.","marker":"[23]"},{"why":"Defines VIF, the reference-based metric that can exceed 1 and captures information fidelity.","marker":"[24]"},{"why":"Defines LPIPS, the deep-feature perceptual reference-based metric included in the study.","marker":"[13]"},{"why":"Defines the reference-free gradient and entropy metrics (TG, NGS, IE, GE) used for comparison.","marker":"[20]"}],"fun_headline_variants":["Reference-based MRI metrics mirror radiologist scores in motion","Brain mask and percentile normalization align MRI metrics with experts","Preprocessing choices decide MRI metric-radiologist agreement","Average Edge Strength leads reference-free MRI metrics under motion","Motion-corrupted MRI: reference-based metrics track expert ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis treats every image volume as an independent observation in the Spearman correlations, although many volumes come from the same participant; shared anatomy and motion behaviour across those volumes could make the effective sample size smaller than the number of volumes, so the reported correlation significance may be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Reference-based MRI metrics mirror radiologist scores in motion","Brain mask and percentile normalization align MRI metrics with experts","Preprocessing choices decide MRI metric-radiologist agreement","Average Edge Strength leads reference-free MRI metrics under motion","Motion-corrupted MRI: reference-based metrics track expert ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000227,"raw_usage":{"total_tokens":1505,"prompt_tokens":1009,"completion_tokens":496,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":419}},"tokens_in":625,"tokens_out":496,"duration_ms":5162,"temperature":1.0,"reasoning_tokens":419,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:41:56.870501+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the Spearman correlations with one averaged data point per participant instead of one per volume; if the reference-based metrics no longer exceed the strong-correlation threshold of |ρ| > 0.7, the paper's central claim is not robust to the non-independence of repeated volumes from the same subject.","supporting_citations":[{"cited_title":"Journal of Magnetic Resonance Imaging 11(2), 174–181 (2000)","cited_arxiv_id":null,"evidence_quote":"Defines the reference-free gradient and entropy metrics (TG, NGS, IE, GE) used for comparison."},{"cited_title":"Web site: https://openneuro.org/datasets/ds004332/versions/1.0.0","cited_arxiv_id":null,"evidence_quote":"Supplies the publicly available dataset of real instructed-motion scans with reference images across four sequences."},{"cited_title":"Magnetic resonance in medicine 90(4), 1297–1315 (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies the private dataset of varied real motion patterns with retrospective correction for MP-RAGE scans."},{"cited_title":"IEEE Transactions on Medical Imaging 39(4), 1064–1072 (2020)","cited_arxiv_id":null,"evidence_quote":"Prior comparison of objective IQMs to radiologist scoring, which this study extends to real-motion data."},{"cited_title":"IEEE Access 11, 14154–14168 (2023) 16","cited_arxiv_id":null,"evidence_quote":"Larger-scale prior evaluation of MR IQMs against expert ratings, providing the baseline that the present metric selection builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines PSNR, another reference-based metric in the comparison."},{"cited_title":"IEEE transactions on Image Processing 20(8), 2378– 2386 (2011)","cited_arxiv_id":null,"evidence_quote":"Defines FSIM, a reference-based metric shown here to correlate strongly with observers."},{"cited_title":"IEEE Transac- tions on image processing 15(2), 430–444 (2006)","cited_arxiv_id":null,"evidence_quote":"Defines VIF, the reference-based metric that can exceed 1 and captures information fidelity."},{"cited_title":"In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, pp","cited_arxiv_id":null,"evidence_quote":"Defines LPIPS, the deep-feature perceptual reference-based metric included in the study."}],"review_version":1}