{"id":"f68f7fe9-4c80-4d04-8706-32d25e767a74","arxiv_id":"2505.07175","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"No-reference image quality metrics frequently fail to detect localised anatomical errors in synthetic brain MRI and can rank generative models inconsistently with downstream segmentation utility.","lead":"This paper stress-tests popular no-reference image quality metrics, such as FID and KID, for evaluating synthetic brain MRI and angiography images under controlled noises, anatomical distortions, and distribution shifts. It finds that most upstream metrics miss localised clinical errors and often rank generative models differently from downstream segmentation performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Evanyseg proxy is the load-bearing yardstick for every upstream-vs-downstream comparison, yet its calibration on synthetic and morphologically perturbed inputs is never validated against ground-truth Dice.","rationale":"The reader's verdict is CONDITIONAL with moderate confidence, and the reader's weakest_assumption is precisely the Evanyseg calibration issue. My stress-test pass independently identifies the same load-bearing concern after reading the full text: the results section repeatedly draws conclusions about 'downstream task performance', but every such conclusion is mediated by the Evanyseg proxy, whose behaviour on synthetic images and on Phase-1 perturbed images is unvalidated. The paper's own Limitations section concedes the dependency on 'the Evanyseg method', which reinforces rather than undermines the concern. I also considered two other candidate concerns: (a) no error bars or significance tests across the score tables, and (b) single-checkpoint generative models. These are real weaknesses but secondary; even with perfect statistics, the interpretation of the downstream scores would remain hostage to Evanyseg calibration. The paper has genuine independent support: a public repository, a systematic and clearly described perturbation design, and specific falsifiable claims such as the morphological-insensitivity results and the VAE/GAN reversal in Table 6. Those claims are credible as observations about the metrics tested, but the central comparative conclusion (that downstream evaluation is more meaningful) is only as strong as the unvalidated proxy. Hence the verdict should remain CONDITIONAL, with the condition being an explicit ground-truth validation of Evanyseg on the exact input regimes used. The required concrete test is feasible because Phase 1 starts from real images with ground-truth masks, so true Dice can be computed directly for the perturbed images; Phase 2 lacks paired ground truth, but the proxy can at least be validated on out-of-distribution real-like synthetic images if any paired reference exists, or by a synthetic-to-real transfer check.","tokens_in":24282,"tokens_out":1996,"duration_ms":17379,"concrete_test":"Reproduce the quality-predictor validation using real IXI and BraTS images with ground-truth masks: apply the exact Phase-1 perturbations (boundary blur, radial intensity gradients, noise, contrast remapping) and the exact Phase-2 generative model outputs if any paired ground truth exists, compute true Dice between the segmentor's output mask and the ground-truth mask, and compare true Dice against Evanyseg's predicted Dice (bias, calibration, rank order). If the rank correlation between predicted and true Dice is high (e.g., Spearman rho > 0.8) and the Table 3 magnitude of effect survives on true Dice, the concern is resolved. If predicted Dice is biased or rank-invariant under the tested perturbations, the central claim needs to be re-baselined against direct downstream metrics.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim is that upstream NRIQMs 'correlate poorly with downstream task suitability' and are 'profoundly insensitive' to localised anatomical changes. Every instance of 'downstream task suitability' is measured with the Evanyseg quality-prediction model (Section 2.2.1). This model is trained on real images paired with synthetically perturbed segmentation masks, using Dice against the real ground-truth mask as the regression target. It is then applied in two out-of-distribution regimes: (1) Phase 1 perturbed real images, where the input image itself is altered by boundary blurring, intensity gradients, noise, and contrast remapping, and (2) Phase 2 fully synthetic images from VAE/GAN/DDPM, whose segmentation masks come from a segmentor applied to the synthetic input. The paper does not report any validation of Evanyseg's predicted Dice against true Dice on held-out images in either regime. The morphology tables (Table 3) are especially vulnerable: the claim that upstream metrics 'completely failed' while the downstream task 'registered a response' rests on predicted-Dice drops from 0.948 to 0.845, but if Evanyseg is biased by the same local blurring it is trained to emulate (erosion/dilation-like perturbations), the apparent downstream sensitivity could be an artefact of the proxy rather than a genuine task-level effect. The paper itself acknowledges this dependency in its Limitations paragraph ('the downstream evaluation relied on... the Evanyseg method'), so the concern is explicitly conceded rather than manufactured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates a suite of no-reference image quality metrics (NRIQMs) for assessing generative medical images, using brain MRI data (BraTS tumour images and IXI vascular MRA). In Phase 1, the authors apply controlled perturbations (noise, morphological manipulations, distribution shifts) to real images and measure (i) upstream metric responses via Z-score heatmaps and (ii) downstream segmentation quality via the Evanyseg predicted-Dice model. In Phase 2, they apply the same metrics and the Evanyseg-based downstream evaluation to outputs from pretrained VAE, GAN, and DDPM models. The central claims are that many upstream NRIQMs correlate poorly with downstream task suitability, are profoundly insensitive to localised morphological alterations, and can yield misleadingly optimistic scores under data memorisation and mode collapse; the authors recommend a multifaceted evaluation framework. The manuscript includes a literature review, detailed methodology, tables of Evanyseg scores, and an open-source GitHub repository, but the downstream yardstick is a regression model whose predictions are not validated for the perturbed and synthetic inputs to which it is applied.","tokens_in":24489,"tokens_out":4289,"duration_ms":43585,"significance":"If the findings are correct, the paper makes a useful empirical contribution by cataloguing the behaviour of a broad set of NRIQMs under clinically motivated perturbations and by demonstrating concrete ranking divergences between upstream metrics and a downstream segmentation proxy. The observation that distance metrics such as FID/KID can appear to improve under data duplication or mode collapse is a valuable caution that aligns with prior literature, and the open-source implementation is a reproducible asset. The clinical-safety framing is appropriate. However, the significance is conditional on the validity of the Evanyseg proxy, because every upstream-versus-downstream comparison in Sections 4.1 and 4.2 rests on it, and the paper does not provide the calibration evidence needed to establish that proxy as a faithful measure of true segmentation quality in the regimes used. The absence of confidence intervals, significance tests, and replicated model instances further limits the strength of the conclusions as stated.","major_comments":[{"comment":"The Evanyseg quality predictor is trained only on pairs of real images with synthetically perturbed segmentation masks, yet it is applied to (a) real images whose image content itself is perturbed (boundary blurring, intensity gradients, noise) and (b) fully synthetic images from VAE, GAN, and DDPM models. The paper provides no validation of Evanyseg's predicted Dice against true Dice in either regime, even though true Dice is computable for all Phase 1 perturbed real images because ground-truth masks exist. Since every upstream-versus-downstream comparison in Sections 4.1 and 4.2 depends on this proxy, a systematic bias in Evanyseg—for example, a learned association between boundary blur and low score from the mask-perturbation training data—could produce the observed drops in Table 3 (0.948 to 0.845) even if real segmentation accuracy were unchanged. The central claim that upstream metrics 'correlate poorly with downstream task suitability' is therefore not adequately supported until the proxy is calibrated.","section":"§2.2.1, §4.1.2, §4.2"},{"comment":"All quantitative comparisons in Phase 1 are based on single runs with no confidence intervals, significance tests, or multiple random seeds. Statements such as 'profound insensitivity' and 'negligible changes' in Section 4.1.2 are not backed by any uncertainty quantification, and several reported differences are very small (e.g., Evanyseg scores of 0.907 versus 0.904 in Table 5, or the near-constant scores in Table 4). Without error bars or a hypothesis test, it is unclear whether these differences are meaningful effects or noise, and the strength of the paper's conclusions exceeds what the evidence supports.","section":"§4.1, Tables 2–5, Figs. 4–7"},{"comment":"The Phase 2 comparison uses a single pretrained instance of each generative architecture (VQ-VAE, WGAN, and latent diffusion model), with weights sourced from prior work or the Medigan library. There is no replication across seeds, training runs, or model variants. Consequently, the claimed VAE/GAN contradiction on IXI (Table 6: 0.86 versus 0.67) and the DDPM superiority could reflect model-specific hyperparameters, training data, or checkpoint quality rather than architecture-level properties. The conclusion that 'upstream metric rankings of generative models do not align with downstream segmentation performance' is stated as a general result but is supported only by an anecdotal comparison of single checkpoints.","section":"§3.2, §4.2"},{"comment":"In the morphological manipulation experiments on real BraTS images, the paper reports only Evanyseg predicted scores and not the actual Dice between the segmentor's output and the ground-truth masks, although those ground-truth masks are available. Reporting true Dice in this setting would (i) provide a direct, validated measure of downstream degradation under boundary blurring and intensity gradients, and (ii) serve as an in-phase calibration check for the Evanyseg proxy. Without this information, the claim that the downstream task 'registered a response' rests entirely on the unvalidated predictor.","section":"§3.1.2, §4.1.2"}],"minor_comments":[{"comment":"The metric is introduced as 'Ct Score (Clusterability Score)' but later referred to as 'CT Score' in Section 2 and Table 1; consistent notation would avoid confusion.","section":"§2.1"},{"comment":"Several entries contain artifact-like editorial notes, e.g., '(Note: Original formula matched FID, likely incorrect for likelihood)' for FLS and '(Note: Original formula seemed incomplete/unclear)' for Ct Score; these should be cleaned up and replaced with accurate formulas or removed.","section":"Table 1"},{"comment":"The diagram contains the typo 'Upsteam' for 'Upstream' in the 'Upsteam' box.","section":"Figure 2"},{"comment":"The abstract claims that metrics 'correlate poorly' with downstream suitability, but no correlation coefficient is computed anywhere; the evidence is a qualitative divergence in rankings across two datasets and three models. Consider qualifying the language to 'rankings diverge' or reporting an actual correlation metric.","section":"Abstract and §5"},{"comment":"Reference [32] (Han et al., GAN-based synthetic brain MR image generation) is cited as the source for the Wasserstein GAN architecture, but the WGAN was originally proposed by Arjovsky et al.; the citation should be corrected or the discrepancy clarified.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper reads more as an empirical benchmark than a methodological advance, which may be appropriate for the venue. The main risk is that the headline claims are phrased more strongly than the evidence supports: the Evanyseg proxy is unvalidated in the two regimes where it is used, and the statistical power is minimal. The authors can address the central issue by (i) computing true Dice for Phase 1 perturbed real images and reporting agreement with Evanyseg, and (ii) adding an out-of-distribution validation of Evanyseg on held-out real images with synthetic image perturbations. If the proxy cannot be validated, the conclusions should be explicitly reframed as applying to the Evanyseg proxy rather than to actual downstream segmentation performance. I would also encourage the authors to tone down the 'profound' and 'comprehensively' language until the supporting evidence is complete."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper does something useful—systematically stress-tests 15 no-reference/feature image quality metrics on brain MRI under controlled noise, morphological, and distribution-shift perturbations, then compares them against a downstream segmentation-quality proxy on real and synthetic images. The headline result is worth having: FID, KID, LPIPS, and most of the other metrics barely react when tumour boundaries are blurred or internal intensity gradients are changed, while segmentation quality drops measurably. That is a concrete, field-relevant negative result, and the perturbation design is thoughtful. The GitHub release helps reproducibility.\n\nThe paper also earns credit for showing the flip side: distance metrics can look better under mode collapse or memorisation, and on IXI the GAN beats the VAE on several distance metrics yet is much worse on downstream vessel segmentation (0.67 vs 0.86). That discrepancy is a good cautionary tale.\n\nNow the soft spots, in proportion. The load-bearing issue is the Evanyseg proxy. It is a learned regressor trained on real images with synthetically perturbed masks, targeting Dice against ground truth. The paper then applies it to fully synthetic images and to perturbed real images without validating that its predictions remain calibrated in those regimes. The stress-test note is correct: if the proxy is biased toward the kind of local blurring it was trained to emulate, the 'downstream registered a response' tables (e.g., 0.948 to 0.845 in Table 3) could be partly an artifact. The paper itself concedes this dependency in the Limitations paragraph, but every upstream-vs-downstream comparison passes through this yardstick. That makes the central claim conditional.\n\nAlso, no confidence intervals, no significance tests, and a single checkpoint per architecture. For BraTS there is no VAE, so the three-way comparison is lopsided. Minor: the FLS and Ct Score entries in Table 1 carry the authors' own notes that the formulas are 'likely incorrect' or 'unclear', which is honest but should be cleaned up.\n\nOverall, the reader's verdict is right. The upstream insensitivity to local morphology is directly visible and does not depend on the proxy; it's solid. But the claim that NRIQMs 'correlate poorly with downstream task suitability' is only as strong as the proxy, and that remains unvalidated. Some quite modest extra work—validate Evanyseg against true Dice on held-out synthetic images, add error bars, include a VAE for BraTS—would firm this up.\n\nWho should read it: anyone evaluating generative models in medical imaging, and people designing evaluation frameworks. It deserves a serious referee; the topic is important, the study is a genuine attempt at systematic evidence, and the weaknesses are addressable. Send it to review.","headline":"A well-designed stress-test of NRIQMs for medical image generation, with a real negative result, but the downstream proxy (Evanyseg) is unvalidated for synthetic inputs, making the central comparison conditional.","tokens_in":25104,"tokens_out":3106,"would_cite":true,"duration_ms":29105,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Many no-reference image quality metrics are blind to localised anatomical errors in synthetic MRI and can rank worse models as better.","keywords":["no-reference image quality metrics","generative medical imaging","FID","KID","downstream segmentation evaluation","brain MRI","data memorisation","mode collapse"],"falsifier":"Take the perturbed real BraTS images used in the morphological experiments, for which ground-truth tumour masks exist, and compute the true Dice between the segmentor's output and the ground truth; if Evanyseg's predicted Dice diverges from these true Dice as boundary blur or intensity gradient strength grows, the paper's downstream yardstick is miscalibrated.","tokens_in":24034,"feed_emoji":"🧠","tokens_out":5552,"duration_ms":50622,"temperature":0.7,"pith_summary":"This paper tries to establish that widely used no-reference image quality metrics are not reliable for judging synthetic medical images, especially when clinical safety is at stake. It compares upstream metric scores with downstream segmentation performance on brain MRI data, under controlled perturbations and across three generative architectures. The paper finds that nearly all tested metrics are insensitive to clinically relevant local changes, such as tumour boundary blurring and internal intensity gradients, that do degrade segmentation quality. It also finds that the metrics can rank a GAN above a VAE for vessel images even though the GAN's output is much worse for the downstream vessel segmentation task.","feed_headline":"Image-quality metrics miss the tumour changes that matter","feed_subtitle":"In brain MRI, FID-style scores can rank a worse generator above a better one for the clinical task.","key_machinery":"The key machinery is a two-phase comparative evaluation framework. Phase 1 applies controlled perturbations, including noise, morphological manipulations, and distribution shifts, to real BraTS and IXI images and measures both upstream NRIQM responses and downstream segmentation quality. Phase 2 applies the same suite of upstream metrics and the same downstream task to synthetic images from a VAE, a GAN, and a DDPM. The downstream yardstick is the Evanyseg quality predictor, a regression model described in Section 2.2.1 that predicts a Dice score from an input image and a segmentation mask without ground truth, and it is the instrument that exposes the divergence between metric scores and task utility.","core_discovery":"The central discovery, on the paper's own terms, is that no-reference image quality metrics can detect global distributional differences but fail to reflect whether a generated medical image is usable for a downstream clinical task, and are largely blind to localised anatomical alterations. Concretely, tumour boundary blurring and internal intensity gradients left upstream NRIQMs essentially unchanged while the predicted tumour segmentation score dropped from 0.948 at baseline to 0.845 at the strongest blur, and on IXI the GAN scored better than the VAE on several distance metrics yet scored 0.67 versus 0.86 on vessel segmentation. This means upstream metric rankings can be inverted relative to task suitability, and a model whose images contain clinically significant structural flaws can pass common no-reference quality checks.","pith_inferences":["A practical remedy implied by the findings is to report an authenticity-oriented metric such as AuthPct alongside distance metrics, because the paper shows the two respond in opposite directions to data duplication and can together flag memorisation.","The same local insensitivity is likely to appear in other imaging modalities such as CT or X-ray, since global feature statistics average away small regional changes; this is a testable extension rather than a claim of the paper.","The morphological perturbation setup could be turned into a calibration tool: for any candidate metric, define the smallest boundary blur or intensity gradient it can detect, and require that detection floor to sit below clinically significant changes.","The Evanyseg proxy could be validated against true Dice on the perturbed real images where ground truth masks exist; if its predictions diverge there, the upstream-versus-downstream comparisons in the paper would need to be re-based on a validated estimator."],"forward_implications":["Distance metrics like FID and KID can appear to improve when the generated set duplicates training data or drops rare anatomical classes, so a model that memorises or collapses modes may look better rather than worse.","No-reference image quality metrics cannot serve as a safety check for local anatomical fidelity: a synthetic tumour with blurred margins or altered internal texture can pass all tested no-reference checks.","Rankings produced by common distance metrics on medical image generations may invert the true task-based ranking, as with the GAN versus VAE comparison on IXI vessel segmentation.","Downstream segmentation alone is also insufficient, since it stayed near baseline under density shifts, mode collapse, and mode invention, so upstream and downstream evaluations must be combined.","Clinical deployment decisions based only on upstream metrics risk selecting models whose outputs mislead downstream decision-support tools."],"supporting_citations":[{"why":"Supplies the ground-truth-free regression method, Evanyseg, used as the downstream yardstick for segmentation quality.","marker":"[71]"},{"why":"Defines FID, the primary distance metric whose behaviour under perturbations and model outputs is under test.","marker":"[33]"},{"why":"Defines KID, the polynomial-kernel MMD distance metric that is a central tested upstream metric.","marker":"[8]"},{"why":"Provides the pre-trained VAE, GAN, and DDPM weights used to generate the IXI synthetic images.","marker":"[19]"},{"why":"Supplies the IXI brain MRA dataset used for vessel experiments and for density-shift scanner mix manipulations.","marker":"[35]"},{"why":"Supplies the BraTS brain tumour MRI data and ground truth masks used for morphological manipulation experiments.","marker":"[7]"},{"why":"Provides the pre-trained vessel segmentation model used as the downstream task for IXI images.","marker":"[18]"},{"why":"Defines the AuthPct authenticity metric whose contrasting behaviour with distance metrics exposes data memorisation.","marker":"[5]"},{"why":"Supplies the pre-trained Wasserstein GAN weights used to generate BraTS synthetic tumour images.","marker":"[59]"},{"why":"Provides the latent diffusion model implementation used to generate BraTS synthetic tumour images.","marker":"[66]"}],"fun_headline_variants":["No-reference metrics blind to tumour edge blurring","Quality scores miss tumour changes that invert model rankings","No-reference IQMs fail to catch clinically relevant errors","Image quality metrics ignore tumour boundary blurring"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The downstream yardstick, the Evanyseg quality predictor, is assumed to give valid predicted Dice scores when applied to perturbed real images and to synthetic images, even though it was trained only on real images with synthetically perturbed segmentation masks.","fun_headline_variants_meta":{"raw":{"variants":["No-reference metrics blind to tumour edge blurring","Quality scores miss tumour changes that invert model rankings","No-reference IQMs fail to catch clinically relevant errors","Image quality metrics ignore tumour boundary blurring"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000369,"raw_usage":{"total_tokens":1978,"prompt_tokens":941,"completion_tokens":1037,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":978}},"tokens_in":557,"tokens_out":1037,"duration_ms":9164,"temperature":1.0,"reasoning_tokens":978,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:22:51.501033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the perturbed real BraTS images used in the morphological experiments, for which ground-truth tumour masks exist, and compute the true Dice between the segmentor's output and the ground truth; if Evanyseg's predicted Dice diverges from these true Dice as boundary blur or intensity gradient strength grows, the paper's downstream yardstick is miscalibrated.","supporting_citations":[{"cited_title":"Towards ground-truth-free evaluation of any seg- mentation in medical images, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the ground-truth-free regression method, Evanyseg, used as the downstream yardstick for segmentation quality."},{"cited_title":"GANs trained by a two time-scale update rule con- verge to a local Nash equilibrium.Adv Neural Inf Process Syst, 30, 2017","cited_arxiv_id":null,"evidence_quote":"Defines FID, the primary distance metric whose behaviour under perturbations and model outputs is under test."},{"cited_title":"Sutherland, Michael Arbel, and Arthur Gretton","cited_arxiv_id":null,"evidence_quote":"Defines KID, the polynomial-kernel MMD distance metric that is a central tested upstream metric."},{"cited_title":"Shape-guided conditional latent diffusion models for synthesising brain vasculature","cited_arxiv_id":null,"evidence_quote":"Provides the pre-trained VAE, GAN, and DDPM weights used to generate the IXI synthetic images."},{"cited_title":"IXI dataset – brain development.https: //brain-development.org/ixi-dataset/","cited_arxiv_id":null,"evidence_quote":"Supplies the IXI brain MRA dataset used for vessel experiments and for density-shift scanner mix manipulations."},{"cited_title":"Identifying the best machine learning algo- rithms for brain tumor segmentation, progression assessment, and overall survival prediction in the BRATS challenge, 2019","cited_arxiv_id":null,"evidence_quote":"Supplies the BraTS brain tumour MRI data and ground truth masks used for morphological manipulation experiments."},{"cited_title":"Learned local 21 attention maps for synthesising vessel segmenta- tions from T2 MRI","cited_arxiv_id":null,"evidence_quote":"Provides the pre-trained vessel segmentation model used as the downstream task for IXI images."},{"cited_title":"How faith- ful is your synthetic data? Sample-level metrics for evaluating and auditing generative models","cited_arxiv_id":null,"evidence_quote":"Defines the AuthPct authenticity metric whose contrasting behaviour with distance metrics exposes data memorisation."},{"cited_title":"medigan: a python li- brary of pretrained generative models for medi- cal image synthesis.Journal of Medical Imaging, 10(6):061403, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the pre-trained Wasserstein GAN weights used to generate BraTS synthetic tumour images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the latent diffusion model implementation used to generate BraTS synthetic tumour images."}],"review_version":1}