{"id":"4e19840a-66a4-4588-a915-6944c4464c66","arxiv_id":"2607.26333","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Chest X-ray AI model rankings and image-quality metric rankings change substantially with the choice of evaluation reference, so benchmark scores are not neutral.","lead":"Chest X-ray AI rankings flip depending on whether models are scored against labels extracted from radiology reports or from experts reading images alone, and common image-quality metrics like PSNR often disagree with clinicians. A new paired expert-labeled dataset quantifies these shifts, showing the evaluation reference can determine which model is selected for deployment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Image-derived reference may be dominated by a single annotator; ranking instability may partly reflect annotation noise rather than systematic label-source differences.","rationale":"The reader's weakest assumption identified the validity of paired image-derived labels as the key load-bearing point. I agree, and sharpen it: the risk is not just that image-derived labels are noisy in the abstract, but that for some pathologies the 'conflict-resolved' reference may effectively be a single annotator's labels, and the MIMIC-CXR reference uses an unfused proxy whose transferability is not established. This threat is concrete because the reported inter-annotator κ for Atelectasis (0.15) and the exact count match between CR and Img2 show the adjudication process can be dominated by one reader. If the ranking instability is primarily a consequence of this annotator noise, the paper's conclusion that references encode different clinical information is weakened, though the more general point—that evaluation choices affect rankings—would still stand. The proposed test is feasible with already-collected labels and would directly estimate how much of the reference effect is annotator-driven. The paper otherwise has many strengths: paired annotations, two datasets, bootstrap CIs for ranking correlations, sensitivity analyses, and an explicit limitations section. These do not fully resolve the concern, but they support a CONDITIONAL rather than a harsher verdict. Therefore the reader's CONDITIONAL verdict remains appropriate, and I do not recommend changing it.","tokens_in":51748,"tokens_out":4161,"duration_ms":46684,"concrete_test":"Using the existing XQA-RR data, compute SRCC between model rankings induced by Img1 alone and by Img2 alone for each of the four selected pathologies, and compare these within-image-source SRCCs with the reported CR-vs-Report SRCCs (e.g., Atelectasis all-models SRCC ≈0.80–0.87). If the Img1-vs-Img2 ranking SRCC is comparable to or lower than the CR-vs-Report SRCC, the reference-effect is not separable from inter-annotator noise, and the central claim would need to be reframed. A confirmatory step on MIMIC-CXR would be to have a third radiologist adjudicate a random 200-image subset and compare FUSE_OR rankings to adjudicated-label rankings, but the XQA-RR analysis is the direct, data-available check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that evaluation-reference choice systematically changes model rankings and conclusions—depends on the image-derived labels being a coherent operationalization of expert image-only judgment, not just one annotator's idiosyncratic readings. Section 3.1.1 reports very low image inter-annotator agreement (Atelectasis κ=0.15 on XQA-RR; Lung opacity κ=0.11 on MIMIC-CXR), and for Atelectasis the conflict-resolved (CR) label counts (165 positives) exactly match Img2 (165) rather than Img1 (44), suggesting adjudication may have simply followed one annotator. On MIMIC-CXR, FUSE_OR is used without adjudication; although validated against CR on XQA-RR (Table F.30), the MIMIC inter-annotator agreement profile differs substantially, so that validation may not transfer. If model rankings induced by the two image annotators separately are about as different as rankings induced by CR versus report labels, then the observed 'reference effect' could be substantially driven by annotator noise rather than by a structural difference between image-derived and report-derived clinical information. The limitations section acknowledges single-institution annotation, but this is load-bearing because the paper's interpretation as systematic, clinically interpretable divergence relies on the image reference being a stable construct. The authors' own low inter-annotator κ values and the one-annotator pattern in CR for Atelectasis are explicit evidence that this condition is not fully secure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper asks whether evaluation-reference choices—image-derived vs report-derived pathology labels, and generic IQA metrics vs expert ratings of diagnostic usability—change model rankings and hence scientific conclusions in chest X-ray machine learning. The authors introduce XQA-Chest, a paired dataset of expert image- and report-derived labels for 650 CUH studies, plus expert IQA ratings for 1,571 degraded images, and they annotate a 735-study MIMIC-CXR subset under similar protocols. They evaluate a broad set of supervised CNNs and vision-language models under different label references, measure agreement and rank correlations (SRCC/KRCC with bootstrap CIs), and similarly compare IQA metrics against expert z-MOS rankings. The central finding is that changing the evaluation reference can reorder models, sometimes dramatically (e.g., VLM-FT rankings for Atelectasis are nearly uncorrelated across references), and that commonly used IQA metrics such as PSNR and CW-SSIM align only weakly with expert judgments.","tokens_in":52049,"tokens_out":8097,"duration_ms":78298,"significance":"If the central claim holds, the paper makes an important methodological contribution: benchmark rankings in CXR machine learning are reference-dependent, and the choice of evaluation reference should be treated as a core component of clinical validity rather than an implementation detail. The paper's strengths include a new paired dataset with expert annotations, extensive agreement statistics with pessimistic/optimistic uncertainty mappings, bootstrap confidence intervals for pathology rank correlations, sensitivity analyses for label-fusion methods, and a qualitative radiology review of disagreements. These are appropriate tools for the claim and provide a useful resource for the community. The main risk is that the image-derived reference is not shown to be a stable construct independent of individual annotator noise; if the apparent reference effect is substantially driven by image-annotator disagreement, the interpretation as a systematic image-vs-report distinction would need to be tempered.","major_comments":[{"comment":"The conflict-resolved (CR) image labels for Atelectasis (165 positives), Cardiomegaly (21), and Support Devices (274) exactly match Img2's counts, while Img1's counts are very different (44, 153, 309). This strongly suggests that adjudication often deferred to one annotator rather than producing a genuine synthesis. The paper does not report model rankings induced by Img1 and Img2 separately. Without these controls, the central claim that changing the label source (image vs report) systematically changes rankings is confounded with image-annotator noise: the image-image rank correlation may be as low as the image-report rank correlation. For the VLM-FT group on Atelectasis, the CR-vs-report SRCC is 0.21±0.22 (Table E.26); if Img1-vs-Img2 SRCC is similar, the effect is not specific to the image/report distinction. Please add per-annotator ranking analyses for XQA-RR (and MIMIC where possi","section":"§3.1.1, Table A.5"},{"comment":"FUSE_OR is validated against CR on XQA-RR and then used as the image-derived reference on MIMIC-CXR, where no adjudication is available. This transfer is not fully supported. First, the high FUSE_OR-vs-CR agreement is partly a consequence of the CR pattern noted above: for pathologies where CR equals one annotator, an optimistic union rule will naturally reproduce the more positive annotator. Second, the MIMIC inter-annotator agreement profile differs substantially from XQA-RR (e.g., Lung opacity κ=0.11 on MIMIC vs 0.42 on XQA-RR; Table B.12 vs B.9), so a fusion rule validated on one institution's disagreement structure may behave differently elsewhere. The MIMIC comparisons labeled 'image-derived vs report-derived' are therefore better described as comparisons between a particular optimistic fusion rule and manually annotated report sections. Please add sensitivity analyses using Img1 a","section":"§2.1.3, Table F.30"},{"comment":"The IQA rank correlations are reported as point estimates without confidence intervals, in contrast to the pathology-ranking results, which use 1000-sample bootstrap CIs. With 1,571 images, the differences among the better-performing metrics (e.g., HaarPSI 0.87, GMSD 0.86, HaarPSImed 0.85, MS-SSIM 0.84) may be within sampling variability. The paper's IQA conclusion—that metric selection determines which processed images appear best—depends on these metric-level differences being reliable. Please provide bootstrap CIs (or paired significance tests) for the SRCC/KRCC values in Table 4.","section":"§3.2, Table 4"}],"minor_comments":[{"comment":"The abstract and Discussion state that SSIM and PSNR 'often fail to align' with expert judgment, but Table 4 gives SSIM-matlab SRCC=0.78, which is moderate agreement. Consider rephrasing to avoid overstatement, e.g., 'SSIM shows weaker alignment than top metrics, and PSNR aligns weakly.'","section":"Abstract, §4"},{"comment":"The handling of 'not mentioned' report labels is not stated. The tables list explicit negative, uncertain, and positive counts, while the positive prevalences in Table A.5 imply that 'not mentioned' was collapsed to negative in the binary analyses. This is a standard CheXpert-like assumption but should be stated explicitly, since it affects all report-derived agreement and ranking results.","section":"§2.1.2, Tables A.7/A.8"},{"comment":"The symbols ✓/✗ in Table 2 are not defined consistently; for example, the 'Images' and 'Reports' columns use different combinations. Add a legend or use explicit 'Yes/No' entries.","section":"Table 2"},{"comment":"The text describes a 14-category schema and then excludes 'No finding' and 'Support devices' to obtain 12 pathologies. State this exclusions before listing the 14 categories to avoid apparent inconsistency.","section":"§2.1.2"},{"comment":"The x-axis labels in Figures 10b–14b use the order 'Impression, FUSE_OR, Findings' while the text and the MIMIC tables usually list 'FUSE_OR, Findings, Impressions.' Use a consistent order for readability.","section":"Figures 10b–14b"},{"comment":"The Discussion states that the observed divergence is 'systematic, clinically interpretable.' Given the single-institution annotation and the CR=Img2 pattern noted above, 'systematic' is too strong; please qualify this claim in light of the per-annotator analysis requested in the major comments.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"This is a valuable dataset and a generally well-conducted study. The stress-test concern about annotator noise is real and load-bearing: the CR labels for several key pathologies appear to equal one annotator's labels, and the paper does not provide the per-annotator ranking control needed to separate label-source effects from annotation-noise effects. I recommend major revision rather than rejection because the central claim is defensible and the missing analyses are well within scope of the existing data. I would also ask the editor to ensure the IQA confidence-interval gap is addressed before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. This is a careful, large-scale empirical study of a question that usually gets ignored: does the choice of evaluation labels—image-derived vs report-derived, or Findings vs Impressions—actually change which CXR models look best? The answer is yes, and the paper shows it systematically across more than 30 models, with bootstrap confidence intervals on ranking correlations. The new paired dataset (XQA-Chest), with independent expert annotations of images and reports for the same studies plus IQA ratings, is a genuine resource. The IQA half is also useful: common metrics like PSNR and SSIM correlate poorly with radiologists' judgments of diagnostic usability, which deserves emphasis in the medical imaging community.\n\nThe paper does not overclaim. It frames the references as answering different questions rather than declaring one true, and the limitations section is candid about single-institution annotation, small sample size for rare pathologies, and synthetic degradations in IQA. The qualitative review of disagreements gives a plausible clinical explanation for the label divergence, and the phenomenon appears across multiple pathologies and also in the Findings-vs-Impressions comparison, which is independent of image-annotator noise.\n\nThe main soft spot is exactly what the stress-test flags. Image inter-annotator agreement is low for Atelectasis (κ=0.15), and the conflict-resolved labels match one annotator's positive counts exactly (165 vs 44). That raises the possibility that some of the ranking instability under image-derived labels is driven by one annotator's idiosyncratic readings rather than a systematic image-vs-report distinction. The paper doesn't report model rankings under each image annotator separately, so we can't tell how much of the effect is annotation noise. This is a moderate concern, not fatal—the conclusions would be weaker for Atelectasis specifically, but the broader pattern across pathologies and report sections stands. The authors should run that analysis before the dataset becomes a community benchmark.\n\nMinor issues: the dataset isn't public yet (pending ReShare review), and Table 4 reports IQA correlations without uncertainty intervals. Both are fixable.\n\nFor whom: anyone building or evaluating CXR classifiers or reconstruction methods, and anyone who uses benchmarks to pick models. This deserves a serious referee—send it to review. I'd cite it if I were writing about evaluation methodology in medical imaging.","headline":"Solid, data-rich empirical study showing evaluation-reference choice can reorder CXR model rankings; the noisy image-reference is a real caveat but doesn't overturn the core message.","tokens_in":52602,"tokens_out":2262,"would_cite":true,"duration_ms":25014,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chest X-ray model rankings depend heavily on whether labels come from images or radiology reports, and on which image-quality metric is used, so evaluation references should be chosen as part of clinical validity, not as an afterthought.","keywords":["chest X-ray","evaluation reference","report-derived labels","image-derived labels","image quality assessment","model ranking","clinical validity","paired dataset"],"falsifier":"An independent multi-institution annotation study with adjudicated image labels that found model rankings under image- and report-derived references to be highly correlated (e.g., SRCC > 0.9) for Atelectasis and Consolidation would directly contradict the claim that reference choice flips rankings. Alternatively, demonstrating that the ranking instability disappears when image labels are produced by a larger, more diverse panel of radiologists would indicate the effect is annotation noise rather than a structural property of the label sources.","tokens_in":51620,"feed_emoji":"🩻","tokens_out":3670,"duration_ms":32865,"temperature":0.7,"pith_summary":"The paper establishes that the choice of evaluation reference is not a technical detail: in chest X-ray machine learning, swapping image-derived labels for report-derived labels, or swapping a generic image-quality metric for expert judgment, can change which models and methods are judged best. The authors built a paired dataset in which expert radiologists independently labeled the same studies from the image alone and from the radiology report, and rated degraded images for diagnostic usability. They show that for pathologies such as Atelectasis and Consolidation, the top-ranked model under one reference can fall far down the ranking under another, and that commonly used quality metrics like PSNR and SSIM correlate weakly with expert assessments. The conclusion is that evaluation-reference selection should be treated as a central component of clinical validity in CXR machine learning, justified per pathology, task, and intended use.","feed_headline":"Swap the labels, swap the winning CXR model","feed_subtitle":"New paired dataset shows image- and report-derived references rank models differently; evaluation needs clinical justification.","key_machinery":"The central object is the paired evaluation-reference dataset XQA-Chest, with its image-only and report-only annotation protocols and an adjudicated (or automatically fused) consensus label. Rank correlation coefficients (SRCC, KRCC) and Cohen's kappa carry the argument by quantifying how label-source disagreement translates into model-ranking instability and how IQA metrics align with expert diagnostic-usability rankings.","core_discovery":"Using 650 paired studies from a clinical cohort and 735 from a public dataset, the paper shows that image-derived and report-derived pathology labels disagree in a pathology-dependent way, and that this disagreement propagates to model rankings: for fine-tuned vision-language models, rankings under the two label sources can be nearly uncorrelated for Atelectasis (SRCC ≈ 0.21), while Pleural effusion remains stable. In image quality assessment, expert rankings of diagnostic usability correlate strongly with some full-reference metrics (HaarPSI, GMSD) but weakly with PSNR, CW-SSIM, and no-reference metrics. The authors conclude that image- and report-derived labels answer different clinical qu","pith_inferences":["If reference-dependence generalizes, single-reference benchmarks in medical imaging can mislead model selection; multi-reference evaluation may become standard practice for clinically meaningful comparisons.","The low image inter-annotator agreement for some pathologies suggests image-derived labels are themselves noisy; a multi-institution adjudicated study could determine whether ranking instability is structural or partly annotation noise.","The IQA results imply that clinically validated, task-specific quality metrics could be developed directly from expert preference data rather than borrowed from generic perceptual metrics."],"forward_implications":["Benchmark leaderboards in CXR classification are reference-dependent; a model's best status is only meaningful relative to the chosen label source.","Evaluation against report-derived labels is internally reproducible but can diverge from image-derived evaluation; studies should state which report section (Findings vs Impression) is used.","IQA metrics used for reconstruction or generative models should be validated against expert diagnostic-usability ratings before deployment, since PSNR and no-reference metrics can misrank outputs.","The paired-references resource enables future work to test label-fusion methods and automated report-extraction strategies against a common expert benchmark."],"fun_headline_variants":["Chest X-ray model leaderboard depends on label source","Image vs report labels flip CXR model rankings","Common CXR quality metrics ignore expert judgment","CXR winner? Depends on your evaluation reference"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paired image-derived labels (conflict-resolved on the clinical cohort, FUSE_OR-fused on the public dataset) accurately represent expert image-only judgment, and their disagreement with report-derived labels reflects systematic differences rather than idiosyncratic annotation noise.","fun_headline_variants_meta":{"raw":{"variants":["Chest X-ray model leaderboard depends on label source","Image vs report labels flip CXR model rankings","Common CXR quality metrics ignore expert judgment","CXR winner? Depends on your evaluation reference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000763,"raw_usage":{"total_tokens":3251,"prompt_tokens":802,"completion_tokens":2449,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":2397}},"tokens_in":546,"tokens_out":2449,"duration_ms":16907,"temperature":1.0,"reasoning_tokens":2397,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:07:12.983768+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent multi-institution annotation study with adjudicated image labels that found model rankings under image- and report-derived references to be highly correlated (e.g., SRCC > 0.9) for Atelectasis and Consolidation would directly contradict the claim that reference choice flips rankings. Alternatively, demonstrating that the ranking instability disappears when image labels are produced by a larger, more diverse panel of radiologists would indicate the effect is annotation noise rather than a structural property of the label sources.","supporting_citations":[],"review_version":1}