{"id":"a836f766-d044-4b58-b955-d718bb83af0c","arxiv_id":"1909.01940","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across four chest radiograph datasets, every model performs best on its own test set, and cross-dataset AUC drops are largest for models trained on ChestX-ray14 and PadChest.","lead":"This paper trains the same chest X-ray classifier on four public datasets and measures how accuracy drops when it is tested on a different dataset. It finds that models trained on CheXpert and MIMIC-CXR transfer better, while a model trained on ChestX-ray14 loses the most performance on other datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-dataset AUC gaps may measure label mismatch more than image domain shift: the eight 'common' labels are produced by different labelers, languages, and mapping rules (Section 3.2), and the paper itself cites 10-30% label-error rates for ChestX-ray14.","rationale":"The reader's weakest_assumption already identifies label equivalence as the key unvalidated premise, and the full text supports that concern: Section 3.2 describes the label merges, Section 4 discusses ChestX-ray14 labeler unreliability, and the datasets use different NLP pipelines and report languages. The central quantitative claim (same-dataset training wins and cross-dataset AUC drops can be large, e.g. ChestX-ray14-trained model on CheXpert) is clear and reproducible in principle, but the interpretation that these gaps measure imaging domain shift rather than label mismatch is not settled. This is exactly the kind of condition that should gate acceptance: the experiment is straightforward, the data are public, and a single expert-annotation study or a labeler-consistency analysis would test it. Because the reader already assigned CONDITIONAL with moderate confidence, my independent read does not move the verdict; it sharpens the condition under which the paper should be accepted as demonstrating image domain shift. I do not see an internal inconsistency or a reason to reject: the descriptive transfer results stand, and the paper is appropriately cautious in places, but the headline conclusion overreaches until the label-equivalence premise is checked.","tokens_in":7440,"tokens_out":3040,"duration_ms":30182,"concrete_test":"Select roughly 500 frontal images per dataset, stratified by original positive labels, and have at least two board-certified radiologists independently annotate the 8 common findings under one written protocol, with adjudication for disagreements. Recompute all 16 train/test AUC entries of Table 2 against these expert labels. If the same-dataset-best pattern and the ChestX-ray14-trained drop of about 0.12 on CheXpert persist, label mismatch is not the main driver; if the gaps shrink materially or reorder, the central claim must be restated as covering combined image and label/annotation shift. A secondary check is to retrain each of the four models with five seeds and report bootstrap confidence intervals, to test whether the 'best mean AUC' rankings in Table 2 are stable rather than split-dependent noise.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference is that the cross-dataset AUC drops in Table 2 quantify image domain shift. For that inference, a positive label in ChestX-ray14 must denote the same clinical finding as the same positive label in CheXpert, MIMIC-CXR, and PadChest. Section 3.2 constructs the shared label set with ad hoc merges: ChestX-ray14 'Nodule' and 'Mass' are combined into 'Lesion', while CheXpert and MIMIC-CXR have their own 'Lesion' label; PadChest subtypes such as 'Atelectasis Basal', 'Total Atelectasis', 'Lobar Atelectasis', and 'Round Atelectasis' are collapsed into 'Atelectasis'. No mapping table or validation of these merges is given. The datasets also use different NLP labelers (CheXpert and MIMIC-CXR share one; ChestX-ray14 and PadChest do not) and different report languages, and Section 4 explicitly cites evidence [17] that ChestX-ray14 labels may be 10-30% less accurate than originally reported. Under these conditions, a cross-dataset drop can be produced by label-threshold or label-semantics mismatch even if the image distribution were unchanged, and the same-dataset baseline is inflated by training and testing on the same noisy labeler. The observed transferability ranking (CheXpert/MIMIC-CXR models generalize better) is also confounded with labeler similarity and report language, so the headline attribution to imaging domain shift is not fully secured. The result remains an honest empirical finding about cross-dataset transfer, but the causal claim that image domain shift alone causes the drop is load-bearing and untested.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a controlled empirical comparison of a fixed DenseNet-121 (CheXNet) multi-label classifier trained on four public chest X-ray datasets (ChestX-ray14, CheXpert, MIMIC-CXR, PadChest) and evaluated on all four test sets. Using eight labels obtained by merging and simplifying the original label sets, it reports per-class and mean AUC in Table 2. The main findings are that same-dataset training yields the best mean AUC on every test set, and that models trained on CheXpert or MIMIC-CXR transfer better to other datasets than models trained on ChestX-ray14 or PadChest. The authors conclude that domain shift causes substantial performance loss and recommend external, case-by-case validation.","tokens_in":7749,"tokens_out":6243,"duration_ms":54104,"significance":"If the headline result is taken only as a statement about cross-dataset transfer of models as deployed, the paper is a useful empirical benchmark and a clear cautionary data point for medical imaging practice. The experimental setup is simple and transparent, the four datasets are public, and the discussion honestly acknowledges labeler-related concerns. The paper also makes a concrete, falsifiable claim about transferability ranking that other groups can test. However, the quantitative contribution is weakened by the absence of uncertainty estimates and by uncontrolled label-generation differences across datasets, so the significance is moderate rather than high.","major_comments":[{"comment":"The central attribution of the observed AUC drops to image domain shift is not fully secured because the eight shared labels are produced by different annotation pipelines. ChestX-ray14's 'Lesion' is a merge of 'Nodule' and 'Mass', PadChest subtypes are collapsed without a mapping table or validation, and the NLP labelers and report languages differ; Section 4 additionally cites evidence that ChestX-ray14 labels may be 10-30% less accurate than originally reported. A positive label therefore need not denote the same clinical finding across datasets, so the cross-dataset gaps in Table 2 can reflect label mismatch and label noise even if the image distribution were unchanged. Please validate label equivalence, or explicitly reframe the conclusions as measuring transfer under the datasets' existing label definitions rather than image domain shift.","section":"Section 3.2 and Table 2"},{"comment":"All results are single point estimates from one training run per dataset with no confidence intervals, error bars, or repeated seeds. The differences that support the transferability ranking (for example, the 0.04 mean-AUC gap between the CheXpert-trained and MIMIC-CXR-trained models on the CheXpert test set) may be within run-to-run variation. Report repeated runs or bootstrap confidence intervals before ranking source datasets.","section":"Section 4 and Table 2"},{"comment":"The mean AUC entries for two rows do not match the arithmetic mean of the eight per-class values shown in the same row. On the CheXpert test set with ChestX-ray14 training, the listed mean is 0.6821 whereas the eight listed per-class AUCs average to 0.6933; on the MIMIC-CXR test set with ChestX-ray14 training, the listed mean is 0.7406 whereas the eight values average to 0.7494. These inconsistencies affect the quantitative claims in the Discussion and should be corrected.","section":"Section 4 and Table 2"}],"minor_comments":[{"comment":"The text says the trained model is evaluated 'with images from the remaining two' datasets, but after training on each of four datasets there are three remaining test sets; please correct this to 'remaining three'.","section":"Section 3.2"},{"comment":"Several entries are typeset without spacing, such as '0.93900.6833' and 'MIMIC-CXR0.7942'; these should be formatted as separate columns.","section":"Table 2"},{"comment":"Reference [20] is cited for ChestX-ray14 but the listed title is 'ChestX-ray8'; please correct the title or the reference.","section":"References"},{"comment":"The pixel-intensity density plot lacks axis labels and units; adding them would allow readers to interpret the distributions.","section":"Figure 2"},{"comment":"The training hyperparameters (optimizer, learning rate, batch size, number of epochs, and loss function) are not stated, which limits reproducibility despite the architecture being specified; please add a short training-details paragraph.","section":"Section 3.2"},{"comment":"The caption contains a typo ('cotains') and the sentence describing the composite image is hard to parse; please revise.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript is a straightforward empirical study with an honest discussion of its own limitations, but the label-semantics confound is load-bearing for the causal claim about image domain shift, and the Table 2 arithmetic errors should be fixed before any final decision. A revision that reframes the conclusions to match what the experiment actually measures, and adds uncertainty estimates, would make the paper acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a useful measurement: training the same DenseNet121 on each of four public chest X-ray datasets and testing on all four gives a clean 4x4 AUC table, and the same-dataset-always-wins result plus the transferability ranking (CheXpert and MIMIC-CXR train the most portable models) is a concrete number people building on these benchmarks will want to cite. Second, the title claim that this quantifies 'domain shift' is only half secured. The AUC drops are real, but they mix image distribution shift with labeler mismatch.\n\nWhat's new: at the time, nobody had done a four-dataset cross-evaluation with these public chest X-ray sets. The experiments are straightforward and reproducible in principle, and the paper is honest about the label merging (Nodule/Mass to Lesion, PadChest subtype merges) and about CheXpert/MIMIC sharing a labeler while ChestX-ray14 and PadChest don't. It also cites Oakden-Rayner's 10-30% label accuracy concern. Credit where due: the paper does not hide its main confound.\n\nSoft spots, in proportion. The biggest is the label equivalence assumption. If a 'positive' means something different in each dataset, or if label noise rates differ, then the cross-dataset drop is not purely image domain shift. The stress-test note is right that this is load-bearing for the causal claim. The authors acknowledge labeler differences as 'one possible cause,' but they still frame the whole paper around domain shift and even advise dataset preference based on transferability, which is confounded with labeler similarity. Also, single runs, no confidence intervals, and randomly re-split test sets for CheXpert/MIMIC/PadChest mean the exact AUC numbers should not be over-read. The rank order is probably stable—the gaps are large—but 0.04 differences are not meaningful without error bars. Minor: the intro says 'three large datasets' when they analyze four, and Figure 1 shows 'consolidation' examples without verifying the labels.\n\nWho's it for: anyone working with public chest X-ray benchmarks, especially people deciding which dataset to pretrain or fine-tune on. It deserves a serious referee—the central observation is important enough, and the authors are clear enough, that a careful review with a request for confidence intervals or repeated runs and a sharper discussion of the label confound would make this a genuinely useful paper.","headline":"Solid four-dataset cross-evaluation of chest X-ray classifiers, but the headline attribution to image domain shift is partly confounded by label mismatch across datasets.","tokens_in":8286,"tokens_out":1862,"would_cite":true,"duration_ms":17889,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chest X-ray AI models drop sharply when used outside their training dataset.","keywords":["domain shift","chest radiograph","deep learning","multi-label classification","model generalization","external validation","medical imaging","dataset bias"],"falsifier":"A reader study that manually re-annotates a random sample of images from all four datasets using a single shared label protocol and then repeats the cross-dataset training; if the AUC gaps vanish or substantially shrink, the paper's attribution of the drops to image domain shift would be refuted.","tokens_in":7236,"feed_emoji":"🩻","tokens_out":6186,"duration_ms":54427,"temperature":0.7,"pith_summary":"This paper asks whether deep learning models for chest radiograph classification can be trusted outside the specific dataset they were trained on. By training the same network on four major public chest X-ray datasets and testing each model on all four test sets, it shows that the best mean performance always comes from same-dataset training. Cross-dataset evaluation produces large drops, e.g., a model trained on ChestX-ray14 loses 0.12 mean AUC on CheXpert compared with a model trained there. Models trained on CheXpert and MIMIC-CXR transfer best to the other datasets. The result raises a caution flag for deploying public-dataset-trained models in other hospitals.","feed_headline":"Chest X-ray AI models drop sharply outside their training dataset","feed_subtitle":"Same-dataset training wins in all four test sets; cross-dataset drops reach 0.12 mean AUC.","key_machinery":"The central object is the controlled cross-dataset evaluation protocol: the same DenseNet121 convolutional neural network, with identical hyperparameters and ImageNet pretraining, is trained separately on each of the four datasets, then evaluated on all four test sets using the eight labels common to every dataset. This isolates domain shift as the cause of performance differences, provided label semantics are comparable. The metric is mean Area Under the ROC Curve (AUC) over the eight shared findings.","core_discovery":"The central discovery is that domain shift—differences in image appearance, acquisition protocols, populations, and label generation between chest X-ray collections—substantially degrades model performance. For all four test datasets, the highest mean AUC is achieved by the model trained on that same dataset. The largest gap observed is 0.12 in mean AUC: a model trained on ChestX-ray14 scores 0.6821 on CheXpert while a CheXpert-trained model scores 0.8042. The authors additionally find that models trained on CheXpert and MIMIC-CXR generalize better to other datasets than models trained on ChestX-ray14 or PadChest, and they attribute part of the transfer difficulty to differences in label extraction methods and label noise.","pith_inferences":["If label noise and label extraction differences are the main drivers of the gap, then the measured 'domain shift' may partly be 'annotation shift'; a direct measurement using manually re-labeled images across datasets could separate image-level shift from label-level shift.","The authors' finding that CheXpert and MIMIC-CXR generalize better might reflect their use of a shared labeler and larger training sizes; a follow-up that matches training set sizes and label distributions could test whether the advantage is intrinsic to those datasets or an artifact of scale.","The same cross-dataset protocol could be extended to other imaging modalities and tasks, such as CT, MRI, or segmentation, to see whether the ranking of dataset representativeness holds beyond chest X-rays."],"forward_implications":["Clinicians and regulators should expect reported accuracy on a source dataset to overstate real-world performance when the deployment population or imaging equipment differs.","Researchers developing chest X-ray classifiers should prefer CheXpert and MIMIC-CXR as training sources, since models trained on them retain more performance across other datasets.","A case-by-case external validation strategy is necessary: models should be validated or fine-tuned on small local datasets from the specific machines and settings where they will be used.","The observed performance gaps are not uniform across findings, so per-finding external evaluation is warranted rather than relying only on a single averaged metric."],"supporting_citations":[{"why":"Supplies the DenseNet121 architecture and training protocol reproduced for all four models.","marker":"[18]"},{"why":"Provides the ChestX-ray14 dataset, its 14 labels, and the original split used in the experiments.","marker":"[20]"},{"why":"Provides the CheXpert dataset, its labels, and the uncertainty labeling scheme treated as negatives.","marker":"[11]"},{"why":"Provides the MIMIC-CXR dataset and labels, which share the CheXpert labeler.","marker":"[13]"},{"why":"Provides the PadChest dataset with 174 labels that are later merged into the common findings.","marker":"[2]"},{"why":"Reports label reliability issues in ChestX-ray14, supporting the label-noise explanation for transfer difficulty.","marker":"[17]"},{"why":"Defines the concept of dataset bias and domain shift that motivates the evaluation.","marker":"[19]"}],"fun_headline_variants":["Chest X-ray AI: cross-dataset AUC drop hits 0.12","Domain shift costs chest X-ray AI 0.12 AUC points","Chest X-ray AI models lose up to 0.12 AUC on new data","Same-dataset training beats cross-dataset in chest X-ray AI","Training data choice is key for chest X-ray AI accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the eight labels mean the same clinical finding in every dataset; different NLP labelers, report languages, and merged label definitions may make the measured gaps reflect label mismatch as well as image domain shift.","fun_headline_variants_meta":{"raw":{"variants":["Chest X-ray AI: cross-dataset AUC drop hits 0.12","Domain shift costs chest X-ray AI 0.12 AUC points","Chest X-ray AI models lose up to 0.12 AUC on new data","Same-dataset training beats cross-dataset in chest X-ray AI","Training data choice is key for chest X-ray AI accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000253,"raw_usage":{"total_tokens":1521,"prompt_tokens":862,"completion_tokens":659,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":478,"completion_tokens_details":{"reasoning_tokens":563}},"tokens_in":478,"tokens_out":659,"duration_ms":6093,"temperature":1.0,"reasoning_tokens":563,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:24:29.206760+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader study that manually re-annotates a random sample of images from all four datasets using a single shared label protocol and then repeats the cross-dataset training; if the AUC gaps vanish or substantially shrink, the paper's attribution of the drops to image domain shift would be refuted.","supporting_citations":[{"cited_title":"In: Proceedings of the Conference on Computer Vision and Pattern Recognition","cited_arxiv_id":null,"evidence_quote":"Defines the concept of dataset bias and domain shift that motivates the evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the PadChest dataset with 174 labels that are later merged into the common findings."}],"review_version":1}