{"id":"335ff409-cde4-437e-a8d9-b79511ce8fe8","arxiv_id":"2506.07984","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"CXR-LT 2024 provides a new large chest X-ray benchmark with 45 labels and three tasks, and reports that top models achieve mAP of 0.28 to 0.53 on long-tailed tasks but only 0.11 to 0.13 on zero-shot unseen diseases.","lead":"This paper reports on the CXR-LT 2024 MICCAI challenge, which expanded a chest X-ray benchmark to 377,110 images and 45 disease labels, including a new zero-shot task on five unseen diseases. The challenge results show that current models reach only modest mAP, especially for rare and unseen findings, highlighting how far automated reading of chest X-rays still has to go.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task 3 zero-shot test labels for the five unseen classes are RadText-generated with no manual validation; if those labels are noisy, the reported mAP 0.129 cannot support the claim that zero-shot detection is unsolved.","rationale":"The reader's weakest assumption identified automatically extracted labels from RadText as the key reliability threat, which is the same underlying issue. This stress test narrows the concern to the five zero-shot classes in Task 3, where the label noise directly contaminates the evaluation rather than only the training signal, and where no manual gold-standard validation exists at all. The paper's Task 2 gold-standard set does not cover these classes, and Table 8's label-quality comparison does not include them, so the central quantitative result for zero-shot performance rests entirely on unvalidated automatic labels. A manual audit of even a few hundred reports would settle whether the reported mAP of 0.129 reflects model failure or label noise. The concern is not that the authors were careless; they transparently discuss label noise and provide the benchmark openly. It is that the load-bearing number for the zero-shot claim has no independent check. Other issues raised by the reader, such as graduate-student annotation of the gold standard and the precision/mAP wording slip in Section 3.6, are less central to the benchmark's key claim. The verdict remains conditional because the proposed label audit is a reasonable condition before the zero-shot benchmark conclusion is treated as established, but no change to the reader's verdict is needed.","tokens_in":22571,"tokens_out":3480,"duration_ms":42166,"concrete_test":"Manually annotate a stratified random sample of 1,000 reports (or the largest feasible subset) from the Task 3 test partition for the five unseen classes, with each report independently labeled by two radiologists and disagreements resolved by consensus. Compute per-class precision and recall of the RadText-derived test labels against this consensus. If average precision or recall falls below 0.7 for any class, recompute the Task 3 mAP on the corrected labels; if the absolute mAP or the ranking of the top three teams changes materially, the zero-shot result must be reported as conditional on label quality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that CXR-LT 2024 is a valid benchmark, and specifically that Task 3 shows zero-shot detection of unseen findings is far from solved, depends on the correctness of the five zero-shot test labels (Bulla, Cardiomyopathy, Hilum, Osteopenia, Scoliosis). These classes are among the 19 new findings whose labels were extracted from MIMIC-CXR radiology reports using RadText (Section 2.2.1). The gold-standard set covers only the 26 CXR-LT 2023 classes (Section 2.2.2), so none of the five unseen classes has any manually validated labels. The paper's own label-noise analysis (Table 8) is limited to the 26 gold-standard classes and reports rule-based precision of only 0.711, not for the zero-shot classes. For Task 3, the evaluation labels themselves are noisy text-mined labels: the leaderboard mAP of 0.129 measures agreement with RadText extractions, not necessarily the ability to detect the actual unseen radiographic findings. Classes such as Hilum (an anatomical region rather than a disease) are especially prone to extraction errors in either direction. If the zero-shot test labels have low precision or recall, the absolute mAP and the conclusion that zero-shot detection is far from solved are not securely established, even though relative ranking among teams may survive. The paper acknowledges label noise as a general limitation (Section 4.2), but that acknowledgement does not address the specific absence of any validation for the five classes that carry the zero-shot claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports the organization and results of the CXR-LT 2024 MICCAI challenge, which comprises three tasks: long-tailed multi-label classification on a 40-class noisy test set, long-tailed classification on a 26-class manually annotated gold-standard set, and zero-shot classification of five unseen findings. The authors describe the construction of the dataset from MIMIC-CXR with RadText-extracted labels, summarize the top nine submitted solutions, report the leaderboard scores (mAP 0.281 for Task 1, 0.526 for Task 2, and 0.129 for Task 3), and discuss methodological themes such as ensembling, loss re-weighting, vision-language models, and synthetic data. The paper also includes a comparison of rule-based and GPT-4o labeling on the 26-class gold standard set.","tokens_in":22862,"tokens_out":5228,"duration_ms":61028,"significance":"If the results are taken at face value, the challenge is a valuable community resource: it extends the CXR-LT benchmark to 377,110 images and 45 classes, introduces a new zero-shot task, provides per-class performance breakdowns, and makes leaderboards and system descriptions publicly available. The observation that zero-shot performance is substantially lower than supervised performance is useful and plausible. The paper also has concrete strengths: a clear evaluation protocol, a public challenge infrastructure, and a direct comparison of rule-based versus LLM-based label extraction. However, the central zero-shot conclusion rests on labels for the five unseen classes that have no manual validation, so the significance is conditional on additional evidence about label quality.","major_comments":[{"comment":"The five Task 3 test classes (Bulla, Cardiomyopathy, Hilum, Osteopenia, Scoliosis) were labeled by RadText from radiology reports, and the gold-standard set described in Section 2.2.2 covers only the 26 CXR-LT 2023 classes. Because no subset of Task 3 labels has manual validation, the reported mAP of 0.129 could substantially reflect agreement with noisy text-mined labels rather than radiographic detection, especially for Hilum, which is an anatomical region rather than a disease. The paper should either provide a manually validated subset for the five unseen classes or explicitly reframe the Task 3 result as performance against unvalidated RadText labels and soften the conclusion that zero-shot detection of unseen findings is far from solved.","section":"Sections 2.2.1, 2.2.2, Table 7"},{"comment":"The text states that Team E placed second with mAP 0.511 and Team A placed third with 0.511, but Table 6 lists Team A second with 0.519 and Team E third with 0.511. Since the challenge rankings are central results, this inconsistency must be resolved and the correct ranking and scores reported consistently throughout the paper.","section":"Section 3.4 and Table 6"},{"comment":"The text claims that GPT-4 produces higher-quality labels \"as evidenced by improved mAP,\" but Table 8 reports micro-precision (and Section 3.6 correctly describes precision), not mAP. As presented, the table supports an improvement in precision (0.711 vs. 0.786), not in mAP. The claim that LLM-based labels improve benchmark quality should be reworded to refer to precision, or the authors should report an mAP-based comparison.","section":"Section 4.2 and Table 8"}],"minor_comments":[{"comment":"The caption says the dataset was formed by adding \"12 new clinical findings (red)\" but Section 2.2.1 states that 19 new findings were added; the legend also appears to say \"Added in CXR-LT 2025,\" which should presumably be 2024.","section":"Figure 1 caption"},{"comment":"The text refers to \"mAUC\" when describing Team C's tie-break; elsewhere the paper uses mAUROC, so the terminology should be unified.","section":"Section 3.3"},{"comment":"The table title says \"micro-precision\" while the text says \"precision\"; use one term consistently and define how the per-class values are aggregated.","section":"Section 3.6 and Table 8"},{"comment":"For Task 3, the development set is listed with 5 labels, but Section 2.1 says participants were provided labels only for the training set; clarify whether these development labels were used solely for leaderboard scoring.","section":"Table 1"},{"comment":"The definition of mAP as \"macro-averaged AP across classes\" is clear, but the paper should state explicitly whether AP is averaged per class without weighting by class frequency.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unvalidated zero-shot labels; the paper acknowledges label noise, but the central zero-shot conclusion is overstated without a small manual validation subset for the five unseen classes. The Task 2 ranking inconsistency and the precision-versus-mAP mismatch in the GPT-4o comparison should be corrected before publication. The challenge itself is a useful contribution and the manuscript is appropriate for the venue if these issues are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: CXR-LT 2024 is a solid challenge report, not a methods paper. The new asset is the expanded dataset (377k images, 45 labels, 19 new rare findings) plus a zero-shot task on five unseen classes. The headline result is the best team hitting mAP 0.129 on those five classes, which the paper reads as 'zero-shot detection is far from solved.' That reading is probably right, but the test labels for those five classes come from RadText text-mining with no manual validation, so the absolute number is shakier than the relative ordering.\n\nWhat the paper does well: it documents task design, data curation, per-class results, and top-team methods clearly, and it owns its limitations—label noise, and a gold standard annotated by graduate students rather than radiologists. That honesty is worth a real read.\n\nSoft spots, in order of size. First, the zero-shot evaluation: the five classes (Bulla, Cardiomyopathy, Hilum, Osteopenia, Scoliosis) are among the 19 new findings whose labels were extracted with RadText, and the gold-standard set covers only the original 26 classes. No manual check touches any of the five zero-shot test labels. The paper's own label-noise analysis reports rule-based precision of 0.711 on the gold-standard set, which does not include the zero-shot classes. So the mAP 0.129 measures agreement with text-mined labels, not necessarily with ground truth. That caveat is missing from the Task 3 discussion. The relative ranking among teams likely still means something, and the broad conclusion that zero-shot CXR disease detection is unsolved is consistent with other evidence, but the absolute number is not secure.\n\nSecond, the gold standard itself is graduate-student review of report text, not radiologist consensus. The paper admits this and proposes future re-annotation, which is fine, but it caps the weight you can put on Task 2 numbers.\n\nThird, minor: Table 8 is labeled 'precision' but the discussion says 'improved mAP'—a wording slip. And the paper does not give explicit dataset access instructions, though they are presumably on the challenge site/GitHub.\n\nCitation pattern is fine; self-citations to RadText, CheXFusion, and the previous CXR-LT papers are legitimate and not load-bearing.\n\nWho's this for? Anyone building or evaluating long-tailed or zero-shot chest X-ray classifiers. It is a benchmark resource, not a methodological advance. It deserves a serious referee: the dataset release and the challenge results are usable, and the zero-shot label concern is addressable with explicit qualification or a small manually validated subset. I would engage with it.","headline":"A genuinely useful benchmark update with an honest write-up, but the zero-shot result rests on unvalidated text-mined labels, so treat the absolute mAP with caution.","tokens_in":23516,"tokens_out":3813,"would_cite":true,"duration_ms":39870,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper describes the second edition of a community benchmark challenge for classifying diseases from chest X-rays, now with 377,110 images, 45 disease labels, and a new zero-shot task on five unseen findings.","keywords":["chest X-ray","long-tailed classification","multi-label classification","zero-shot learning","disease detection","benchmark","label noise","vision-language models"],"falsifier":"Select a random sample of the new rare findings from the training set and have a panel of radiologists re-read the original reports; if agreement with the parsed labels is low for those classes, the benchmark's training signal and evaluation for rare diseases are not trustworthy.","tokens_in":22397,"feed_emoji":"🩻","tokens_out":5147,"duration_ms":52861,"temperature":0.7,"pith_summary":"The paper describes the second edition of a community benchmark challenge for classifying diseases from chest X-rays, now with 377,110 images, 45 disease labels, and a new zero-shot task on five unseen findings. The reported results place the best models at 0.281 mean average precision on a 40-class noisy test set, 0.526 on a 26-class manually labeled gold standard, and 0.129 on the five unseen classes. The authors are trying to establish a shared testbed that quantifies how far long-tailed and zero-shot chest X-ray diagnosis has come, and to show that rare and unseen findings remain the biggest gap.","feed_headline":"Zero-shot chest X-ray diagnosis still weak: top mAP 0.13","feed_subtitle":"New 377k-image challenge adds 19 rare findings; best model misses most unseen diseases.","key_machinery":"The load-bearing machinery is the long-tailed label distribution generated by automatically parsing MIMIC-CXR radiology reports with RadText, together with a manually annotated gold standard subset and macro-averaged mAP as the primary metric. This setup separates the effect of label noise from the effect of class rarity, and the gold standard subset provides a human-verified check on the noisy test set.","core_discovery":"The central claim is that CXR-LT 2024 provides a valid benchmark for long-tailed, multi-label, and zero-shot chest X-ray disease classification, and that current state-of-the-art models, while making progress on common findings, perform poorly on rare and unseen ones. Evidence for this is the gap between the 0.28 mAP on the full 40-class test set and the 0.53 mAP on the 26-class gold standard, plus the 0.13 mAP on the five zero-shot classes. The paper also argues that ensemble methods, loss re-weighting, and synthetic data help tail classes, and that vision-language models are the emerging tool for zero-shot generalization.","pith_inferences":["A natural next step would be to have attending radiologists re-annotate the gold standard subset, since the current annotations were made by graduate students reading report text; if rankings shift, the noisy test set may be less predictive for rare classes.","The low zero-shot mAP suggests that text-only descriptions of unseen diseases may be insufficient; grounding vision-language models on anatomical knowledge or visual exemplars could be a testable improvement.","Because Task 1 and Task 2 mAP were highly correlated, the noisy test set may suffice for ranking on common findings, but this correlation should be checked specifically for the 19 new rare classes."],"forward_implications":["The addition of 19 rare findings drops top mAP to 0.28, showing that rare diseases are the current bottleneck in chest X-ray classification.","Zero-shot classification of five unseen findings hovers near 0.13 mAP, implying that models cannot yet generalize to novel radiological abnormalities.","GPT-4-based labeling reached higher precision than the rule-based parser on the gold standard, suggesting LLM-based pipelines could reduce label noise at scale.","The benchmark offers a reusable public dataset of 377,110 images and 45 labels, enabling future comparisons on long-tailed and zero-shot chest X-ray classification."],"supporting_citations":[{"why":"Supplies the base MIMIC-CXR images and reports that the challenge dataset is built from.","marker":"[14]"},{"why":"Provides the RadText tool used to parse radiology reports into disease labels.","marker":"[17]"},{"why":"Describes the CXR-LT 2023 challenge and the gold standard annotation protocol reused in Task 2.","marker":"[2]"},{"why":"Fleischner glossary is a source for selecting the 19 new clinical findings.","marker":"[16]"},{"why":"Provides the MIMIC-CXR-JPG image format used to distribute the dataset.","marker":"[18]"},{"why":"Supplies the GPT-4 prompting approach compared against rule-based labeling in Table 8.","marker":"[47]"}],"fun_headline_variants":["CXR-LT 2024: 377k X-rays, 19 rare diseases, still 0.13 zero-shot mAP","Rare disease X-ray AI fails: zero-shot mAP only 0.13","New chest X-ray benchmark exposes weak rare disease detection","Long-tail chest X-ray challenge: gains on common, fails on rare","Zero-shot chest X-ray gap: 0.13 mAP on unseen diseases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's reliability depends on the automatically extracted labels from radiology reports; if those labels are substantially wrong, especially for the 19 new rare findings, then the reported per-class results are unreliable.","fun_headline_variants_meta":{"raw":{"variants":["CXR-LT 2024: 377k X-rays, 19 rare diseases, still 0.13 zero-shot mAP","Rare disease X-ray AI fails: zero-shot mAP only 0.13","New chest X-ray benchmark exposes weak rare disease detection","Long-tail chest X-ray challenge: gains on common, fails on rare","Zero-shot chest X-ray gap: 0.13 mAP on unseen diseases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1449,"prompt_tokens":997,"completion_tokens":452,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":341}},"tokens_in":613,"tokens_out":452,"duration_ms":5256,"temperature":1.0,"reasoning_tokens":341,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:20:30.363049+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of the new rare findings from the training set and have a panel of radiologists re-read the original reports; if agreement with the parsed labels is low for those classes, the benchmark's training signal and evaluation for rare diseases are not trustworthy.","supporting_citations":[{"cited_title":"Fleischner society: glossary of terms for thoracic imaging.Radiology, 246(3): 697–722, 2008","cited_arxiv_id":null,"evidence_quote":"Fleischner glossary is a source for selecting the 19 new clinical findings."},{"cited_title":"Enhancing disease detection in radiology reports through fine-tuning lightweight LLM on weak labels","cited_arxiv_id":"2409.16563","evidence_quote":"Supplies the GPT-4 prompting approach compared against rule-based labeling in Table 8."}],"review_version":1}