Pith. sign in

REVIEW 26 cited by

CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.09167 v3 pith:LYV7NO4E submitted 2020-04-20 cs.CL cs.IRcs.LG

classification cs.CLcs.IRcs.LG
keywords annotationslabelingreportexpertmedicalbertchexbertlabeler
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The extraction of labels from radiology text reports enables large-scale training of medical imaging models. Existing approaches to report labeling typically rely either on sophisticated feature engineering based on medical domain knowledge or manual annotations by experts. In this work, we introduce a BERT-based approach to medical image report labeling that exploits both the scale of available rule-based systems and the quality of expert annotations. We demonstrate superior performance of a biomedically pretrained BERT model first trained on annotations of a rule-based labeler and then finetuned on a small set of expert annotations augmented with automated backtranslation. We find that our final model, CheXbert, is able to outperform the previous best rules-based labeler with statistical significance, setting a new SOTA for report labeling on one of the largest datasets of chest x-rays.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation

    cs.CV 2026-01 conditional novelty 7.0 of 10

    AnatomiX, a two-stage anatomy-first multimodal LLM for chest X-ray interpretation, reports >25% relative gains on anatomy grounding and grounded captioning, but some aggregate benchmark numbers are internally inconsis...

  2. Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    DPO-Clin improves medical report generation by focusing preference optimization on clinical findings, adding visual-context preference inversion and counterfactual training for uncertain predictions.

  3. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  4. NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

    cs.NE 2026-08 conditional novelty 6.0 of 10

    NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...

  5. Scaling medical imaging report generation with multimodal reinforcement learning

    cs.CV 2026-01 conditional novelty 6.0 of 10

    UniRG-CXR, a Qwen3-VL-8B model trained with SFT plus GRPO reinforcement learning that directly optimizes the ReXrank metric components, reports state-of-the-art 1/RadCliQ-v1 results on all four ReXrank chest X-ray dat...

  6. Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Domain-adapted LLM encoders trained with masked token prediction and supervised contrastive learning improve chest X-ray image-text retrieval and external generalization, reaching GREEN scores of 0.308 on MIMIC-CXR an...

  7. Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis

    cs.CV 2025-07 reject novelty 6.0 of 10

    RadGazeIntent, a transformer model, predicts per-fixation diagnostic intention from radiologist gaze on chest X-rays, evaluated on three newly constructed intention-labeled datasets.

  8. Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation

    stat.ME 2025-07 conditional novelty 6.0 of 10

    REVTAF, a retrieval-augmented radiology report generator, reports average gains of 7.4 points on MIMIC-CXR and 2.9 points on IU X-Ray across nine metrics.

  9. Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 8-stage chest X-ray VQA benchmark and a context-aware model trained on it.

  10. Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Medical CLIP training with text, clinical, and graph soft labels plus negation hard negatives improves chest X-ray zero-shot and fine-tuned performance.

  11. CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction

    cs.AI 2025-05 conditional novelty 6.0 of 10

    CorBenchX provides a large-scale synthetic error dataset and benchmark for chest X-ray report error detection and correction, plus a multi-step RL method that improves model performance.

  12. Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A co-evolutionary framework that alternates between chest X-ray abnormality detection and radiology report generation, using each task to refine the other's pseudo-labels, achieves state-of-the-art results on MS-CXR a...

  13. Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Libra introduces a Temporal Alignment Connector for multimodal LLMs that fuses current and prior chest X-ray features and reports improved radiology report generation on MIMIC-CXR.

  14. ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ReXrank is a standardized public leaderboard for AI chest X-ray report generation, built on a 10,000-study private test set and 8 metrics, and it ranks 16 models with MedVersa as the current top performer.

  15. RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores

    cs.CL 2025-08 conditional novelty 5.0 of 10

    RadReason trains a 7B language model with GRPO to output six radiology error sub-scores plus textual reasons, reporting Kendall tau 0.730 on ReXVal, best among offline metrics.

  16. AMRG: Extend Vision Language Models for Automatic Mammography Report Generation

    eess.IV 2025-08 conditional novelty 5.0 of 10

    A LoRA-tuned MedGemma VLM generates mammography reports on the public DMID dataset, with ROUGE-L 0.5691, METEOR 0.6152, CIDEr 0.5818, and BI-RADS accuracy 0.5582.

  17. RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

    cs.CV 2025-07 reject novelty 5.0 of 10

    A video-based eye-gaze prompt improved report generation and diagnosis for one general-purpose vision-language model, LLaVA-OneVision, but hurt or barely helped two others, and the main comparison to medical models re...

  18. MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning

    cs.CV 2025-01 conditional novelty 5.0 of 10

    MoColl uses an LLM agent to query a trainable VQA model and synthesize captions, and the agent curates synthetic training pairs to improve the specialist; it reports top scores on radiology report generation.

  19. ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation

    cs.CL 2024-11 conditional novelty 5.0 of 10

    ReFINE is a fine-tuned Llama3 reward model that scores radiology reports on multiple criteria through a margin-based loss, showing higher correlation with human ratings than prior metrics.

  20. Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.

  21. Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration

    cs.CV 2025-05 reject novelty 4.0 of 10

    A fine-tuned 8B medical vision-language model that claims explainable grounding, uncertainty estimates and cancer prognosis, but its core uncertainty formula is mathematically inconsistent and key comparisons use the ...

  22. Clinically-Inspired Hierarchical Multi-Label Classification of Chest X-rays with a Penalty-Based Loss Function

    cs.CV 2025-02 reject novelty 4.0 of 10

    A hierarchical chest X-ray classifier reports AUROC 0.903 on CheXpert with a penalty loss, but the penalty uses a hard indicator with zero gradient, so the claimed dependency enforcement is not realized by the equations.

  23. GIT-CXR: End-to-End Transformer for Chest X-Ray Report Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Length-based curriculum learning improves an end-to-end GIT transformer for chest X-ray report generation, yielding high METEOR and clinical F1 scores; the state-of-the-art claim is however weakened by inconsistent ev...

  24. High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers

    cs.CL 2025-05 reject novelty 3.0 of 10

    A DeBERTa-Base model trained on GPT-4 pseudo-labels for 200k chest X-ray reports reports Macro F1 0.9120 on MIMIC-500, but the 'distillation' loss reduces to ordinary cross-entropy on hard labels.

  25. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

  26. Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation

    cs.CV 2024-12 conditional novelty 3.0 of 10

    A LLaVA-style radiology report generator using LoRA fine-tuning and stitched chest X-ray inputs placed fourth in the RRG24 shared task.

Pith tools