REVIEW 26 cited by
CheXbert: Combining Automatic Labelers and Expert Annotations for Accurate Radiology Report Labeling Using BERT
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The extraction of labels from radiology text reports enables large-scale training of medical imaging models. Existing approaches to report labeling typically rely either on sophisticated feature engineering based on medical domain knowledge or manual annotations by experts. In this work, we introduce a BERT-based approach to medical image report labeling that exploits both the scale of available rule-based systems and the quality of expert annotations. We demonstrate superior performance of a biomedically pretrained BERT model first trained on annotations of a rule-based labeler and then finetuned on a small set of expert annotations augmented with automated backtranslation. We find that our final model, CheXbert, is able to outperform the previous best rules-based labeler with statistical significance, setting a new SOTA for report labeling on one of the largest datasets of chest x-rays.
Forward citations
Cited by 26 Pith papers
-
AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretation
AnatomiX, a two-stage anatomy-first multimodal LLM for chest X-ray interpretation, reports >25% relative gains on anatomy grounding and grounded captioning, but some aggregate benchmark numbers are internally inconsis...
-
Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation
DPO-Clin improves medical report generation by focusing preference optimization on clinical findings, adding visual-context preference inversion and counterfactual training for uncertain predictions.
-
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...
-
NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives
NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...
-
Scaling medical imaging report generation with multimodal reinforcement learning
UniRG-CXR, a Qwen3-VL-8B model trained with SFT plus GRPO reinforcement learning that directly optimizes the ReXrank metric components, reports state-of-the-art 1/RadCliQ-v1 results on all four ReXrank chest X-ray dat...
-
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
Domain-adapted LLM encoders trained with masked token prediction and supervised contrastive learning improve chest X-ray image-text retrieval and external generalization, reaching GREEN scores of 0.308 on MIMIC-CXR an...
-
Interpreting Radiologist's Intention from Eye Movements in Chest X-ray Diagnosis
RadGazeIntent, a transformer model, predicts per-fixation diagnostic intention from radiologist gaze on chest X-rays, evaluated on three newly constructed intention-labeled datasets.
-
Learnable Retrieval Enhanced Visual-Text Alignment and Fusion for Radiology Report Generation
REVTAF, a retrieval-augmented radiology report generator, reports average gains of 7.4 points on MIMIC-CXR and 2.9 points on IU X-Ray across nine metrics.
-
Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning
A new 8-stage chest X-ray VQA benchmark and a context-aware model trained on it.
-
Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis
Medical CLIP training with text, clinical, and graph soft labels plus negation hard negatives improves chest X-ray zero-shot and fine-tuned performance.
-
CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction
CorBenchX provides a large-scale synthetic error dataset and benchmark for chest X-ray report error detection and correction, plus a multi-step RL method that improves model performance.
-
Unlocking the Potential of Weakly Labeled Data: A Co-Evolutionary Learning Framework for Abnormality Detection and Report Generation
A co-evolutionary framework that alternates between chest X-ray abnormality detection and radiology report generation, using each task to refine the other's pseudo-labels, achieves state-of-the-art results on MS-CXR a...
-
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis
Libra introduces a Temporal Alignment Connector for multimodal LLMs that fuses current and prior chest X-ray features and reports improved radiology report generation on MIMIC-CXR.
-
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
ReXrank is a standardized public leaderboard for AI chest X-ray report generation, built on a 10,000-study private test set and 8 metrics, and it ranks 16 models with MedVersa as the current top performer.
-
RadReason: Radiology Report Evaluation Metric with Reasons and Sub-Scores
RadReason trains a 7B language model with GRPO to output six radiology error sub-scores plus textual reasons, reporting Kendall tau 0.730 on ReXVal, best among offline metrics.
-
AMRG: Extend Vision Language Models for Automatic Mammography Report Generation
A LoRA-tuned MedGemma VLM generates mammography reports on the public DMID dataset, with ROUGE-L 0.5691, METEOR 0.6152, CIDEr 0.5818, and BI-RADS accuracy 0.5582.
-
RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze
A video-based eye-gaze prompt improved report generation and diagnosis for one general-purpose vision-language model, LLaVA-OneVision, but hurt or barely helped two others, and the main comparison to medical models re...
-
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
MoColl uses an LLM agent to query a trainable VQA model and synthesize captions, and the agent curates synthetic training pairs to improve the specialist; it reports top scores on radiology report generation.
-
ReFINE: A Reward-Based Framework for Interpretable and Nuanced Evaluation of Radiology Report Generation
ReFINE is a fine-tuned Llama3 reward model that scores radiology reports on multiple criteria through a margin-based loss, showing higher correlation with human ratings than prior metrics.
-
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
A fine-tuned 8B medical vision-language model that claims explainable grounding, uncertainty estimates and cancer prognosis, but its core uncertainty formula is mathematically inconsistent and key comparisons use the ...
-
Clinically-Inspired Hierarchical Multi-Label Classification of Chest X-rays with a Penalty-Based Loss Function
A hierarchical chest X-ray classifier reports AUROC 0.903 on CheXpert with a penalty loss, but the penalty uses a hard indicator with zero gradient, so the claimed dependency enforcement is not realized by the equations.
-
GIT-CXR: End-to-End Transformer for Chest X-Ray Report Generation
Length-based curriculum learning improves an end-to-end GIT transformer for chest X-ray report generation, yielding high METEOR and clinical F1 scores; the state-of-the-art claim is however weakened by inconsistent ev...
-
High-Fidelity Pseudo-label Generation by Large Language Models for Training Robust Radiology Report Classifiers
A DeBERTa-Base model trained on GPT-4 pseudo-labels for 200k chest X-ray reports reports Macro F1 0.9120 on MIMIC-500, but the 'distillation' loss reduces to ordinary cross-entropy on hard labels.
-
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.
-
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
A LLaVA-style radiology report generator using LoRA fine-tuning and stitched chest X-ray inputs placed fourth in the RRG24 shared task.
Discussion (0). Continue with ORCID to comment.