REVIEW 27 cited by
CheXpert Plus: Augmenting a Large Chest X-ray Dataset with Text Radiology Reports, Patient Demographics and Additional Image Formats
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Since the release of the original CheXpert paper five years ago, CheXpert has become one of the most widely used and cited clinical AI datasets. The emergence of vision language models has sparked an increase in demands for sharing reports linked to CheXpert images, along with a growing interest among AI fairness researchers in obtaining demographic data. To address this, CheXpert Plus serves as a new collection of radiology data sources, made publicly available to enhance the scaling, performance, robustness, and fairness of models for all subsequent machine learning tasks in the field of radiology. CheXpert Plus is the largest text dataset publicly released in radiology, with a total of 36 million text tokens, including 13 million impression tokens. To the best of our knowledge, it represents the largest text de-identification effort in radiology, with almost 1 million PHI spans anonymized. It is only the second time that a large-scale English paired dataset has been released in radiology, thereby enabling, for the first time, cross-institution training at scale. All reports are paired with high-quality images in DICOM format, along with numerous image and patient metadata covering various clinical and socio-economic groups, as well as many pathology labels and RadGraph annotations. We hope this dataset will boost research for AI models that can further assist radiologists and help improve medical care. Data is available at the following URL: https://stanfordaimi.azurewebsites.net/datasets/5158c524-d3ab-4e02-96e9-6ee9efc110a1 Models are available at the following URL: https://github.com/Stanford-AIMI/chexpert-plus
Forward citations
Cited by 27 Pith papers
-
Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation
LePaX enables high-resolution chest X-ray report generation by learning to allocate resolution to diagnostically relevant regions and fusing high-res patches back into global features without increasing token count.
-
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.
-
Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos
An agentic multi-stage pipeline produces Colon-Bench—528 densely annotated colonoscopy windows spanning 14 lesions, 300k boxes, 213k masks and clinical text—then benchmarks MLLMs and a colon-skill prompt that gains up...
-
UniMod: Enhancing Multi-Modal Medical Diagnosis through Cross-Modality and Within-Modality Alignment
UniMod improves multi-modal medical diagnosis by adding independent image-only and text-only classification losses, plus cross- and within-modality alignment, reducing over-reliance on text.
-
Positive-Unlabeled Preference Optimization For Chest X-ray Report Generation
PU-DPO applies positive-unlabeled learning to preference optimization so that report generators learn to mention findings that are present but missing from noisy training reports.
-
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...
-
RadHarmony: Radiological Data Handling in the Era of Agentic AI
RadHarmony is a unified Python library and AI-agent workflow for harmonizing 24 public radiology datasets, demonstrated by a multi-dataset self-supervised chest X-ray model with no dataset-specific code.
-
Reconfigurable Radiology Labels Without Relabeling
Cached structured report annotations let radiology label schemas be changed with dictionary edits instead of relabeling the corpus, recovering long-tail findings at near-zero marginal cost.
-
Scaling medical imaging report generation with multimodal reinforcement learning
UniRG-CXR, a Qwen3-VL-8B model trained with SFT plus GRPO reinforcement learning that directly optimizes the ReXrank metric components, reports state-of-the-art 1/RadCliQ-v1 results on all four ReXrank chest X-ray dat...
-
MedVision: Benchmarking Quantitative Medical Image Analysis
A 30.8M-pair medical imaging benchmark shows pretrained vision-language models are poor at detection, tumor-size, and angle/distance measurement, and that fine-tuning on the benchmark substantially improves them.
-
Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
Domain-adapted LLM encoders trained with masked token prediction and supervised contrastive learning improve chest X-ray image-text retrieval and external generalization, reaching GREEN scores of 0.308 on MIMIC-CXR an...
-
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
Decomposing clinical terms into visual attributes lets 0.23B-2B vision-language models match or beat much larger medical VLMs for abnormality grounding with only 16k training pairs.
-
BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP
A continued-pretrained ModernBERT encoder for biomedical and clinical text claims SOTA on several clinical NLP tasks, with caveats about data overlap between pretraining and evaluation.
-
Bringing CLIP to the Clinic: Dynamic Soft Labels and Negation-Aware Learning for Medical Analysis
Medical CLIP training with text, clinical, and graph soft labels plus negation hard negatives improves chest X-ray zero-shot and fine-tuned performance.
-
ReXGradient-160K: A Large-Scale Publicly Available Dataset of Chest Radiographs with Free-text Reports
ReXGradient-160K provides 160,000 chest X-ray studies with free-text reports from 109,487 patients across 79 medical sites, with public and private test splits.
-
Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-ray Report Generation
AM-MRG combines disease-region extraction with two Hopfield memory retrievers to improve LLM-generated chest X-ray reports on IU X-ray, MIMIC-CXR, and Chexpert Plus.
-
FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models
A modular query-code-update pipeline reduces measurement hallucinations in chest X-ray reports by replacing model-generated numbers with measurements from specialized vision tools.
-
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
ReXrank is a standardized public leaderboard for AI chest X-ray report generation, built on a 10,000-study private test set and 8 metrics, and it ranks 16 models with MedVersa as the current top performer.
-
R2GenKG: Hierarchical Multi-modal Knowledge Graph for LLM-based Radiology Report Generation
R2GenKG generates X-ray reports with an LLM conditioned on a GPT-4o-built multi-modal knowledge graph, reporting small metric gains on IU-Xray and CheXpert Plus.
-
Label-free estimation of clinically relevant performance metrics under distribution shifts
New methods CM-ATC and CM-DoC estimate a classifier's confusion matrix on unlabeled medical images by applying class-specific confidence thresholds and offsets, outperforming the prior baseline on real-world shifts bu...
-
MCA-RG: Enhancing LLMs with Medical Concept Alignment for Radiology Report Generation
MCA-RG uses concept alignment, contrastive learning, matching loss, and feature gating to generate radiology reports, reporting SOTA on MIMIC-CXR and CheXpert Plus.
-
RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
RADAR filters an LLM's radiology findings by agreement with an expert classifier and retrieves only the missing observations, reporting improved clinical accuracy on three datasets.
-
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI
GMAI-VL-5.5M is a new 5.5M-sample medical image-text dataset built from 219 datasets via GPT-4o annotation-guided generation, and GMAI-VL is a three-stage LLaVA-style model reporting SOTA numbers, though the evaluatio...
-
Medical Video Generation for Disease Progression Simulation
MVG generates synthetic disease-progression videos from one medical image and a text prompt, using repeated diffusion editing plus video interpolation, evaluated on chest X-ray, retina, and skin images.
-
RadPhi-3: Small Language Models for Radiology
RadPhi-3, a 3.8B parameter instruction-tuned model, handles radiology QA and chest X-ray report utilities and posts a marginal SOTA score on the RaLEs benchmark.
-
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
A LLaVA-style radiology report generator using LoRA fine-tuning and stitched chest X-ray inputs placed fourth in the RRG24 shared task.
-
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2
On the IU-Xray dataset, SWIN-BART outperforms ViT-B16-BART, SWIN-GPT-2, and ViT-B16-GPT-2 on n-gram and embedding-based report metrics.
Discussion (0). Continue with ORCID to comment.