REVIEW 20 cited by
RoentGen: Vision-Language Foundation Model for Chest X-ray Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Multimodal models trained on large natural image-text pair datasets have exhibited astounding abilities in generating high-quality images. Medical imaging data is fundamentally different to natural images, and the language used to succinctly capture relevant details in medical data uses a different, narrow but semantically rich, domain-specific vocabulary. Not surprisingly, multi-modal models trained on natural image-text pairs do not tend to generalize well to the medical domain. Developing generative imaging models faithfully representing medical concepts while providing compositional diversity could mitigate the existing paucity of high-quality, annotated medical imaging datasets. In this work, we develop a strategy to overcome the large natural-medical distributional shift by adapting a pre-trained latent diffusion model on a corpus of publicly available chest x-rays (CXR) and their corresponding radiology (text) reports. We investigate the model's ability to generate high-fidelity, diverse synthetic CXR conditioned on text prompts. We assess the model outputs quantitatively using image quality metrics, and evaluate image quality and text-image alignment by human domain experts. We present evidence that the resulting model (RoentGen) is able to create visually convincing, diverse synthetic CXR images, and that the output can be controlled to a new extent by using free-form text prompts including radiology-specific language. Fine-tuning this model on a fixed training set and using it as a data augmentation method, we measure a 5% improvement of a classifier trained jointly on synthetic and real images, and a 3% improvement when trained on a larger but purely synthetic training set. Finally, we observe that this fine-tuning distills in-domain knowledge in the text-encoder and can improve its representation capabilities of certain diseases like pneumothorax by 25%.
Forward citations
Cited by 20 Pith papers
-
UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation
UniNDM detects sexual intent from early-stage diffusion noise and mitigates it via LLM-generated negative prompts and initial-noise optimization, across U-Net and DiT models.
-
Steering Optimisation Trajectories in Diffusion Representation Learning
SteeringDRL identifies two optimization regimes in diffusion autoencoders and uses gated residual U-Nets with a log SNR curriculum to steer training toward disentangled representations, improving performance across mu...
-
CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation
Hierarchical compositional conditioning lets a diffusion model generate higher-quality, fairer medical images and generalize to unseen demographic intersections without extra training data.
-
Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation
A new Vietnamese PET/CT-report dataset improves medical VLM report generation and VQA, but clinical F1 scores remain modest.
-
Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.
-
RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts
RoentMod edits chest X-rays to add specific diseases, exposing and partly correcting shortcut learning in AI interpretation models.
-
MedDiff-FT: Data-Efficient Diffusion Model Fine-tuning with Structural Guidance for Controllable Medical Image Synthesis
Fine-tuning Stable Diffusion with mask guidance yields synthetic medical image-mask pairs that raise nnU-Net Dice scores by about 1% on average across five datasets.
-
ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer
An anatomy- and clinical-conditioned diffusion model that synthesizes post-treatment CT and uses its features to improve immunotherapy response prediction in NSCLC.
-
Can Modern NLP Systems Reliably Annotate Chest Radiography Exams? A Pre-Purchase Evaluation and Comparative Study of Solutions from AWS, Google, Azure, John Snow Labs, and Open-Source Models on an Independent Pediatric Dataset
On 95,008 pediatric chest X-ray reports, six NLP labeling systems varied widely in entity extraction and assertion classification, with consensus-based accuracy between 50% and 76%.
-
Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images
Radiologists' eye movements differ slightly between real and AI-generated chest X-rays, especially for longest and shortest fixations, but none of the differences are backed by inferential statistics.
-
The Devil is in the Prompts: De-Identification Traces Enhance Memorization Risks in Synthetic Chest X-Ray Generation
In MIMIC-CXR, prompts containing the de-identification marker '___' show the highest text-conditional memorization scores, and the marker itself is the top contributing token.
-
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants
A new 1.47M-instance biomedical instruction-tuning dataset improves a mixed-modal 7B model's medical VQA accuracy by 18 to 26 percentage points over GPT-4o and Chameleon.
-
Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches
LLIFT generates semi-synthetic brain MRIs with user-placed lesion-like patches using weak labels, though its image-level FID results do not show the patch itself is pathology-realistic.
-
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
A synthetic-pretrained COVID-19 CT classifier achieves higher worst-cell fairness than full-real training with 1% of the real data, and synthetic test cohorts reproduce the oracle's subgroup ranking (Spearman ρ=1.00).
-
CARL-CXR: Continual Adapter-Based Routing for Task-Unknown Chest Radiograph Classification
Across sequential MIMIC-CXR then CheXpert learning, CARL-CXR keeps MIMIC AUROC at 0.740 (0.012 forgetting) and routes 75% of test images correctly without task labels.
-
Perceptual Evaluation of GANs and Diffusion Models for Generating X-rays
In a three-radiologist reader study on MIMIC-CXR images, diffusion-generated chest X-rays looked more realistic overall, while GANs were better for the absence of an enlarged cardiac silhouette.
-
Prompt Mechanisms in Medical Imaging: A Comprehensive Survey
A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.
-
MRI Image Generation Based on Text Prompts
Fine-tuning Stable Diffusion with MRI-text pairs yields plausible brain MRI images by field strength and modality, and synthetic images appear to improve a small MRI contrast classification task.
-
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
Fine-tuning large diffusion models with LoRA yielded the best FID scores for text-to-image generation on colonoscopy and radiology data, while a compact Stable-Diffusion-derived model (MSDM) remained competitive at lo...
-
Diffusion-Based Data Augmentation for Medical Image Segmentation
DiffAug augments medical training data with text-and-mask guided diffusion inpainting, filtered by a latent-space segmentation network, improving polyp and optic-disc segmentation.
Discussion (0). Continue with ORCID to comment.