Pith. sign in

REVIEW 20 cited by

RoentGen: Vision-Language Foundation Model for Chest X-ray Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.12737 v1 pith:45NHDQGG submitted 2022-11-23 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords modelmedicalimagessynthetictraineddataimagingmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Multimodal models trained on large natural image-text pair datasets have exhibited astounding abilities in generating high-quality images. Medical imaging data is fundamentally different to natural images, and the language used to succinctly capture relevant details in medical data uses a different, narrow but semantically rich, domain-specific vocabulary. Not surprisingly, multi-modal models trained on natural image-text pairs do not tend to generalize well to the medical domain. Developing generative imaging models faithfully representing medical concepts while providing compositional diversity could mitigate the existing paucity of high-quality, annotated medical imaging datasets. In this work, we develop a strategy to overcome the large natural-medical distributional shift by adapting a pre-trained latent diffusion model on a corpus of publicly available chest x-rays (CXR) and their corresponding radiology (text) reports. We investigate the model's ability to generate high-fidelity, diverse synthetic CXR conditioned on text prompts. We assess the model outputs quantitatively using image quality metrics, and evaluate image quality and text-image alignment by human domain experts. We present evidence that the resulting model (RoentGen) is able to create visually convincing, diverse synthetic CXR images, and that the output can be controlled to a new extent by using free-form text prompts including radiology-specific language. Fine-tuning this model on a fixed training set and using it as a data augmentation method, we measure a 5% improvement of a classifier trained jointly on synthetic and real images, and a 3% improvement when trained on a larger but purely synthetic training set. Finally, we observe that this fine-tuning distills in-domain knowledge in the text-encoder and can improve its representation capabilities of certain diseases like pneumothorax by 25%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniNDM: A Unified Noise-driven Detection and Mitigation Framework Against Sexual Content in Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    UniNDM detects sexual intent from early-stage diffusion noise and mitigates it via LLM-generated negative prompts and initial-noise optimization, across U-Net and DiT models.

  2. Steering Optimisation Trajectories in Diffusion Representation Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SteeringDRL identifies two optimization regimes in diffusion autoencoders and uses gated residual U-Nets with a log SNR curriculum to steer training toward disentangled representations, improving performance across mu...

  3. CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Hierarchical compositional conditioning lets a diffusion model generate higher-quality, fairer medical images and generalize to unseen demographic intersections without extra training data.

  4. Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new Vietnamese PET/CT-report dataset improves medical VLM report generation and VQA, but clinical F1 scores remain modest.

  5. Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A 3D encoder pretrained with GPT-4V slice captions and partial optimal transport alignment beats vision-only SSL baselines on several medical tasks, but a key evaluation dataset may overlap with pretraining.

  6. RoentMod: A Synthetic Chest X-Ray Modification Model to Identify and Correct Image Interpretation Model Shortcuts

    eess.IV 2025-09 conditional novelty 6.0 of 10

    RoentMod edits chest X-rays to add specific diseases, exposing and partly correcting shortcut learning in AI interpretation models.

  7. MedDiff-FT: Data-Efficient Diffusion Model Fine-tuning with Structural Guidance for Controllable Medical Image Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Fine-tuning Stable Diffusion with mask guidance yields synthetic medical image-mask pairs that raise nnU-Net Dice scores by about 1% on average across five datasets.

  8. ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer

    eess.IV 2025-05 conditional novelty 6.0 of 10

    An anatomy- and clinical-conditioned diffusion model that synthesizes post-treatment CT and uses its features to improve immunotherapy response prediction in NSCLC.

  9. Can Modern NLP Systems Reliably Annotate Chest Radiography Exams? A Pre-Purchase Evaluation and Comparative Study of Solutions from AWS, Google, Azure, John Snow Labs, and Open-Source Models on an Independent Pediatric Dataset

    cs.CL 2025-05 conditional novelty 6.0 of 10

    On 95,008 pediatric chest X-ray reports, six NLP labeling systems varied widely in entity extraction and assertion classification, with consensus-based accuracy between 50% and 76%.

  10. Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images

    cs.CV 2025-04 reject novelty 6.0 of 10

    Radiologists' eye movements differ slightly between real and AI-generated chest X-rays, especially for longest and shortest fixations, but none of the differences are backed by inferential statistics.

  11. The Devil is in the Prompts: De-Identification Traces Enhance Memorization Risks in Synthetic Chest X-Ray Generation

    eess.IV 2025-02 reject novelty 6.0 of 10

    In MIMIC-CXR, prompts containing the de-identification marker '___' show the highest text-conditional memorization scores, and the marker itself is the top contributing token.

  12. MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A new 1.47M-instance biomedical instruction-tuning dataset improves a mixed-modal 7B model's medical VQA accuracy by 18 to 26 percentage points over GPT-4o and Chameleon.

  13. Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

    cs.CV 2026-07 conditional novelty 5.0 of 10

    LLIFT generates semi-synthetic brain MRIs with user-placed lesion-like patches using weak labels, though its image-level FID results do not show the patch itself is pathology-realistic.

  14. Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers

    cs.AI 2026-07 reject novelty 5.0 of 10

    A synthetic-pretrained COVID-19 CT classifier achieves higher worst-cell fairness than full-real training with 1% of the real data, and synthetic test cohorts reproduce the oracle's subgroup ranking (Spearman ρ=1.00).

  15. CARL-CXR: Continual Adapter-Based Routing for Task-Unknown Chest Radiograph Classification

    cs.CV 2026-02 reject novelty 5.0 of 10

    Across sequential MIMIC-CXR then CheXpert learning, CARL-CXR keeps MIMIC AUROC at 0.740 (0.012 forgetting) and routes 75% of test images correctly without task labels.

  16. Perceptual Evaluation of GANs and Diffusion Models for Generating X-rays

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    In a three-radiologist reader study on MIMIC-CXR images, diffusion-generated chest X-rays looked more realistic overall, while GANs were better for the absence of an enlarged cardiac silhouette.

  17. Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.

  18. MRI Image Generation Based on Text Prompts

    eess.IV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning Stable Diffusion with MRI-text pairs yields plausible brain MRI images by field strength and modality, and synthetic images appear to improve a small MRI contrast classification task.

  19. Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning large diffusion models with LoRA yielded the best FID scores for text-to-image generation on colonoscopy and radiology data, while a compact Stable-Diffusion-derived model (MSDM) remained competitive at lo...

  20. Diffusion-Based Data Augmentation for Medical Image Segmentation

    cs.CV 2025-08 conditional novelty 3.0 of 10

    DiffAug augments medical training data with text-and-mask guided diffusion inpainting, filtered by a latent-space segmentation network, improving polyp and optic-disc segmentation.

Pith tools