Pith. sign in

REVIEW 4 cited by

Adapting Pretrained Vision-Language Foundational Models to Medical Imaging Domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04133 v1 pith:GYFFK6S7 submitted 2022-10-09 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords imagesmedicalmodelsdiffusionmodelstableabnormalitycapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal foundation models are typically trained on millions of pairs of natural images and text captions, frequently obtained through web-crawling approaches. Although such models depict excellent generative capabilities, they do not typically generalize well to specific domains such as medical images that have fundamentally shifted distributions compared to natural images. Building generative models for medical images that faithfully depict clinical context may help alleviate the paucity of healthcare datasets. Thus, in this study, we seek to research and expand the representational capabilities of large pretrained foundation models to medical concepts, specifically for leveraging the Stable Diffusion model to generate domain specific images found in medical imaging. We explore the sub-components of the Stable Diffusion pipeline (the variational autoencoder, the U-Net and the text-encoder) to fine-tune the model to generate medical images. We benchmark the efficacy of these efforts using quantitative image quality metrics and qualitative radiologist-driven evaluations that accurately represent the clinical content of conditional text prompts. Our best-performing model improves upon the stable diffusion baseline and can be conditioned to insert a realistic-looking abnormality on a synthetic radiology image, while maintaining a 95% accuracy on a classifier trained to detect the abnormality.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  2. Data Can Speak for Itself: Quality-guided Utilization of Wireless Synthetic Data

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SynCheck uses margin-based quality metrics and semi-supervised pseudo-labeling to filter and relabel wireless synthetic data, improving task accuracy over naive mixture training.

  3. Anatomy-Grounded Weakly Supervised Prompt Tuning for Chest X-ray Latent Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Anatomy-token prompt tuning with Gaussian location priors improves phrase grounding in a chest X-ray latent diffusion model, beating prior models on MS-CXR-loc and VinDr-CXR.

  4. Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.

Pith tools