Pith. sign in

REVIEW 10 cited by

Does CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.13906 v1 pith:LB3S2DVN submitted 2021-12-27 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords pubmedclipvisualclipmedvqamedicalansweringdatasetsdomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language--Image Pre-training (CLIP) has shown remarkable success in learning with cross-modal supervision from extensive amounts of image--text pairs collected online. Thus far, the effectiveness of CLIP has been investigated primarily in general-domain multimodal problems. This work evaluates the effectiveness of CLIP for the task of Medical Visual Question Answering (MedVQA). To this end, we present PubMedCLIP, a fine-tuned version of CLIP for the medical domain based on PubMed articles. Our experiments are conducted on two MedVQA benchmark datasets and investigate two MedVQA methods, MEVF (Mixture of Enhanced Visual Features) and QCR (Question answering via Conditional Reasoning). For each of these, we assess the merits of visual representation learning using PubMedCLIP, the original CLIP, and state-of-the-art MAML (Model-Agnostic Meta-Learning) networks pre-trained only on visual data. We open source the code for our MedVQA pipeline and pre-training PubMedCLIP. CLIP and PubMedCLIP achieve improvements in comparison to MAML's visual encoder. PubMedCLIP achieves the best results with gains in the overall accuracy of up to 3%. Individual examples illustrate the strengths of PubMedCLIP in comparison to the previously widely used MAML networks. Visual representation learning with language supervision in PubMedCLIP leads to noticeable improvements for MedVQA. Our experiments reveal distributional differences in the two MedVQA benchmark datasets that have not been imparted in previous work and cause different back-end visual encoders in PubMedCLIP to exhibit different behavior on these datasets. Moreover, we witness fundamental performance differences of VQA in general versus medical domains.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Symmetric Alignment: Spectral Diagnostics of Modality Imbalance in Vision-Language Models in the Medical Domain

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    SAS reveals that medical images retain richer structural information than paired clinical reports in VLMs, an asymmetry hidden from symmetric metrics, with strongest correlation to retrieval performance.

  2. BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs

    cs.CV 2023-03 conditional novelty 7.0 of 10

    BiomedCLIP, pretrained on the new 15-million-pair PMC-15M dataset, achieves state-of-the-art performance on diverse biomedical vision-language tasks and even outperforms radiology-specific models on chest X-ray pneumo...

  3. Geometry-Aware Distillation for Prompt Tuning Biomedical Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    OGKD injects inter-class geometry into teacher targets for two distillation losses (GAD on global tokens, LGD on patches) and reports 1.7-2.8% average accuracy gains over prior VLM adaptation methods on 11 medical datasets.

  4. LangMamba: A Language-driven Mamba Framework for Low-dose CT Denoising with Vision-language Models

    eess.IV 2025-07 conditional novelty 6.0 of 10

    A language-driven Mamba framework for low-dose CT denoising uses a frozen vision-language model to provide semantic supervision, achieving marginal but consistent quantitative gains over previous methods.

  5. Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Brain-Adapter combines a text-conditioned attention stream and a visual MIL stream with consistency constraints and uncertainty-aware refinement to classify acute intracranial pathologies in 3D CT scans from 2D VLMs a...

  6. BiomedAP: A Vision-Informed Dual-Anchor Framework with Gated Cross-Modal Fusion for Robust Medical Vision-Language Adaptation

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    BiomedAP improves robustness of biomedical VLMs to prompt variations using gated cross-modal fusion and dual-anchor constraints, outperforming baselines on 11 benchmarks.

  7. VSF-Med:A Vulnerability Scoring Framework for Medical Vision-Language Models

    cs.CV 2025-06 reject novelty 5.0 of 10

    VSF-Med introduces an eight-dimension, judge-scored vulnerability score for medical VLMs and reports that all five tested models are most vulnerable to persistent attack effects, with Llama-3.2 showing the largest drop.

  8. Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

    cs.CV 2025-06 unverdicted novelty 5.0 of 10

    Introduces Hybrid Tuning adapter with frequency filtering and noise estimation to adapt CLIP for ultrasound segmentation and classification, claiming outperformance on six multi-center datasets.

  9. Prompt Mechanisms in Medical Imaging: A Comprehensive Survey

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A broad survey that organizes prompt mechanisms for medical image generation, segmentation, and classification into a two-dimensional taxonomy of core technologies and clinical applications.

  10. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0 of 10

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.

Pith tools