Pith. sign in

REVIEW 2 cited by

DeViDe: Faceted medical knowledge for improved medical vision-language pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03618 v1 pith:QVIPY6RU submitted 2024-04-04 cs.CV

classification cs.CV
keywords knowledgedevidemedicaldescriptionsradiologyreportsabstractdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language pre-training for chest X-rays has made significant strides, primarily by utilizing paired radiographs and radiology reports. However, existing approaches often face challenges in encoding medical knowledge effectively. While radiology reports provide insights into the current disease manifestation, medical definitions (as used by contemporary methods) tend to be overly abstract, creating a gap in knowledge. To address this, we propose DeViDe, a novel transformer-based method that leverages radiographic descriptions from the open web. These descriptions outline general visual characteristics of diseases in radiographs, and when combined with abstract definitions and radiology reports, provide a holistic snapshot of knowledge. DeViDe incorporates three key features for knowledge-augmented vision language alignment: First, a large-language model-based augmentation is employed to homogenise medical knowledge from diverse sources. Second, this knowledge is aligned with image information at various levels of granularity. Third, a novel projection layer is proposed to handle the complexity of aligning each image with multiple descriptions arising in a multi-label setting. In zero-shot settings, DeViDe performs comparably to fully supervised models on external datasets and achieves state-of-the-art results on three large-scale datasets. Additionally, fine-tuning DeViDe on four downstream tasks and six segmentation tasks showcases its superior performance across data from diverse distributions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Decomposing clinical terms into visual attributes lets 0.23B-2B vision-language models match or beat much larger medical VLMs for abnormality grounding with only 16k training pairs.

  2. TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A closed-form, multi-layer prototype residual corrects medical VLM logits from few support samples and improves OOD balanced accuracy over zero-shot and most trained adapters.

Pith tools