Pith. sign in

REVIEW 4 cited by

Contrastive Learning of Medical Visual Representations from Paired Images and Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.00747 v2 pith:4Z2HBQNA submitted 2020-10-02 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords medicalimageimagesdatapairedrepresentationscontrastivelearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning visual representations of medical images (e.g., X-rays) is core to medical image understanding but its progress has been held back by the scarcity of human annotations. Existing work commonly relies on fine-tuning weights transferred from ImageNet pretraining, which is suboptimal due to drastically different image characteristics, or rule-based label extraction from the textual report data paired with medical images, which is inaccurate and hard to generalize. Meanwhile, several recent studies show exciting results from unsupervised contrastive learning from natural images, but we find these methods help little on medical images because of their high inter-class similarity. We propose ConVIRT, an alternative unsupervised strategy to learn medical visual representations by exploiting naturally occurring paired descriptive text. Our new method of pretraining medical image encoders with the paired text data via a bidirectional contrastive objective between the two modalities is domain-agnostic, and requires no additional expert input. We test ConVIRT by transferring our pretrained weights to 4 medical image classification tasks and 2 zero-shot retrieval tasks, and show that it leads to image representations that considerably outperform strong baselines in most settings. Notably, in all 4 classification tasks, our method requires only 10\% as much labeled training data as an ImageNet initialized counterpart to achieve better or comparable performance, demonstrating superior data efficiency.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives

    cs.NE 2026-08 conditional novelty 6.0 of 10

    NeuroMosaic links MRI regions to diagnostic language via an anatomical graph router and concept memory, reporting external macro-F1 up to 0.784, IDH AUROC 0.918, and 0.703 pointing accuracy, with a 0.036 macro-F1 gain...

  2. Prospective clinical indication, post-hoc report leakage, and fusion design in multi-image chest radiograph classification: a patient-clustered evaluation

    cs.CV 2026-07 accept novelty 6.0 of 10

    Report-derived chest X-ray labels are almost perfectly predicted by the report text itself (AUROC 0.98), while images plus prospective indication reach only AUROC 0.78, quantifying report-label circularity.

  3. CARL-CXR: Continual Adapter-Based Routing for Task-Unknown Chest Radiograph Classification

    cs.CV 2026-02 reject novelty 5.0 of 10

    Across sequential MIMIC-CXR then CheXpert learning, CARL-CXR keeps MIMIC AUROC at 0.740 (0.012 forgetting) and routes 75% of test images correctly without task labels.

  4. Toward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data

    eess.IV 2026-07 conditional novelty 2.0 of 10

    On 668 public chest X-rays, SVM with HOG/GLCM features achieves 75.4% accuracy (AUC 0.755) for COVID-19 vs other pneumonia, modestly above the 71.6% majority baseline.

Pith tools