Pith. sign in

REVIEW 2 cited by

EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.06644 v2 pith:3OIM6XRP submitted 2024-09-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords diseaseseyeclipfoundationmulti-modallearningmodalitiesophthalmiccontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Early detection of eye diseases like glaucoma, macular degeneration, and diabetic retinopathy is crucial for preventing vision loss. While artificial intelligence (AI) foundation models hold significant promise for addressing these challenges, existing ophthalmic foundation models primarily focus on a single modality, whereas diagnosing eye diseases requires multiple modalities. A critical yet often overlooked aspect is harnessing the multi-view information across various modalities for the same patient. Additionally, due to the long-tail nature of ophthalmic diseases, standard fully supervised or unsupervised learning approaches often struggle. Therefore, it is essential to integrate clinical text to capture a broader spectrum of diseases. We propose EyeCLIP, a visual-language foundation model developed using over 2.77 million multi-modal ophthalmology images with partial text data. To fully leverage the large multi-modal unlabeled and labeled data, we introduced a pretraining strategy that combines self-supervised reconstructions, multi-modal image contrastive learning, and image-text contrastive learning to learn a shared representation of multiple modalities. Through evaluation using 14 benchmark datasets, EyeCLIP can be transferred to a wide range of downstream tasks involving ocular and systemic diseases, achieving state-of-the-art performance in disease classification, visual question answering, and cross-modal retrieval. EyeCLIP represents a significant advancement over previous methods, especially showcasing few-shot, even zero-shot capabilities in real-world long-tail scenarios.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning

    cs.AI 2025-07 conditional novelty 6.0 of 10

    FundusExpert, an 8B ophthalmic MLLM trained on region-grounded cognitive-chain instructions, reports state-of-the-art QA and report-generation results, with a fitted data-scaling exponent of 0.068.

  2. Color Fundus Photography Analysis: Co-evolution of Data, Preprocessing, and Modeling toward Multimodal AI

    cs.CV 2026-07 conditional novelty 3.0 of 10

    A review that organizes color fundus photography AI as the co-evolution of datasets, preprocessing, and models, concluding that performance ceilings are set by joint optimization of data, hygiene, and multimodal context.

Pith tools