Pith. sign in

REVIEW 2 cited by

Understanding Transferable Representation Learning and Zero-shot Transfer in CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.00927 v2 pith:3JU5LVTW submitted 2023-10-02 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords cliplearningperformancezero-shotapproachdifferentimagerepresentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modal learning has become increasingly popular due to its ability to leverage information from different data sources (e.g., text and images) to improve the model performance. Recently, CLIP has emerged as an effective approach that employs vision-language contrastive pretraining to learn joint image and text representations and exhibits remarkable performance in zero-shot learning and text-guided natural image generation. Despite the huge practical success of CLIP, its theoretical understanding remains elusive. In this paper, we formally study transferrable representation learning underlying CLIP and demonstrate how features from different modalities get aligned. We also analyze its zero-shot transfer performance on the downstream tasks. Inspired by our analysis, we propose a new CLIP-type approach, which achieves better performance than CLIP and other state-of-the-art methods on benchmark datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal Preference Optimization for Long-Form Video Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    TPO trains video-LMMs to prefer answers generated from complete, relevant frames over answers from incomplete or irrelevant frames, improving temporal grounding on LongVideoBench, MLVU, and Video-MME.

  2. Expanding Event Modality Applications through a Robust CLIP-Based Encoder

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A CLIP-based encoder for event cameras, trained with contrastive, consistency, and KL losses, improves zero-shot and few-shot object recognition and extends to video anomaly detection and cross-modal retrieval.

Pith tools