Pith. sign in

REVIEW 6 cited by

CLIP2Video: Mastering Video-Text Retrieval via Image CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11097 v1 pith:2HNO52MG submitted 2021-06-21 cs.CV

classification cs.CV
keywords modelretrievaltemporalvideovideo-textblockclipclip2video
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present CLIP2Video network to transfer the image-language pre-training model to video-text retrieval in an end-to-end manner. Leading approaches in the domain of video-and-language learning try to distill the spatio-temporal video features and multi-modal interaction between videos and languages from a large-scale video-text dataset. Different from them, we leverage pretrained image-language model, simplify it as a two-stage framework with co-learning of image-text and enhancing temporal relations between video frames and video-text respectively, make it able to train on comparatively small datasets. Specifically, based on the spatial semantics captured by Contrastive Language-Image Pretraining (CLIP) model, our model involves a Temporal Difference Block to capture motions at fine temporal video frames, and a Temporal Alignment Block to re-align the tokens of video clips and phrases and enhance the multi-modal correlation. We conduct thorough ablation studies, and achieve state-of-the-art performance on major text-to-video and video-to-text retrieval benchmarks, including new records of retrieval accuracy on MSR-VTT, MSVD and VATEX.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Hard negatives selected by visual confusability in sign embeddings, not linguistic similarity, substantially raise fine-grained sign-language retrieval accuracy without collapsing coarse performance.

  2. Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval

    cs.MM 2025-07 conditional novelty 6.0 of 10

    ProCLIP selects query-relevant frames via prompt-aware cross-attention and prunes candidates with a CLIP-distilled lightweight model, matching top accuracy at a fraction of the latency.

  3. Learning Speaker-Invariant Visual Features for Lipreading

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SIFLip improves lipreading on unseen speakers by combining text-aligned contrastive learning with gradient-reversal-based speaker disentanglement.

  4. A Mathematical Perspective On Contrastive Learning

    stat.ML 2025-05 conditional novelty 6.0 of 10

    A probabilistic tilting framework for contrastive learning yields closed-form Gaussian results showing which conditional statistics each loss can recover.

  5. Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.

  6. ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 7B multimodal model that fuses audio and visual signals with explicit timestamps achieves strong measured comprehension of real-world short videos on the authors' new ShortVid-Bench benchmark.

Pith tools