Pith. sign in

REVIEW 21 cited by

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08860 v2 pith:NLXDZ2HY submitted 2021-04-18 cs.CV

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

classification cs.CV
keywords clipmodelretrievalvideo-textclip4clipdatasetsempiricalpre-training
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training), an image-language pre-training model, has demonstrated the power of visual concepts learning from web collected image-text datasets. In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner. Several questions are investigated via empirical studies: 1) Whether image feature is enough for video-text retrieval? 2) How a post-pretraining on a large-scale video-text dataset based on the CLIP affect the performance? 3) What is the practical mechanism to model temporal dependency between video frames? And 4) The Hyper-parameters sensitivity of the model on video-text retrieval task. Extensive experimental results present that the CLIP4Clip model transferred from the CLIP can achieve SOTA results on various video-text retrieval datasets, including MSR-VTT, MSVC, LSMDC, ActivityNet, and DiDeMo. We release our code at https://github.com/ArrowLuo/CLIP4Clip.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Trajectory-aware Cross-view Geo-localization with Sequential Observations

    cs.CV 2026-07 conditional novelty 7.0

    TrajLoc combines video and text route descriptions with trajectory geometry to retrieve satellite images, introducing the SeqGeo-VL benchmark and improving state-of-the-art on both video and text geo-localization.

  2. Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

    cs.CV 2026-07 accept novelty 7.0

    A camera-trap TVR benchmark of 135 ethology queries plus an interpretable SALMA-to-JSON plus constrained-LLM-parser pipeline yields 34% set F1, beating zero-shot VLMs at 18%.

  3. Reasoning Text-to-Video Retrieval for Operating Room Clips via Action-Driven Digital Twins

    cs.CV 2026-06 conditional novelty 7.0

    OR3 converts OR clips to action-driven digital twins, uses LLM imagination for hypothetical ActDTs, and achieves 57.6 R@1 and 77.3 R@5 on 276 implicit queries from 386 robotic knee procedure clips, outperforming baselines.

  4. OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

    cs.CV 2026-05 unverdicted novelty 7.0

    OmniRetriever-7B uses fusion-as-teacher distillation plus Tuple-InfoNCE to improve any-to-any audio-video-text retrieval over prior open and closed models.

  5. Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations

    cs.CV 2026-05 unverdicted novelty 7.0

    CoDAAR creates a unified discrete representation space for multimodal sequences by aligning modality-specific codebooks through index-level semantic consensus, enabling both specificity and cross-modal generalization.

  6. Cross-Modal-Domain Generalization Through Semantically Aligned Discrete Representations

    cs.CV 2026-05 unverdicted novelty 7.0

    CoDAAR aligns modality-specific codebooks at the index level using Discrete Temporal Alignment and Cascading Semantic Alignment to achieve cross-modal generalization while preserving unique structures, reporting state...

  7. Adapting MLLMs for Nuanced Video Retrieval

    cs.CV 2025-12 unverdicted novelty 7.0

    Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.

  8. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

    cs.CV 2022-04 unverdicted novelty 7.0

    Socratic Models compose zero-shot multimodal reasoning by prompting pretrained language and vision models to exchange information and enable new capabilities without finetuning.

  9. LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

    cs.CV 2026-07 conditional novelty 6.0

    Latent-action modeling (inverse dynamics plus forward world model) with a patch-level anti-collapse regularizer improves surgical action-triplet recognition and makes encoder change features land more on instrument-ti...

  10. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    cs.CV 2026-07 unverdicted novelty 6.0

    VideoSearch-R1 achieves SOTA on VCMR across three datasets via iterative retrieval, latent-space soft query refinement, and GRPO training.

  11. Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path Distillation

    cs.CV 2026-04 unverdicted novelty 6.0

    MDPD mutually distills knowledge between a frozen backbone and a learnable side network during fine-tuning, then discards the side network at inference to accelerate speed by at least 25% while preserving accuracy.

  12. MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning

    cs.LG 2026-04 unverdicted novelty 6.0

    MP-ISMoE uses Gaussian noise perturbed iterative quantization and interactive side mixture-of-experts to deliver higher accuracy than prior memory-efficient transfer learning methods while keeping similar parameter an...

  13. VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG

    cs.CV 2026-04 unverdicted novelty 6.0

    VideoStir introduces a spatio-temporal graph-based structure and intent-aware retrieval for long-video RAG, achieving competitive performance with SOTA methods via a new IR-600K dataset.

  14. LLaVA-Video: Video Instruction Tuning With Synthetic Data

    cs.CV 2024-10 unverdicted novelty 6.0

    LLaVA-Video-178K is a new synthetic video instruction dataset that, when combined with existing data to train LLaVA-Video, produces strong results on video understanding benchmarks.

  15. Demystifying CLIP Data

    cs.CV 2023-09 accept novelty 6.0

    MetaCLIP curates balanced 400M-pair subsets from CommonCrawl that outperform CLIP data, reaching 70.8% zero-shot ImageNet accuracy on ViT-B versus CLIP's 68.3%.

  16. Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

    cs.CV 2026-07 conditional novelty 5.0

    LLM-generated atomic action descriptions injected into scene-graph node features improve prompt-guided disentanglement for multi-label action recognition, with reported gains on Charades-oracle and SportsHHI but not o...

  17. Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

    cs.CV 2026-07 conditional novelty 5.0

    A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.

  18. LARE: Low-Attention Region Encoding for Text-Image Retrieval

    cs.CV 2026-06 unverdicted novelty 5.0

    LARE uses parallel encoding of full images and low-attention regions to improve text-image retrieval, shown on a new Dense-Set subset of COCO and Flickr30K with re-captioned overlooked areas.

  19. Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search

    cs.CV 2026-05 unverdicted novelty 5.0

    Inverse attention embeddings combined with standard visual features improve recall in video semantic search for crowded scenes without additional training.

  20. SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

    cs.CV 2024-01 unverdicted novelty 5.0

    SRL-CLIP uses rule-based captions derived from semantic role labels to adapt CLIP via contrastive fine-tuning on 23k pairs, matching or exceeding larger models trained on far more data across video tasks.

  21. Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing

    cs.CV 2026-07 conditional novelty 4.5

    Multi-scale temporal convolutions plus capsule routing improve video-text moment localization to 42.9% R@0.5 and 41.1% mIoU on ActivityNet Captions.