Pith. sign in

REVIEW 7 cited by

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.06430 v4 pith:K3QLGRWC submitted 2022-09-14 cs.CV

CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

classification cs.CV
keywords clipclip-vippre-trainedimage-textmodelrepresentationresultsvideo-language
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The pre-trained image-text models, like CLIP, have demonstrated the strong power of vision-language representation learned from a large scale of web-collected image-text data. In light of the well-learned visual features, some existing works transfer image representation to video domain and achieve good results. However, how to utilize image-language pre-trained model (e.g., CLIP) for video-language pre-training (post-pretraining) is still under explored. In this paper, we investigate two questions: 1) what are the factors hindering post-pretraining CLIP to further improve the performance on video-language tasks? and 2) how to mitigate the impact of these factors? Through a series of comparative experiments and analyses, we find that the data scale and domain gap between language sources have great impacts. Motivated by these, we propose a Omnisource Cross-modal Learning method equipped with a Video Proxy mechanism on the basis of CLIP, namely CLIP-ViP. Extensive results show that our approach improves the performance of CLIP on video-text retrieval by a large margin. Our model also achieves SOTA results on a variety of datasets, including MSR-VTT, DiDeMo, LSMDC, and ActivityNet. We will release our code and pre-trained CLIP-ViP models at https://github.com/microsoft/XPretrain/tree/main/CLIP-ViP.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Adapting MLLMs for Nuanced Video Retrieval

    cs.CV 2025-12 unverdicted novelty 7.0

    Text-only contrastive fine-tuning of an MLLM with hard negatives produces embeddings that handle temporal, negation, and multimodal nuances in video retrieval and achieves SOTA performance.

  2. Adversarial Video Promotion Against Text-to-Video Retrieval

    cs.CV 2025-08 unverdicted novelty 7.0

    Pioneers ViPro, the first attack to adversarially promote videos in text-to-video retrieval, using Modal Refinement to improve black-box transferability across multiple targets.

  3. VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement

    cs.CV 2026-07 unverdicted novelty 6.0

    VideoSearch-R1 achieves SOTA on VCMR across three datasets via iterative retrieval, latent-space soft query refinement, and GRPO training.

  4. CoVR-R:Reason-Aware Composed Video Retrieval

    cs.CV 2026-03 conditional novelty 6.0

    Zero-shot LMM reasoning over edit after-effects (states, phases, camera, tempo) plus a new CoVR-R benchmark yields large recall gains on implicit-effect composed video retrieval without task-specific training.

  5. Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

    cs.CV 2026-07 conditional novelty 5.0

    LLM-generated atomic action descriptions injected into scene-graph node features improve prompt-guided disentanglement for multi-label action recognition, with reported gains on Charades-oracle and SportsHHI but not o...

  6. Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing

    cs.CV 2026-07 conditional novelty 4.5

    Multi-scale temporal convolutions plus capsule routing improve video-text moment localization to 42.9% R@0.5 and 41.1% mIoU on ActivityNet Captions.

  7. A Survey on Foundation Models for Personalized Federated Intelligence

    cs.AI 2025-05 unverdicted novelty 3.0

    The survey introduces personalized federated intelligence (PFI) as a framework integrating federated learning and foundation models to support privacy-aware personalization of AI models.