Pith. sign in

REVIEW 5 cited by

A CLIP-Hitchhiker's Guide to Long Video Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.08508 v1 pith:IDKVDK63 submitted 2022-05-17 cs.CV

A CLIP-Hitchhiker's Guide to Long Video Retrieval

classification cs.CV
keywords videoretrievalbaselinelongclipframeimage-textmean-pooling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the image-text representation for video tasks. However, there has been limited success in learning temporal aggregation that outperform mean-pooling the image-level representations extracted per frame by CLIP. We find that the simple yet effective baseline of weighted-mean of frame embeddings via query-scoring is a significant improvement above all prior temporal modelling attempts and mean-pooling. In doing so, we provide an improved baseline for others to compare to and demonstrate state-of-the-art performance of this simple baseline on a suite of long video retrieval benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering

    cs.CV 2026-04 unverdicted novelty 7.0

    Neighborhood re-ranking via Hungarian matching and query-conditioned local steering improve CLIP retrieval on attribute-binding and compositional tasks by addressing local geometric inconsistencies.

  2. Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

    cs.IR 2026-03 unverdicted novelty 6.0

    Short, simple captions describing single actions achieve higher retrieval recall than complex multi-step or fine-grained scene descriptions across all tested models.

  3. SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels

    cs.CV 2024-01 unverdicted novelty 5.0

    SRL-CLIP uses rule-based captions derived from semantic role labels to adapt CLIP via contrastive fine-tuning on 23k pairs, matching or exceeding larger models trained on far more data across video tasks.

  4. LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces

    cs.CV 2026-06 unverdicted novelty 4.0

    LEViL performs annotation-free video pretraining via VLM-generated pseudo-label spaces and zero-shot distillation, then uses target-aware fine-tuning to outperform semi-supervised baselines on UCF101 and HMDB51 in lim...

  5. Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

    cs.CL 2024-06 unverdicted novelty 2.0

    Survey summarizing video-language understanding tasks, challenges, and methods from architecture, training, and data perspectives, including performance comparisons and future directions.