REVIEW 5 cited by
A CLIP-Hitchhiker's Guide to Long Video Retrieval
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A CLIP-Hitchhiker's Guide to Long Video Retrieval
read the original abstract
Our goal in this paper is the adaptation of image-text models for long video retrieval. Recent works have demonstrated state-of-the-art performance in video retrieval by adopting CLIP, effectively hitchhiking on the image-text representation for video tasks. However, there has been limited success in learning temporal aggregation that outperform mean-pooling the image-level representations extracted per frame by CLIP. We find that the simple yet effective baseline of weighted-mean of frame embeddings via query-scoring is a significant improvement above all prior temporal modelling attempts and mean-pooling. In doing so, we provide an improved baseline for others to compare to and demonstrate state-of-the-art performance of this simple baseline on a suite of long video retrieval benchmarks.
Forward citations
Cited by 5 Pith papers
-
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
Neighborhood re-ranking via Hungarian matching and query-conditioned local steering improve CLIP retrieval on attribute-binding and compositional tasks by addressing local geometric inconsistencies.
-
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
Short, simple captions describing single actions achieve higher retrieval recall than complex multi-step or fine-grained scene descriptions across all tested models.
-
SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels
SRL-CLIP uses rule-based captions derived from semantic role labels to adapt CLIP via contrastive fine-tuning on 23k pairs, matching or exceeding larger models trained on far more data across video tasks.
-
LEViL: Label-Efficient Video Learning via Zero-Shot Distillation over VLM-Generated Pseudo-Label Spaces
LEViL performs annotation-free video pretraining via VLM-generated pseudo-label spaces and zero-shot distillation, then uses target-aware fine-tuning to outperform semi-supervised baselines on UCF101 and HMDB51 in lim...
-
Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives
Survey summarizing video-language understanding tasks, challenges, and methods from architecture, training, and data perspectives, including performance comparisons and future directions.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.