Pith. sign in

REVIEW 1 cited by

An Empirical Study of Frame Selection for Text-to-Video Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.00298 v1 pith:U7T2JKTP submitted 2023-11-01 cs.CV

classification cs.CV
keywords videoframeretrievalefficiencyframesselectionselectionsempirical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-video retrieval (TVR) aims to find the most relevant video in a large video gallery given a query text. The intricate and abundant context of the video challenges the performance and efficiency of TVR. To handle the serialized video contexts, existing methods typically select a subset of frames within a video to represent the video content for TVR. How to select the most representative frames is a crucial issue, whereby the selected frames are required to not only retain the semantic information of the video but also promote retrieval efficiency by excluding temporally redundant frames. In this paper, we make the first empirical study of frame selection for TVR. We systemically classify existing frame selection methods into text-free and text-guided ones, under which we detailedly analyze six different frame selections in terms of effectiveness and efficiency. Among them, two frame selections are first developed in this paper. According to the comprehensive analysis on multiple TVR benchmarks, we empirically conclude that the TVR with proper frame selections can significantly improve the retrieval efficiency without sacrificing the retrieval performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SVTA is a large synthetic video-text benchmark with 41,315 videos covering 68 anomaly types, created by LLM-generated captions and text-to-video generation, and evaluated with three retrieval baselines.

Pith tools