Pith. sign in

REVIEW 4 major objections 5 minor 64 references

A Flexible and Scalable Framework for Video Moment Search

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Video moment search drops to under one second per query

desk verdict SPR is a clean, honest framework for ranked video moment retrieval, but the SOTA claim needs independent baselines and a real long-video test before I trust the margin. read the letter →

arxiv 2501.05072 v1 pith:GT2ZKDWD submitted 2025-01-09 cs.IR cs.CV

classification cs.IRcs.CV
keywords rankedvideomomentretrievalsearchsegment-proposal-rankingdenseapproximatenearestneighborcontrastivelearningTVR-Rankingcorpus
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that ranked video moment search can be reduced to a three-stage pipeline: retrieve the most relevant 4-second segments from an offline index, merge adjacent retrieved segments into coarse proposals, then refine the timestamps and re-rank the few dozen proposals. The authors argue that this decomposition removes the 'single perfect match' assumption of earlier video moment retrieval, handles hour-long videos by construction, and keeps online cost near real time even on corpora of hundreds of thousands of segments. On the TVR-Ranking benchmark their SPRReLo instantiation reports NDCG@10 at IoU≥0.3 of 0.5509 on test, against 0.4353 for the ReLoCLNet baseline, with total query processing of roughly 0.7–1 second, and the coarse proposals alone already exceed the baseline at the IoU≥0.3 threshold. The flexible stage separation means any of the three components can be upgraded independently, which is the main practical payoff.

What carries the argument

The load-bearing object is the ordered trio of stages, Segment-Proposal-Ranking: (1) segment retrieval with a shared embedding space and cosine similarity, (2) rule-based proposal generation that merges temporally adjacent segments from the same video, and (3) a refinement and re-ranking head that outputs a matching score φ(Q,M) and predicted start/end timestamps (t_s,t_e). The mechanism that carries the efficiency claim is the fixed segment length of 4 seconds with offline precomputation and approximate nearest neighbor search, so online search cost depends on the number of retrieved segments (k=200), not on video length or corpus size. Training uses the MIL-NCE contrastive loss to align queries with the possibly multiple relevant segments in each batch, and the refinement module is trained with the same losses as ReLoCLNet—video contrastive and hinge loss for the score, frame contrastive and localization loss for boundaries—with an added 8-second context padding to raise the coverage ceiling.

What would settle it

Take a set of queries deliberately constructed so that relevance is only apparent across several segments (for example, 'the scene where she realizes the clue and then acts on it' with the ground truth spanning both parts), retrieve the top-200 segments, and compute how many ground-truth moments are contained in the resulting proposals; if that coverage falls far below the paper's reported upper bound of NDCG@10 IoU≥0.7 = 0.8842, then the assumption that segment-level retrieval suffices is false.

Watch

Extended reading notes

Core claim

The central claim is that a ranked list of video moments for a text query can be obtained by first treating each fixed-length video segment as an independent retrieval unit. Segments are embedded with CLIP, a frozen vision-language model, projected into a shared query–segment space learned with the MIL-NCE contrastive objective, and stored in a Faiss ANN index built offline; at query time only the top-200 segments are fetched, so no matter how long the video or how large the corpus, online work is bounded by 200 segments plus the proposals they form. Adjacent segments from the same video are merged into coarse proposals, and a light refinement module—instantiated either with CLIP features or with the ReLoCLNet encoder trained with contrastive and hinge losses—adjusts boundaries and re-scores moments. The paper's evidence is that this two-level (coarse then fine) design beats all prior RVMR baselines on TVR-Ranking by a substantial margin, cuts average per-query time to under one second across 19,614 videos, and degrades only slightly when the corpus is tripled with irrelevant videos.

Load-bearing premise

The whole pipeline assumes that any ground-truth moment is well covered by a handful of individually query-relevant 4-second segments, so that searching the top-200 segments never discards the truth before refinement begins.

Editorial extensions

If this is right

  • If the framework's claims hold, ranked moment search becomes practical at the scale of tens of thousands of videos, with per-query latency of about one second on a single GPU.
  • Because the three stages are independent, improvements in embedding models, ANN indexes, or refinement architectures can be dropped in without re-architecting the rest of the system.
  • The upper-bound analysis shows that even a perfect refinement module cannot recover moments that fall outside the top-200 retrieved segments, so improving segment recall is the direct path to raising the performance ceiling.
  • The scalability experiments indicate that adding large numbers of unrelated videos does not hurt ranking quality when using Flat or IVF indexes, supporting deployment on growing corpora.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the same segment-proposal-rank recipe should transfer to other sub-second temporal units and other modalities (audio, subtitles) as long as the segment embedding captures query-relevant content; the paper only evaluates visual features, mostly CLIP frames and I3D/RoBERTa subtitle streams in one instantiation.
  • Beyond the paper: the framework's dependence on per-segment relevance means it is likely to miss 'emergent' moments whose relevance is defined by a multi-segment relationship, such as cause-effect chains across scenes; a natural test is whether top-200 segment recall on such queries falls well below the reported upper bound of NDCG@10 IoU≥0.7 = 0.8842.
  • Beyond the paper: the reported collapse of the IVFPQ index (NDCG dropping from roughly 0.44 to 0.09) suggests that product quantization is unsafe for this embedding distribution; a testable extension is to add a two-stage re-rank over an IVF shortlist to recover accuracy at comparable speed.
  • Beyond the paper: the pseudo-training set built with SimCSE similarity is a weak supervision signal, and a dataset with manual relevance grades across moments would likely close part of the gap between current results and the practical upper bound.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes SPR, a three-stage framework for Ranked Video Moment Retrieval (RVMR): offline indexing of fixed-length video segments, online top-k segment retrieval with approximate nearest neighbor search, rule-based merging of retrieved segments into coarse proposals, and a refinement/re-ranking module that adjusts timestamps and re-ranks proposals. The framework is evaluated on the TVR-Ranking dataset, with two instantiations (SPRReLo and SPRCLIP) reporting improvements over published baselines, together with efficiency and scalability measurements and an upper-bound analysis of the segment-retrieval stage. The central claims are that SPR reaches state-of-the-art performance on TVR-Ranking while processing a query in about one second and scaling to nearly a million indexed segments.

Significance. If the empirical claims hold, the paper makes a useful contribution: it addresses a realistic retrieval setting where multiple moments can match a query, and it does so with a modular pipeline that is substantially faster than end-to-end moment-localization models. The strengths of the paper are its clean decomposition into three stages, the inclusion of an explicit upper-bound analysis (Table 4) that separates the retrieval bottleneck from the refinement bottleneck, the efficiency and scalability measurements (Tables 2 and 3), detailed ablations of the training strategies (Table 5), and the release of code. The main reservations concern verification: the SOTA comparison relies entirely on baseline numbers taken from the same group's dataset paper, and no variance or significance information is reported, so the margin claimed in the abstract is not yet fully established.

major comments (4)
  1. [4.3, Table 1] The headline state-of-the-art claim rests entirely on baseline numbers sourced from [31], a paper that shares authors with this submission. Since no baseline is re-run and no variance or significance information is reported, the large margin (e.g., 0.5509 vs 0.4353 for NDCG@10 with IoU>=0.3) cannot be distinguished from favorable configuration or evaluation differences. Please provide at least one independent re-run of the baselines, or report multiple seeds with means and standard deviations for all methods, and explicitly verify that the pseudo-training configuration is identical for all models.
  2. [4.4, Table 2] The abstract claims significant reductions in computational cost and processing time, but Table 2 reports only absolute timings for the SPR stages. There is no comparison with the inference time of XML, CONQUER, or ReLoCLNet under the same hardware. Without such a comparison, the efficiency advantage over existing RVMR methods is not established. Please add wall-clock timings for the baselines or restrict the claim to absolute throughput.
  3. [3.3.2 and 4.2] The training procedure for the refinement and re-ranking module is underspecified. The text says it follows [59], but [59] is trained on full-video ground-truth moments, whereas here the module consumes coarse proposals produced by the first two stages. The paper does not define how training proposals are generated from the pseudo-training set, how positive and negative proposals are sampled within a batch, or how the strong/weak positive selection described in Section 4.7 is performed at training time. This information is needed to reproduce Table 5 and to audit the gains attributed to the added hinge losses.
  4. [3.1, Table 4] The framework's key premise is that every relevant moment is covered by a small set of individually query-relevant segments. On TVR-Ranking this premise is empirically well-supported by the upper-bound analysis (UB NDCG@10 with IoU>=0.7 reaches 0.8842 with tau_C=8), so I do not regard it as a fatal flaw. However, the paper should explicitly acknowledge that for queries whose relevance is only visible at the multi-segment level, top-k segment retrieval cannot recover the moment, and it should either report failure cases or soften the claim that the framework handles videos of any length in general.
minor comments (5)
  1. [Throughout] There are several typos: 'additinal' in Table 3's caption, 'an text query' in Section 1, and repeated 'e.g.,,' in the introduction and Figure 1 caption.
  2. [4.4, Table 2] The formatting of Table 2 is confusing: the SPIVF and SPIVFPQ rows appear to have a different number of columns than the other rows, making the per-stage time columns ambiguous.
  3. [5, Table 5] In Table 5, the row labeled 'SP' has a dash in the Model group column; it would be clearer to label it 'SP (coarse, no refinement)' to avoid confusion with the SPR variants.
  4. [1, Table 1] The caption of Table 1 says 'SP, SPR CLIP and SPR ReLo' but the table uses underscores; please make the notation consistent across text, tables, and figures.
  5. [4.5, Table 3] The caption of Table 3 contains a fragment 'T: TVR-Ranking validation dataset additinal videos from C: Charades'; please rephrase it as a complete sentence and fix the typo.

Circularity Check

0 steps flagged · score 2.0 of 10

No derivation-level circularity; only same-group baseline sourcing in Table 1, a verification gap, not a by-construction equivalence.

full rationale

This is an empirical systems paper, not a derivation: SPR is a three-stage pipeline, and its reported numbers are measurements on the TVR-Ranking test set. I walked the claimed chain: segment retrieval trains two projectors with MIL-NCE on pseudo-labels from [31]; proposal generation is a fixed merging rule; refinement/re-ranking reuses ReLoCLNet [59] and CLIP instantiations. None of the predicted quantities is defined in terms of the metric it is scored by, and no fitted parameter is renamed as a prediction. The only load-bearing external input is the benchmark itself: 'All baseline model results are sourced from [31]' (Sec. 4.3), and [31] shares three authors with this paper (Liang et al., including Chongzhi Zhang, Xizhou Zhu, Aixin Sun). That makes the SOTA comparison self-referential in provenance and would benefit from independent reproduction or error bars, but the baseline numbers are fixed empirical values, not consequences of SPR's assumptions. The paper's key coverage assumption is also explicitly quantified as the upper bound in Table 4, so it is not smuggled in as a result. Hence no circular step can be exhibited; score 2 reflects the minor self-citation caveat, not a by-construction circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The paper is a system paper, so the axiom ledger is dominated by tuned hyperparameters and dataset assumptions. The five free parameters (segment length, number of retrieved segments, context padding, number of pseudo moments, overlap threshold) directly shape the reported NDCG values and were selected on the validation split. The key structural assumption is that relevant moments are covered by individually relevant segments; Table 4's upper bound quantifies the ceiling this assumption imposes. No invented entities are introduced.

free parameters (5)
  • segment length tau_S = 4 seconds
    Sets the minimum granularity of proposals; chosen by hand with no ablation reported (Section 4.2).
  • number of retrieved segments k = 200
    Tuned on validation via Table 6; performance peaks at 200.
  • context padding tau_C = 8 seconds
    Chosen based on upper bound analysis in Table 4, where UB peaks at tau_C=8.
  • pseudo-training overlap threshold = 0.3
    Tuned on validation in Table 7; filters segments with less than 30% overlap with moment annotations.
  • number of pseudo-positive moments N = 40
    Tuned on validation in Table 7; top-40 moments per query are used to train the segment retrieval module.
assumptions (4)
  • domain assumption TVR-Ranking relevance annotations are a valid ground truth for ranked video moment retrieval
    The dataset and task are introduced in [31] by the same group; all evaluation treats the manual relevance levels 1-4 and IoU matching as ground truth.
  • domain assumption The pseudo-training set provides a useful training signal for segment retrieval
    Section 4.1: pseudo-queries are assigned up to 40 moments by SimCSE query-caption similarity without manual relevance; the segment retrieval projectors (Section 3.1.3) are trained with MIL-NCE on this noisy signal, and the overlap threshold (0.3) is tuned on validation.
  • domain assumption Relevant moments decompose into individually query-relevant 4-second segments
    The pipeline's ceiling is set by top-k segment retrieval; the upper-bound analysis in Table 4 assumes any ground-truth moment contained in a coarse proposal can be perfectly refined. If a query's relevance is only apparent across multiple segments, the retrieval stage can miss the moment.
  • domain assumption Frozen CLIP ViT-L/14 provides sufficient visual-text alignment for both segment and query encoding
    Section 3.1.3 and 4.2 use frozen CLIP for frame and query features; the learned projectors are trained on top of it. The ablation in Table 7 shows untrained CLIP performs poorly, so the assumption is that training the projectors fixes the alignment gap, which is empirically validated only on TVR-Ranking.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Flexible and Scalable Framework for Video Moment Search." pith.science (2026). https://pith.science/paper/GT2ZKDWD

@misc{pith2026250105072,
  author       = {Pith},
  title        = {Pith review of: A Flexible and Scalable Framework for Video Moment Search},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GT2ZKDWD}},
  note         = {Machine review of arXiv:2501.05072}
}
read the original abstract

Video moment search, the process of finding relevant moments in a video corpus to match a user's query, is crucial for various applications. Existing solutions, however, often assume a single perfect matching moment, struggle with inefficient inference, and have limitations with hour-long videos. This paper introduces a flexible and scalable framework for retrieving a ranked list of moments from collection of videos in any length to match a text query, a task termed Ranked Video Moment Retrieval (RVMR). Our framework, called Segment-Proposal-Ranking (SPR), simplifies the search process into three independent stages: segment retrieval, proposal generation, and moment refinement with re-ranking. Specifically, videos are divided into equal-length segments with precomputed embeddings indexed offline, allowing efficient retrieval regardless of video length. For scalable online retrieval, both segments and queries are projected into a shared feature space to enable approximate nearest neighbor (ANN) search. Retrieved segments are then merged into coarse-grained moment proposals. Then a refinement and re-ranking module is designed to reorder and adjust timestamps of the coarse-grained proposals. Evaluations on the TVR-Ranking dataset demonstrate that our framework achieves state-of-the-art performance with significant reductions in computational cost and processing time. The flexible design also allows for independent improvements to each stage, making SPR highly adaptable for large-scale applications.

Figures

Figures reproduced from arXiv: 2501.05072 by the authors.

Figure 1
Figure 1. The Segment-Proposal-Ranking (SPR) framework. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Segment retrieval. With the offline constructed [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Refinement and re-ranking. This module computes [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Visualization of the ground truth moment and the results from SP and SPR for two example queries from the [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

64 extracted references · 20 canonical work pages

  1. [31]

    Renjie Liang, Li Li, Chongzhi Zhang, Jing Wang, Xizhou Zhu, and Aixin Sun. 2024. TVR-Ranking: A Dataset for Ranked Video Moment Retrieval with Imprecise Queries. arXiv:2407.06597 [cs.AI] https://arxiv.org/abs/2407.06597

  2. [59]

    Hao Zhang, Aixin Sun, Wei Jing, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh. 2021. Video Corpus Moment Retrieval with Contrastive Learning. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, Fernando Diaz, Chirag Shah, Torsten Suel...

  3. [2]

    Bruce Croft, and Yanlei Diao

    Michael Bendersky, W. Bruce Croft, and Yanlei Diao. 2011. Quality-biased ranking of web documents. In Proceedings of the Forth International Conference on Web Search and Web Data Mining, WSDM 2011, Hong Kong, China, February 9-12, 2011, Irwin King, Wolfgang Nejdl, and Hang Li (Eds.). ACM, 95–104. https: //doi.org/10.1145/1935826.1935849

  4. [3]

    Rae, Erich Elsen, and Laurent Sifre

    Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Ruther- ford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, O...

  5. [4]

    Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005. Learning to rank using gradient descent. InProceedings of the 22nd International Conference on Machine Learning (Bonn, Germany) (ICML ’05). Association for Computing Machinery, New York, NY, USA, 89–96. https: //doi.org/10.1145/1102351.1102363

  6. [5]

    Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021. On Pursuit of Designing Multi-modal Transformer for Video Grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Francine Moens, Xuanjing Huang, Lucia ...

  7. [6]

    Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308

  8. [7]

    Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W. Cohen. 2023. Re- Imagen: Retrieval-Augmented Text-to-Image Generator. InThe Eleventh Interna- tional Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=XSEBx0iSjFQ

Show all 64 references
  1. [8]

    Crandall, Mohit Bansal, and Gedas Bertasius

    Feng Cheng, Xizi Wang, Jie Lei, David J. Crandall, Mohit Bansal, and Gedas Bertasius. 2023. VindLU: A Recipe for Effective Video-and-Language Pretraining. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEE...

  2. [9]

    Ritendra Datta, Dhiraj Joshi, Jia Li, and James Ze Wang. 2008. Image retrieval: Ideas, influences, and trends of the new age. ACM Comput. Surv. 40, 2 (2008), 5:1–5:60. https://doi.org/10.1145/1348246.1348248

  3. [10]

    Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  4. [11]

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2024. The Faiss library. CoRR abs/2401.08281 (2024). https://doi.org/10.48550/ARXIV. 2401.08281 arXiv:2401.08281

  5. [12]

    Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan C. Russell

  6. [13]

    Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017. TALL: Temporal Activity Localization via Language Query. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . IEEE Computer Society, 5277–5285. https://doi.org/10.1109/I...

  7. [14]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech...

  8. [15]

    Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, ...

  9. [16]

    Deepak Gupta, Kush Attal, and Dina Demner-Fushman. 2024. Towards Answering Health-related Questions from Medical Videos: Datasets and Approaches. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COL...

  10. [17]

    Tanveer Hannan, Md Mohaiminul Islam, Thomas Seidl, and Gedas Bertasius

  11. [18]

    Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017. Localizing Moments in Video with Natural Language. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017. IEEE Computer Societ...

  12. [19]

    Zhijian Hou, Chong-Wah Ngo, and Wing Kwong Chan. 2021. CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval. In MM ’21: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021 , Heng Tao Shen, Yueting Zhuang, John R. Smith, Yang Yang, Pablo ...

  13. [20]

    Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wing Kwong Chan, Chong-Wah Ngo, Mike Zheng Shou, and Nan Duan. 2023. CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding. In Proceedings of the 61st Annual Meeting of the Association for...

  14. [21]

    Piotr Indyk and Rajeev Motwani. 1998. Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality. In Proceedings of the Thirtieth Annual ACM Symposium on the Theory of Computing, Dallas, Texas, USA, May 23-26, 1998 , Jeffrey Scott Vitter (Ed.). ACM, 604–613. h...

  15. [22]

    Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Trans. Pattern Anal. Mach. Intell. 33, 1 (2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57

  16. [23]

    Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Super- vision. In Proceedings of the 38th International Conference on...

  17. [24]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ...

  18. [25]

    Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, C...

  19. [26]

    Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles

  20. [27]

    Lafferty and ChengXiang Zhai

    John D. Lafferty and ChengXiang Zhai. 2001. Document Language Models, Query Models, and Risk Minimization for Information Retrieval. In SIGIR 2001: Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, September...

  21. [28]

    Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2020. Tvr: A large-scale dataset for video-subtitle moment retrieval. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16 . Springer, 447–463

  22. [29]

    Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...

  23. [30]

    Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu

  24. [32]

    Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692 (2019)

  25. [33]

    Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu, and Ge Yu. 2023. Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal Retrieval. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwand...

  26. [34]

    Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Or- leans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum? id=Bkg6RiCqY7

  27. [35]

    Siyu Lou, Xuenan Xu, Mengyue Wu, and Kai Yu. 2022. Audio-Text Retrieval in Context. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022 . IEEE, 4793–4797. https://doi.org/10.1109/ICASSP43922.2022.9746786

  28. [36]

    Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning. Neurocomputing 508 (2022), 293–304. https://doi.org/10.1016/J. NEUCOM.2022.07.028

  29. [37]

    Craig Macdonald, Rodrygo L. T. Santos, and Iadh Ounis. 2013. The whens and hows of learning to rank for web search. Inf. Retr. 16, 5 (2013), 584–628. https://doi.org/10.1007/S10791-012-9209-9

  30. [38]

    Malkov and Dmitry A

    Yury A. Malkov and Dmitry A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs.IEEE Trans. Pattern Anal. Mach. Intell. 42, 4 (2020), 824–836. https://doi.org/10.1109/ TPAMI.2018.2889473

  31. [39]

    Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to-End Learning of Visual Representations From Uncurated Instructional Videos. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seatt...

  32. [40]

    Roy- Chowdhury

    Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K. Roy- Chowdhury. 2018. Learning Joint Embedding with Multimodal Cues for Cross- Modal Video-Text Retrieval. In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, ICMR 2018, Yokoham...

  33. [41]

    Behnam Neyshabur and Nathan Srebro. 2015. On Symmetric and Asymmetric LSHs for Inner Product Search. InProceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 (JMLR Workshop and Conference Proceedings, Vol. 37) , Francis R...

  34. [43]

    Xiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng, Jianfeng Dong, Pan Zhou, and Zichuan Xu. 2020. Fine-grained Iterative Attention Network for Tem- poral Language Localization in Videos. In MM ’20: The 28th ACM Interna- tional Conference on Multimedia, Virtual Event / Seattle, W ...

  35. [44]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...

  36. [45]

    Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013. Grounding Action Descriptions in Videos. Trans. Assoc. Comput. Linguistics 1 (2013), 25–36. https://doi.org/10.1162/TACL_ A_00207

  37. [46]

    Robertson and Steve Walker

    Stephen E. Robertson and Steve Walker. 1997. On Relevance Weights with Little Relevance Information. In SIGIR ’97: Proceedings of the 20th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 27-31, 1997, Philadelphia, PA, USA , ...

  38. [47]

    Robertson and Hugo Zaragoza

    Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. 3, 4 (2009), 333–389. https://doi.org/10.1561/1500000019

  39. [48]

    Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta

    Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14...

  40. [49]

    Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron, Chen Zhao, Silvio Giancola, and Bernard Ghanem. 2022. MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions. In IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  41. [50]

    Nicola Tonellotto. 2022. Lecture Notes on Neural Information Retrieval. CoRR abs/2207.13443 (2022). https://doi.org/10.48550/ARXIV.2207.13443 arXiv:2207.13443

  42. [51]

    Hao Wang, Zheng-Jun Zha, Liang Li, Dong Liu, and Jiebo Luo. 2021. Structured Multi-Level Interaction Network for Video Moment Localization via Language Query. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision ...

  43. [52]

    Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. In IEEE International Conference on Acoustics, Speech and Signal Processi...

  44. [54]

    Bennett, Junaid Ahmed, and Arnold Overwijk

    Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In 9th International Conference on Learning Representations, ICLR 2021, V...

  45. [55]

    Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. In Proceedings of the 2021 Conference on Empirical Methods in Nat...

  46. [56]

    Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko

    Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019. Multilevel Language and Vision Integration for Text-to-Clip Re- trieval. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Application...

  47. [57]

    Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, and Zhen- zhen Jiao. 2024. Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W...

  48. [58]

    Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma

  49. [60]

    Hao Zhang, Aixin Sun, Wei Jing, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh. 2021. Parallel Attention Network with Sequence Matching for Video Grounding. In Findings of the Association for Computational Lin- guistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 ...

  50. [61]

    Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2020. Span-based Lo- calizing Network for Natural Language Video Localization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Jo...

  51. [62]

    Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, and Ziwei Liu. 2023. ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model. InIEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 . IEEE, ...

  52. [63]

    Top-20" and “Top-40

    Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. 2020. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artific...

  53. [2017]

    In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017

    Dense-Captioning Events in Videos. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . IEEE Computer Society, 706–715. https://doi.org/10.1109/ICCV.2017.83

  54. [2019]

    CoRR abs/1907.12763 (2019)

    Temporal Localization of Moments in Video Collections with Natural Language. CoRR abs/1907.12763 (2019). arXiv:1907.12763 http://arxiv.org/abs/ 1907.12763

  55. [2020]

    HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Asso...

  56. [2025]

    In Computer Vision – ECCV 2024 , Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.)

    RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos. In Computer Vision – ECCV 2024 , Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.). Springer Nature Switzerland, Cham, 352–369

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.