REVIEW 4 major objections 5 minor 64 references
A Flexible and Scalable Framework for Video Moment Search
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Video moment search drops to under one second per query
desk verdict SPR is a clean, honest framework for ranked video moment retrieval, but the SOTA claim needs independent baselines and a real long-video test before I trust the margin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the ordered trio of stages, Segment-Proposal-Ranking: (1) segment retrieval with a shared embedding space and cosine similarity, (2) rule-based proposal generation that merges temporally adjacent segments from the same video, and (3) a refinement and re-ranking head that outputs a matching score φ(Q,M) and predicted start/end timestamps (t_s,t_e). The mechanism that carries the efficiency claim is the fixed segment length of 4 seconds with offline precomputation and approximate nearest neighbor search, so online search cost depends on the number of retrieved segments (k=200), not on video length or corpus size. Training uses the MIL-NCE contrastive loss to align queries with the possibly multiple relevant segments in each batch, and the refinement module is trained with the same losses as ReLoCLNet—video contrastive and hinge loss for the score, frame contrastive and localization loss for boundaries—with an added 8-second context padding to raise the coverage ceiling.
What would settle it
Take a set of queries deliberately constructed so that relevance is only apparent across several segments (for example, 'the scene where she realizes the clue and then acts on it' with the ground truth spanning both parts), retrieve the top-200 segments, and compute how many ground-truth moments are contained in the resulting proposals; if that coverage falls far below the paper's reported upper bound of NDCG@10 IoU≥0.7 = 0.8842, then the assumption that segment-level retrieval suffices is false.
Extended reading notes
Core claim
The central claim is that a ranked list of video moments for a text query can be obtained by first treating each fixed-length video segment as an independent retrieval unit. Segments are embedded with CLIP, a frozen vision-language model, projected into a shared query–segment space learned with the MIL-NCE contrastive objective, and stored in a Faiss ANN index built offline; at query time only the top-200 segments are fetched, so no matter how long the video or how large the corpus, online work is bounded by 200 segments plus the proposals they form. Adjacent segments from the same video are merged into coarse proposals, and a light refinement module—instantiated either with CLIP features or with the ReLoCLNet encoder trained with contrastive and hinge losses—adjusts boundaries and re-scores moments. The paper's evidence is that this two-level (coarse then fine) design beats all prior RVMR baselines on TVR-Ranking by a substantial margin, cuts average per-query time to under one second across 19,614 videos, and degrades only slightly when the corpus is tripled with irrelevant videos.
Load-bearing premise
The whole pipeline assumes that any ground-truth moment is well covered by a handful of individually query-relevant 4-second segments, so that searching the top-200 segments never discards the truth before refinement begins.
Editorial extensions
If this is right
- If the framework's claims hold, ranked moment search becomes practical at the scale of tens of thousands of videos, with per-query latency of about one second on a single GPU.
- Because the three stages are independent, improvements in embedding models, ANN indexes, or refinement architectures can be dropped in without re-architecting the rest of the system.
- The upper-bound analysis shows that even a perfect refinement module cannot recover moments that fall outside the top-200 retrieved segments, so improving segment recall is the direct path to raising the performance ceiling.
- The scalability experiments indicate that adding large numbers of unrelated videos does not hurt ranking quality when using Flat or IVF indexes, supporting deployment on growing corpora.
Reading between the lines
- Beyond the paper: the same segment-proposal-rank recipe should transfer to other sub-second temporal units and other modalities (audio, subtitles) as long as the segment embedding captures query-relevant content; the paper only evaluates visual features, mostly CLIP frames and I3D/RoBERTa subtitle streams in one instantiation.
- Beyond the paper: the framework's dependence on per-segment relevance means it is likely to miss 'emergent' moments whose relevance is defined by a multi-segment relationship, such as cause-effect chains across scenes; a natural test is whether top-200 segment recall on such queries falls well below the reported upper bound of NDCG@10 IoU≥0.7 = 0.8842.
- Beyond the paper: the reported collapse of the IVFPQ index (NDCG dropping from roughly 0.44 to 0.09) suggests that product quantization is unsafe for this embedding distribution; a testable extension is to add a two-stage re-rank over an IVF shortlist to recover accuracy at comparable speed.
- Beyond the paper: the pseudo-training set built with SimCSE similarity is a weak supervision signal, and a dataset with manual relevance grades across moments would likely close part of the gap between current results and the practical upper bound.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SPR, a three-stage framework for Ranked Video Moment Retrieval (RVMR): offline indexing of fixed-length video segments, online top-k segment retrieval with approximate nearest neighbor search, rule-based merging of retrieved segments into coarse proposals, and a refinement/re-ranking module that adjusts timestamps and re-ranks proposals. The framework is evaluated on the TVR-Ranking dataset, with two instantiations (SPRReLo and SPRCLIP) reporting improvements over published baselines, together with efficiency and scalability measurements and an upper-bound analysis of the segment-retrieval stage. The central claims are that SPR reaches state-of-the-art performance on TVR-Ranking while processing a query in about one second and scaling to nearly a million indexed segments.
Significance. If the empirical claims hold, the paper makes a useful contribution: it addresses a realistic retrieval setting where multiple moments can match a query, and it does so with a modular pipeline that is substantially faster than end-to-end moment-localization models. The strengths of the paper are its clean decomposition into three stages, the inclusion of an explicit upper-bound analysis (Table 4) that separates the retrieval bottleneck from the refinement bottleneck, the efficiency and scalability measurements (Tables 2 and 3), detailed ablations of the training strategies (Table 5), and the release of code. The main reservations concern verification: the SOTA comparison relies entirely on baseline numbers taken from the same group's dataset paper, and no variance or significance information is reported, so the margin claimed in the abstract is not yet fully established.
major comments (4)
- [4.3, Table 1] The headline state-of-the-art claim rests entirely on baseline numbers sourced from [31], a paper that shares authors with this submission. Since no baseline is re-run and no variance or significance information is reported, the large margin (e.g., 0.5509 vs 0.4353 for NDCG@10 with IoU>=0.3) cannot be distinguished from favorable configuration or evaluation differences. Please provide at least one independent re-run of the baselines, or report multiple seeds with means and standard deviations for all methods, and explicitly verify that the pseudo-training configuration is identical for all models.
- [4.4, Table 2] The abstract claims significant reductions in computational cost and processing time, but Table 2 reports only absolute timings for the SPR stages. There is no comparison with the inference time of XML, CONQUER, or ReLoCLNet under the same hardware. Without such a comparison, the efficiency advantage over existing RVMR methods is not established. Please add wall-clock timings for the baselines or restrict the claim to absolute throughput.
- [3.3.2 and 4.2] The training procedure for the refinement and re-ranking module is underspecified. The text says it follows [59], but [59] is trained on full-video ground-truth moments, whereas here the module consumes coarse proposals produced by the first two stages. The paper does not define how training proposals are generated from the pseudo-training set, how positive and negative proposals are sampled within a batch, or how the strong/weak positive selection described in Section 4.7 is performed at training time. This information is needed to reproduce Table 5 and to audit the gains attributed to the added hinge losses.
- [3.1, Table 4] The framework's key premise is that every relevant moment is covered by a small set of individually query-relevant segments. On TVR-Ranking this premise is empirically well-supported by the upper-bound analysis (UB NDCG@10 with IoU>=0.7 reaches 0.8842 with tau_C=8), so I do not regard it as a fatal flaw. However, the paper should explicitly acknowledge that for queries whose relevance is only visible at the multi-segment level, top-k segment retrieval cannot recover the moment, and it should either report failure cases or soften the claim that the framework handles videos of any length in general.
minor comments (5)
- [Throughout] There are several typos: 'additinal' in Table 3's caption, 'an text query' in Section 1, and repeated 'e.g.,,' in the introduction and Figure 1 caption.
- [4.4, Table 2] The formatting of Table 2 is confusing: the SPIVF and SPIVFPQ rows appear to have a different number of columns than the other rows, making the per-stage time columns ambiguous.
- [5, Table 5] In Table 5, the row labeled 'SP' has a dash in the Model group column; it would be clearer to label it 'SP (coarse, no refinement)' to avoid confusion with the SPR variants.
- [1, Table 1] The caption of Table 1 says 'SP, SPR CLIP and SPR ReLo' but the table uses underscores; please make the notation consistent across text, tables, and figures.
- [4.5, Table 3] The caption of Table 3 contains a fragment 'T: TVR-Ranking validation dataset additinal videos from C: Charades'; please rephrase it as a complete sentence and fix the typo.
Circularity Check
No derivation-level circularity; only same-group baseline sourcing in Table 1, a verification gap, not a by-construction equivalence.
full rationale
This is an empirical systems paper, not a derivation: SPR is a three-stage pipeline, and its reported numbers are measurements on the TVR-Ranking test set. I walked the claimed chain: segment retrieval trains two projectors with MIL-NCE on pseudo-labels from [31]; proposal generation is a fixed merging rule; refinement/re-ranking reuses ReLoCLNet [59] and CLIP instantiations. None of the predicted quantities is defined in terms of the metric it is scored by, and no fitted parameter is renamed as a prediction. The only load-bearing external input is the benchmark itself: 'All baseline model results are sourced from [31]' (Sec. 4.3), and [31] shares three authors with this paper (Liang et al., including Chongzhi Zhang, Xizhou Zhu, Aixin Sun). That makes the SOTA comparison self-referential in provenance and would benefit from independent reproduction or error bars, but the baseline numbers are fixed empirical values, not consequences of SPR's assumptions. The paper's key coverage assumption is also explicitly quantified as the upper bound in Table 4, so it is not smuggled in as a result. Hence no circular step can be exhibited; score 2 reflects the minor self-citation caveat, not a by-construction circularity.
Assumptions & free parameters
free parameters (5)
- segment length tau_S =
4 seconds
- number of retrieved segments k =
200
- context padding tau_C =
8 seconds
- pseudo-training overlap threshold =
0.3
- number of pseudo-positive moments N =
40
assumptions (4)
- domain assumption TVR-Ranking relevance annotations are a valid ground truth for ranked video moment retrieval
- domain assumption The pseudo-training set provides a useful training signal for segment retrieval
- domain assumption Relevant moments decompose into individually query-relevant 4-second segments
- domain assumption Frozen CLIP ViT-L/14 provides sufficient visual-text alignment for both segment and query encoding
Cite this review
Pith. "Pith review of A Flexible and Scalable Framework for Video Moment Search." pith.science (2026). https://pith.science/paper/GT2ZKDWD
@misc{pith2026250105072,
author = {Pith},
title = {Pith review of: A Flexible and Scalable Framework for Video Moment Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/GT2ZKDWD}},
note = {Machine review of arXiv:2501.05072}
}
read the original abstract
Video moment search, the process of finding relevant moments in a video corpus to match a user's query, is crucial for various applications. Existing solutions, however, often assume a single perfect matching moment, struggle with inefficient inference, and have limitations with hour-long videos. This paper introduces a flexible and scalable framework for retrieving a ranked list of moments from collection of videos in any length to match a text query, a task termed Ranked Video Moment Retrieval (RVMR). Our framework, called Segment-Proposal-Ranking (SPR), simplifies the search process into three independent stages: segment retrieval, proposal generation, and moment refinement with re-ranking. Specifically, videos are divided into equal-length segments with precomputed embeddings indexed offline, allowing efficient retrieval regardless of video length. For scalable online retrieval, both segments and queries are projected into a shared feature space to enable approximate nearest neighbor (ANN) search. Retrieved segments are then merged into coarse-grained moment proposals. Then a refinement and re-ranking module is designed to reorder and adjust timestamps of the coarse-grained proposals. Evaluations on the TVR-Ranking dataset demonstrate that our framework achieves state-of-the-art performance with significant reductions in computational cost and processing time. The flexible design also allows for independent improvements to each stage, making SPR highly adaptable for large-scale applications.
Figures
Reference graph
Works this paper leans on
-
[31]
Renjie Liang, Li Li, Chongzhi Zhang, Jing Wang, Xizhou Zhu, and Aixin Sun. 2024. TVR-Ranking: A Dataset for Ranked Video Moment Retrieval with Imprecise Queries. arXiv:2407.06597 [cs.AI] https://arxiv.org/abs/2407.06597
work page Pith review arXiv 2024
-
[59]
Hao Zhang, Aixin Sun, Wei Jing, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh. 2021. Video Corpus Moment Retrieval with Contrastive Learning. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual Event, Canada, July 11-15, 2021, Fernando Diaz, Chirag Shah, Torsten Suel...
arXiv 2021
-
[2]
Michael Bendersky, W. Bruce Croft, and Yanlei Diao. 2011. Quality-biased ranking of web documents. In Proceedings of the Forth International Conference on Web Search and Web Data Mining, WSDM 2011, Hong Kong, China, February 9-12, 2011, Irwin King, Wolfgang Nejdl, and Hang Li (Eds.). ACM, 95–104. https: //doi.org/10.1145/1935826.1935849
arXiv 2011
-
[3]
Rae, Erich Elsen, and Laurent Sifre
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Ruther- ford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, O...
work page 2022
-
[4]
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005. Learning to rank using gradient descent. InProceedings of the 22nd International Conference on Machine Learning (Bonn, Germany) (ICML ’05). Association for Computing Machinery, New York, NY, USA, 89–96. https: //doi.org/10.1145/1102351.1102363
arXiv 2005
-
[5]
Meng Cao, Long Chen, Mike Zheng Shou, Can Zhang, and Yuexian Zou. 2021. On Pursuit of Designing Multi-modal Transformer for Video Grounding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Francine Moens, Xuanjing Huang, Lucia ...
-
[6]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6299–6308
2017
-
[7]
Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W. Cohen. 2023. Re- Imagen: Retrieval-Augmented Text-to-Image Generator. InThe Eleventh Interna- tional Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=XSEBx0iSjFQ
work page 2023
Show all 64 references
-
[8]
Crandall, Mohit Bansal, and Gedas Bertasius
Feng Cheng, Xizi Wang, Jie Lei, David J. Crandall, Mohit Bansal, and Gedas Bertasius. 2023. VindLU: A Recipe for Effective Video-and-Language Pretraining. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEE...
2023
-
[9]
Ritendra Datta, Dhiraj Joshi, Jia Li, and James Ze Wang. 2008. Image retrieval: Ideas, influences, and trends of the new age. ACM Comput. Surv. 40, 2 (2008), 5:1–5:60. https://doi.org/10.1145/1348246.1348248
2008
-
[10]
Jacob Devlin. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
- [11]
-
[12]
Victor Escorcia, Mattia Soldan, Josef Sivic, Bernard Ghanem, and Bryan C. Russell
-
[13]
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017. TALL: Temporal Activity Localization via Language Query. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . IEEE Computer Society, 5277–5285. https://doi.org/10.1109/I...
2017 doi
-
[14]
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech...
2021
-
[15]
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Proceedings of the 2021 Conference on Em- pirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, ...
2021 doi
-
[16]
Deepak Gupta, Kush Attal, and Dina Demner-Fushman. 2024. Towards Answering Health-related Questions from Medical Videos: Datasets and Approaches. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COL...
2024
-
[17]
Tanveer Hannan, Md Mohaiminul Islam, Thomas Seidl, and Gedas Bertasius
-
[18]
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan C. Russell. 2017. Localizing Moments in Video with Natural Language. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017. IEEE Computer Societ...
2017
-
[19]
Zhijian Hou, Chong-Wah Ngo, and Wing Kwong Chan. 2021. CONQUER: Contextual Query-aware Ranking for Video Corpus Moment Retrieval. In MM ’21: ACM Multimedia Conference, Virtual Event, China, October 20 - 24, 2021 , Heng Tao Shen, Yueting Zhuang, John R. Smith, Yang Yang, Pablo ...
2021
-
[20]
Zhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao, Kun Yan, Wing Kwong Chan, Chong-Wah Ngo, Mike Zheng Shou, and Nan Duan. 2023. CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding. In Proceedings of the 61st Annual Meeting of the Association for...
2023
-
[21]
Piotr Indyk and Rajeev Motwani. 1998. Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality. In Proceedings of the Thirtieth Annual ACM Symposium on the Theory of Computing, Dallas, Texas, USA, May 23-26, 1998 , Jeffrey Scott Vitter (Ed.). ACM, 604–613. h...
1998
-
[22]
Hervé Jégou, Matthijs Douze, and Cordelia Schmid. 2011. Product Quantization for Nearest Neighbor Search. IEEE Trans. Pattern Anal. Mach. Intell. 33, 1 (2011), 117–128. https://doi.org/10.1109/TPAMI.2010.57
2011 doi
-
[23]
Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Super- vision. In Proceedings of the 38th International Conference on...
2021
-
[24]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ...
2020
-
[25]
Omar Khattab and Matei Zaharia. 2020. ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, C...
2020
-
[26]
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles
-
[27]
Lafferty and ChengXiang Zhai
John D. Lafferty and ChengXiang Zhai. 2001. Document Language Models, Query Models, and Risk Minimization for Information Retrieval. In SIGIR 2001: Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, September...
2001
-
[28]
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal. 2020. Tvr: A large-scale dataset for video-subtitle moment retrieval. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXI 16 . Springer, 447–463
2020
-
[29]
Patrick S. H. Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advanc...
2020
-
[30]
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu
-
[32]
Yinhan Liu. 2019. Roberta: A robustly optimized bert pretraining approach.arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[33]
Zhenghao Liu, Chenyan Xiong, Yuanhuiyi Lv, Zhiyuan Liu, and Ge Yu. 2023. Universal Vision-Language Dense Retrieval: Learning A Unified Representation Space for Multi-Modal Retrieval. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwand...
2023
-
[34]
Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Or- leans, LA, USA, May 6-9, 2019 . OpenReview.net. https://openreview.net/forum? id=Bkg6RiCqY7
2019
-
[35]
Siyu Lou, Xuenan Xu, Mengyue Wu, and Kai Yu. 2022. Audio-Text Retrieval in Context. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022 . IEEE, 4793–4797. https://doi.org/10.1109/ICASSP43922.2022.9746786
2022
-
[36]
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning. Neurocomputing 508 (2022), 293–304. https://doi.org/10.1016/J. NEUCOM.2022.07.028
2022 doi
-
[37]
Craig Macdonald, Rodrygo L. T. Santos, and Iadh Ounis. 2013. The whens and hows of learning to rank for web search. Inf. Retr. 16, 5 (2013), 584–628. https://doi.org/10.1007/S10791-012-9209-9
2013 doi
-
[38]
Malkov and Dmitry A
Yury A. Malkov and Dmitry A. Yashunin. 2020. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs.IEEE Trans. Pattern Anal. Mach. Intell. 42, 4 (2020), 824–836. https://doi.org/10.1109/ TPAMI.2018.2889473
2020
-
[39]
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. 2020. End-to-End Learning of Visual Representations From Uncurated Instructional Videos. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seatt...
2020
-
[40]
Roy- Chowdhury
Niluthpol Chowdhury Mithun, Juncheng Li, Florian Metze, and Amit K. Roy- Chowdhury. 2018. Learning Joint Embedding with Multimodal Cues for Cross- Modal Video-Text Retrieval. In Proceedings of the 2018 ACM on International Conference on Multimedia Retrieval, ICMR 2018, Yokoham...
2018
-
[41]
Behnam Neyshabur and Nathan Srebro. 2015. On Symmetric and Asymmetric LSHs for Inner Product Search. InProceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015 (JMLR Workshop and Conference Proceedings, Vol. 37) , Francis R...
2015
-
[43]
Xiaoye Qu, Pengwei Tang, Zhikang Zou, Yu Cheng, Jianfeng Dong, Pan Zhou, and Zichuan Xu. 2020. Fine-grained Iterative Attention Network for Tem- poral Language Localization in Videos. In MM ’20: The 28th ACM Interna- tional Conference on Multimedia, Virtual Event / Seattle, W ...
2020
-
[44]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learni...
2021
-
[45]
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013. Grounding Action Descriptions in Videos. Trans. Assoc. Comput. Linguistics 1 (2013), 25–36. https://doi.org/10.1162/TACL_ A_00207
2013 doi
-
[46]
Robertson and Steve Walker
Stephen E. Robertson and Steve Walker. 1997. On Relevance Weights with Little Relevance Information. In SIGIR ’97: Proceedings of the 20th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 27-31, 1997, Philadelphia, PA, USA , ...
1997
-
[47]
Robertson and Hugo Zaragoza
Stephen E. Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Found. Trends Inf. Retr. 3, 4 (2009), 333–389. https://doi.org/10.1561/1500000019
2009 doi
-
[48]
Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14...
2016 doi
-
[49]
Mattia Soldan, Alejandro Pardo, Juan León Alcázar, Fabian Caba Heilbron, Chen Zhao, Silvio Giancola, and Bernard Ghanem. 2022. MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions. In IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2022
- [50]
-
[51]
Hao Wang, Zheng-Jun Zha, Liang Li, Dong Liu, and Jiebo Luo. 2021. Structured Multi-Level Interaction Network for Video Moment Localization via Language Query. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision ...
2021
-
[52]
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. In IEEE International Conference on Acoustics, Speech and Signal Processi...
2023
-
[54]
Bennett, Junaid Ahmed, and Arnold Overwijk
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In 9th International Conference on Learning Representations, ICLR 2021, V...
2021
-
[55]
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. In Proceedings of the 2021 Conference on Empirical Methods in Nat...
2021
-
[56]
Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko
Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko. 2019. Multilevel Language and Vision Integration for Text-to-Clip Re- trieval. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Application...
2019
-
[57]
Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, and Zhen- zhen Jiao. 2024. Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W...
2024
-
[58]
Jingtao Zhan, Jiaxin Mao, Yiqun Liu, Jiafeng Guo, Min Zhang, and Shaoping Ma
-
[60]
Hao Zhang, Aixin Sun, Wei Jing, Liangli Zhen, Joey Tianyi Zhou, and Rick Siow Mong Goh. 2021. Parallel Attention Network with Sequence Matching for Video Grounding. In Findings of the Association for Computational Lin- guistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 ...
2021
-
[61]
Hao Zhang, Aixin Sun, Wei Jing, and Joey Tianyi Zhou. 2020. Span-based Lo- calizing Network for Natural Language Video Localization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020 , Dan Jurafsky, Jo...
2020 doi
-
[62]
Mingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai, Fangzhou Hong, Huirong Li, Lei Yang, and Ziwei Liu. 2023. ReMoDiffuse: Retrieval-Augmented Motion Diffusion Model. InIEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023 . IEEE, ...
2023
-
[63]
Top-20" and “Top-40
Songyang Zhang, Houwen Peng, Jianlong Fu, and Jiebo Luo. 2020. Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artific...
2020 doi
-
[2017]
In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017
Dense-Captioning Events in Videos. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017 . IEEE Computer Society, 706–715. https://doi.org/10.1109/ICCV.2017.83
2017 doi
-
[2019]
CoRR abs/1907.12763 (2019)
Temporal Localization of Moments in Video Collections with Natural Language. CoRR abs/1907.12763 (2019). arXiv:1907.12763 http://arxiv.org/abs/ 1907.12763
2019 arXiv
-
[2020]
HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 , Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Asso...
2020 doi
-
[2025]
In Computer Vision – ECCV 2024 , Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.)
RGNet: A Unified Clip Retrieval and Grounding Network for Long Videos. In Computer Vision – ECCV 2024 , Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.). Springer Nature Switzerland, Cham, 352–369
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.