REVIEW 4 cited by
Finding Moments in Video Collections Using Natural Language
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce the task of retrieving relevant video moments from a large corpus of untrimmed, unsegmented videos given a natural language query. Our task poses unique challenges as a system must efficiently identify both the relevant videos and localize the relevant moments in the videos. To address these challenges, we propose SpatioTemporal Alignment with Language (STAL), a model that represents a video moment as a set of regions within a series of short video clips and aligns a natural language query to the moment's regions. Our alignment cost compares variable-length language and video features using symmetric squared Chamfer distance, which allows for efficient indexing and retrieval of the video moments. Moreover, aligning language features to regions within a video moment allows for finer alignment compared to methods that extract only an aggregate feature from the entire video moment. We evaluate our approach on two recently proposed datasets for temporal localization of moments in video with natural language (DiDeMo and Charades-STA) extended to our video corpus moment retrieval setting. We show that our STAL re-ranking model outperforms the recently proposed Moment Context Network on all criteria across all datasets on our proposed task, obtaining relative gains of 37% - 118% for average recall and up to 30% for median rank. Moreover, our approach achieves more than 130x faster retrieval and 8x smaller index size with a 1M video corpus in an approximate setting.
Forward citations
Cited by 4 Pith papers
-
Moment of Untruth: Dealing with Negative Queries in Video Moment Retrieval
The paper introduces Negative-Aware Video Moment Retrieval, adding rejection of in-domain and out-of-domain irrelevant queries to moment retrieval, and shows a UniVTG adaptation rejects most negatives while retaining ...
-
A Flexible and Scalable Framework for Video Moment Search
SPR, a segment-proposal-ranking pipeline with offline indexing and approximate nearest neighbor retrieval, achieves state-of-the-art NDCG on TVR-Ranking while processing a query in under one second.
-
Text-Video Multi-Grained Integration for Video Moment Montage
A new task, dataset, and multi-grained text-video fusion model for automatically assembling video montages from a narration script.
-
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
DMR-JRG jointly learns video paragraph retrieval and weakly supervised sentence-level grounding through mutually reinforcing retrieval and grounding branches.
Discussion (0). Continue with ORCID to comment.