Pith. sign in

REVIEW 2 cited by

MomentDiff: Generative Video Moment Retrieval from Random to Real

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.02869 v2 pith:FWNCLUL5 submitted 2023-07-06 cs.CV

classification cs.CV
keywords randommomentdifftemporaldatasetsvideoanti-biaslocationreal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description. To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typical human retrieval process from random browsing to gradual localization. Specifically, we first diffuse the real span to random noise, and learn to denoise the random noise to the original span with the guidance of similarity between text and video. This allows the model to learn a mapping from arbitrary random locations to real moments, enabling the ability to locate segments from random initialization. Once trained, MomentDiff could sample random temporal segments as initial guesses and iteratively refine them to generate an accurate temporal boundary. Different from discriminative works (e.g., based on learnable proposals or queries), MomentDiff with random initialized spans could resist the temporal location biases from datasets. To evaluate the influence of the temporal location biases, we propose two anti-bias datasets with location distribution shifts, named Charades-STA-Len and Charades-STA-Mom. The experimental results demonstrate that our efficient framework consistently outperforms state-of-the-art methods on three public benchmarks, and exhibits better generalization and robustness on the proposed anti-bias datasets. The code, model, and anti-bias evaluation datasets are available at https://github.com/IMCCretrieval/MomentDiff.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models

    cs.MM 2025-01 conditional novelty 6.0 of 10

    A tuning-free pipeline using frozen LLaMA-3, MiniGPT-v2, and Video-ChatGPT reports state-of-the-art zero-shot video moment retrieval on three benchmarks.

  2. Multi-modal Fusion and Query Refinement Network for Video Moment Retrieval and Highlight Detection

    cs.CV 2025-01 conditional novelty 4.0 of 10

    MRNet fuses RGB, optical flow, and depth features with word-, phrase-, and sentence-level query features, and reports improved moment retrieval and highlight detection scores on QVHighlights and Charades-STA.

Pith tools