Pith. sign in

REVIEW 1 cited by

Localizing Moments in Long Video Via Multimodal Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.13372 v2 pith:VFAR4VIH submitted 2023-02-26 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords groundingmodelguidancelongvideowindowsbasecurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent introduction of the large-scale, long-form MAD and Ego4D datasets has enabled researchers to investigate the performance of current state-of-the-art methods for video grounding in the long-form setup, with interesting findings: current grounding methods alone fail at tackling this challenging task and setup due to their inability to process long video sequences. In this paper, we propose a method for improving the performance of natural language grounding in long videos by identifying and pruning out non-describable windows. We design a guided grounding framework consisting of a Guidance Model and a base grounding model. The Guidance Model emphasizes describable windows, while the base grounding model analyzes short temporal windows to determine which segments accurately match a given language query. We offer two designs for the Guidance Model: Query-Agnostic and Query-Dependent, which balance efficiency and accuracy. Experiments demonstrate that our proposed method outperforms state-of-the-art models by 4.1% in MAD and 4.52% in Ego4D (NLQ), respectively. Code, data and MAD's audio features necessary to reproduce our experiments are available at: https://github.com/waybarrios/guidance-based-video-grounding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DeCafNet reduces long-video temporal grounding cost by up to 47 percent while improving accuracy, using a lightweight sidekick encoder to select salient clips for a heavy expert encoder.

Pith tools