Pith. sign in

REVIEW 1 cited by

CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.10918 v2 pith:557OFCM4 submitted 2022-09-22 cs.CV cs.CLcs.IR

classification cs.CVcs.CLcs.IR
keywords longconealignmentvideoscoarse-to-fineframeworkinferencemechanism
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper tackles an emerging and challenging problem of long video temporal grounding~(VTG) that localizes video moments related to a natural language (NL) query. Compared with short videos, long videos are also highly demanded but less explored, which brings new challenges in higher inference computation cost and weaker multi-modal alignment. To address these challenges, we propose CONE, an efficient COarse-to-fiNE alignment framework. CONE is a plug-and-play framework on top of existing VTG models to handle long videos through a sliding window mechanism. Specifically, CONE (1) introduces a query-guided window selection strategy to speed up inference, and (2) proposes a coarse-to-fine mechanism via a novel incorporation of contrastive learning to enhance multi-modal alignment for long videos. Extensive experiments on two large-scale long VTG benchmarks consistently show both substantial performance gains (e.g., from 3.13% to 6.87% on MAD) and state-of-the-art results. Analyses also reveal higher efficiency as the query-guided window selection mechanism accelerates inference time by 2x on Ego4D-NLQ and 15x on MAD while keeping SOTA results. Codes have been released at https://github.com/houzhijian/CONE.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DeCafNet reduces long-video temporal grounding cost by up to 47 percent while improving accuracy, using a lightweight sidekick encoder to select salient clips for a heavy expert encoder.

Pith tools