Pith. sign in

REVIEW 1 cited by

ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12813 v1 pith:N45XDU3H submitted 2024-10-01 cs.MM cs.CV

classification cs.MMcs.CV
keywords videochatvtggroundingtemporaldialoguelanguagecaptionsdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video Temporal Grounding (VTG) aims to ground specific segments within an untrimmed video corresponding to the given natural language query. Existing VTG methods largely depend on supervised learning and extensive annotated data, which is labor-intensive and prone to human biases. To address these challenges, we present ChatVTG, a novel approach that utilizes Video Dialogue Large Language Models (LLMs) for zero-shot video temporal grounding. Our ChatVTG leverages Video Dialogue LLMs to generate multi-granularity segment captions and matches these captions with the given query for coarse temporal grounding, circumventing the need for paired annotation data. Furthermore, to obtain more precise temporal grounding results, we employ moment refinement for fine-grained caption proposals. Extensive experiments on three mainstream VTG datasets, including Charades-STA, ActivityNet-Captions, and TACoS, demonstrate the effectiveness of ChatVTG. Our ChatVTG surpasses the performance of current zero-shot methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TimeRefine: Temporal Grounding with Time Refining Video LLM

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A training reformulation that turns timestamp prediction into iterative offset refinement, with an auxiliary L1 loss, improves Video LLM temporal grounding.

Pith tools