Pith. sign in

REVIEW 1 cited by

Multi-object event graph representation learning for Video Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.07747 v1 pith:DELEF2EV submitted 2024-09-12 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords eventgraphlearningmethodmultipleobjectsquestionrepresentation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships among objects extracted from videos to perform causal and temporal reasoning. While prior works have focused on modeling individual object movements using transformer-based methods, they falter when capturing complex scenarios involving multiple objects (e.g., "a boy is throwing a ball in a hoop"). We propose a contrastive language event graph representation learning method called CLanG to address this limitation. Aiming to capture event representations associated with multiple objects, our method employs a multi-layer GNN-cluster module for adversarial graph representation learning, enabling contrastive learning between the question text and its relevant multi-object event graph. Our method outperforms a strong baseline, achieving up to 2.2% higher accuracy on two challenging VideoQA datasets, NExT-QA and TGIF-QA-R. In particular, it is 2.8% better than baselines in handling causal and temporal questions, highlighting its strength in reasoning multiple object-based events.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RAVU builds a spatio-temporal graph memory of a video and uses compositional reasoning over the graph to retrieve a few key frames, achieving state-of-the-art zero-shot video QA on NExT-QA and EgoSchema with 5-10 frames.

Pith tools