Pith. sign in

REVIEW 1 cited by

Towards Fine-Grained Video Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06820 v1 pith:6KGAXQFF submitted 2025-03-10 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords temporalvideoansweringfine-grainedmodelmoma-qaquestionunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial granularity, which consequently limits the capabilities of existing VideoQA methods. This paper introduces the Multi-Object Multi-Actor Question Answering (MOMA-QA) dataset, which is designed to address these shortcomings by emphasizing temporal localization, spatial relationship reasoning, and entity-centric queries. With ground truth scene graphs and temporal interval annotations, MOMA-QA is ideal for developing models for fine-grained video understanding. Furthermore, we present a novel video-language model, SGVLM, which incorporates a scene graph predictor, an efficient frame retriever, and a pre-trained large language model for temporal localization and fine-grained relationship understanding. Evaluations on MOMA-QA and other public datasets demonstrate the superior performance of our model, setting new benchmarks for VideoQA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

    cs.CL 2026-08 conditional novelty 5.0 of 10

    A new video benchmark, DrivelHub+, shows that current video-language models can describe what happens in social media clips but largely fail to infer the implicit humour, irony, or cultural meaning.

Pith tools