Pith. sign in

REVIEW 2 cited by

iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.07735 v1 pith:AP3YRGTK submitted 2020-11-16 cs.CV cs.AIcs.LGcs.MMeess.IV

classification cs.CVcs.AIcs.LGcs.MMeess.IV
keywords videoiperceiveeventvideoqavisualansweringcaptioningcommon-sense
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most prior art in visual understanding relies solely on analyzing the "what" (e.g., event recognition) and "where" (e.g., event localization), which in some cases, fails to describe correct contextual relationships between events or leads to incorrect underlying visual attention. Part of what defines us as human and fundamentally different from machines is our instinct to seek causality behind any association, say an event Y that happened as a direct result of event X. To this end, we propose iPerceive, a framework capable of understanding the "why" between events in a video by building a common-sense knowledge base using contextual cues to infer causal relationships between objects in the video. We demonstrate the effectiveness of our technique using the dense video captioning (DVC) and video question answering (VideoQA) tasks. Furthermore, while most prior work in DVC and VideoQA relies solely on visual information, other modalities such as audio and speech are vital for a human observer's perception of an environment. We formulate DVC and VideoQA tasks as machine translation problems that utilize multiple modalities. By evaluating the performance of iPerceive DVC and iPerceive VideoQA on the ActivityNet Captions and TVQA datasets respectively, we show that our approach furthers the state-of-the-art. Code and samples are available at: iperceive.amanchadha.com.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PR-DETR injects k-means-derived position anchors and an overlap-aware relation mask into a DETR decoder, improving dense video captioning on two benchmarks.

  2. FocusedAD: Character-centric Movie Audio Description

    cs.CV 2025-04 conditional novelty 5.0 of 10

    FocusedAD uses face tracking, soft prompts, and a video language model to generate character-focused movie audio descriptions, reporting state-of-the-art scores on MAD-eval-Named and the new Cinepile-AD benchmark.

Pith tools