Pith. sign in

REVIEW 1 cited by

Learning to Locate Visual Answer in Video Corpus Using Question

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.05423 v4 pith:W7TUPEH5 submitted 2022-10-11 cs.CV cs.CL

classification cs.CVcs.CL
keywords answervisualvideocorpuslocalizationretrievaltaskvcval
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a new task, named video corpus visual answer localization (VCVAL), which aims to locate the visual answer in a large collection of untrimmed instructional videos using a natural language question. This task requires a range of skills - the interaction between vision and language, video retrieval, passage comprehension, and visual answer localization. In this paper, we propose a cross-modal contrastive global-span (CCGS) method for the VCVAL, jointly training the video corpus retrieval and visual answer localization subtasks with the global-span matrix. We have reconstructed a dataset named MedVidCQA, on which the VCVAL task is benchmarked. Experimental results show that the proposed method outperforms other competitive methods both in the video corpus retrieval and visual answer localization subtasks. Most importantly, we perform detailed analyses on extensive experiments, paving a new path for understanding the instructional videos, which ushers in further research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M$^3$-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding

    cs.CV 2025-07 conditional novelty 7.0 of 10

    M3-Med is a new bilingual benchmark that tests video-language AI models on multi-hop temporal reasoning in medical instructional videos, and it shows a large gap between current models and human performance.

Pith tools