VideoNavQA pairs 101,000 template-generated questions with near-optimal navigation videos in House3D, and adapted VQA models beat language-only baselines by about 8 accuracy points.
Embodied Question Answering
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2019 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
VideoNavQA pairs 101,000 template-generated questions with near-optimal navigation videos in House3D, and adapted VQA models beat language-only baselines by about 8 accuracy points.