MUPA combines three ordering-based reasoning paths with a reflection agent that verifies and fuses answer-evidence pairs, reaching 30.3% and 47.4% grounded QA accuracy on NExT-GQA and DeVE-QA with a 7B model.
Can I trust your answer? Visually grounded video question answering,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
MUPA combines three ordering-based reasoning paths with a reflection agent that verifies and fuses answer-evidence pairs, reaching 30.3% and 47.4% grounded QA accuracy on NExT-GQA and DeVE-QA with a 7B model.