VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.
Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
abstract
We used a 3D simulator to create artificial video data with standardized annotations, aiming to aid in the development of Embodied AI. Our question answering (QA) dataset measures the extent to which a robot can understand human behavior and the environment in a home setting. Preliminary experiments suggest our dataset is useful in measuring AI's comprehension of daily life. \end{abstract}
citation-role summary
dataset 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
VUDG: A Dataset for Video Understanding Domain Generalization
VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.