Pith. sign in

Multimodal Datasets and Benchmarks for Reasoning about Dynamic Spatio-Temporality in Everyday Environments

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We used a 3D simulator to create artificial video data with standardized annotations, aiming to aid in the development of Embodied AI. Our question answering (QA) dataset measures the extent to which a robot can understand human behavior and the environment in a home setting. Preliminary experiments suggest our dataset is useful in measuring AI's comprehension of daily life. \end{abstract}

citation-role summary

dataset 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

dataset 1

polarities

use dataset 1

representative citing papers

VUDG: A Dataset for Video Understanding Domain Generalization

cs.CV · 2025-05-30 · conditional · novelty 6.0

VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.

citing papers explorer

Showing 1 of 1 citing paper.

  • VUDG: A Dataset for Video Understanding Domain Generalization cs.CV · 2025-05-30 · conditional · none · ref 19 · internal anchor

    VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.