Pith. sign in

Pano-avqa: Grounded audio-visual question answering on 360 ◦ videos

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it
abstract

360$^\circ$ videos convey holistic views for the surroundings of a scene. It provides audio-visual cues beyond pre-determined normal field of views and displays distinctive spatial relations on a sphere. However, previous benchmark tasks for panoramic videos are still limited to evaluate the semantic understanding of audio-visual relationships or spherical spatial property in surroundings. We propose a novel benchmark named Pano-AVQA as a large-scale grounded audio-visual question answering dataset on panoramic videos. Using 5.4K 360$^\circ$ video clips harvested online, we collect two types of novel question-answer pairs with bounding-box grounding: spherical spatial relation QAs and audio-visual relation QAs. We train several transformer-based models from Pano-AVQA, where the results suggest that our proposed spherical spatial embeddings and multimodal training objectives fairly contribute to a better semantic understanding of the panoramic surroundings on the dataset.

fields

cs.CV 1 cs.RO 1

years

2026 2

representative citing papers

EAGOR: Embodied Reasoning in Omni-direction

cs.RO · 2026-07-07 · conditional · novelty 7.0

EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

citing papers explorer

Showing 2 of 2 citing papers.

  • EAGOR: Embodied Reasoning in Omni-direction cs.RO · 2026-07-07 · conditional · none · ref 15 · internal anchor

    EAGOR reformulates embodied 360-degree directional reasoning as recursive Bayesian estimation on a spherical manifold using spherical harmonics, achieving training-free, rotation-equivariant target tracking.

  • PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World cs.CV · 2026-05-13 · unverdicted · none · ref 54 · 2 links

    PanoWorld adds spherical spatial cross-attention and pano-native training data to MLLMs for improved spatial reasoning on ERP panoramas, outperforming baselines on new and existing benchmarks.