Pith. sign in

Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.

fields

eess.AS 1

years

2025 1

verdicts

ACCEPT 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Deep Learning for Personalized Binaural Audio Reproduction eess.AS · 2025-08-30 · accept · none · ref 234 · internal anchor

    A structured survey of deep learning for personalized binaural audio, covering explicit HRTF prediction and end-to-end synthesis, datasets, metrics, and open challenges.