Pith. sign in

REVIEW 3 cited by

Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02988 v1 pith:5ITFBNTN submitted 2025-04-03 cs.SD eess.AS

classification cs.SDeess.AS
keywords audio-visualdatalocalizationseldsynthetictooldetectionevent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stereo Sound Event Localization and Detection with Onscreen/offscreen Classification

    cs.SD 2025-07 conditional novelty 6.0 of 10

    It introduces a stereo-audio sound event localization benchmark with onscreen/offscreen classification, and finds the audiovisual baseline's onscreen judgments are near chance.

  2. Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Fusing frozen CLAP and OWL-ViT embeddings via a Cross-Modal Conformer, plus autocorrelation-based features, improves stereo SELD over DCASE 2025 baselines.

  3. Deep Learning for Personalized Binaural Audio Reproduction

    eess.AS 2025-08 accept novelty 4.0 of 10

    A structured survey of deep learning for personalized binaural audio, covering explicit HRTF prediction and end-to-end synthesis, datasets, metrics, and open challenges.

Pith tools