Pith. sign in

REVIEW 3 cited by

Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13676 v1 pith:UJPATLNC submitted 2024-07-18 cs.MM cs.CVcs.SDeess.AS

classification cs.MMcs.CVcs.SDeess.AS
keywords localizationcross-modalsoundsourceinteractionbenchmarksevaluationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. Cross-modal interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or off-screen sounds. In this paper, we first comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. Then, we identify the limitations of previous studies and make several contributions to overcome the limitations. First, we introduce a new synthetic benchmark for interactive sound source localization. Second, we introduce new evaluation metrics to rigorously assess sound source localization methods, focusing on accurately evaluating both localization performance and cross-modal interaction ability. Third, we propose a learning framework with a cross-modal alignment strategy to enhance cross-modal interaction. Lastly, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks together to thoroughly assess cross-modal interaction capabilities and benchmark competing methods. Our new benchmarks and evaluation metrics reveal previously overlooked issues in sound source localization studies. Our proposed novel method, with enhanced cross-modal alignment, shows superior sound source localization performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using new and standard evaluation metrics.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning from Silence and Noise for Visual Sound Source Localization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Adding silence and Gaussian noise as negative training pairs improves self-supervised visual sound source localization, and the authors provide IS3+ and a separability metric.

  2. What's Making That Sound Right Now? Video-centric Audio-Visual Localization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new video-level benchmark and a temporally aware model show that tracking sound sources over time is necessary for robust audio-visual localization.

  3. Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual ...

Pith tools