REVIEW 3 cited by
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. Cross-modal interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or off-screen sounds. In this paper, we first comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. Then, we identify the limitations of previous studies and make several contributions to overcome the limitations. First, we introduce a new synthetic benchmark for interactive sound source localization. Second, we introduce new evaluation metrics to rigorously assess sound source localization methods, focusing on accurately evaluating both localization performance and cross-modal interaction ability. Third, we propose a learning framework with a cross-modal alignment strategy to enhance cross-modal interaction. Lastly, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks together to thoroughly assess cross-modal interaction capabilities and benchmark competing methods. Our new benchmarks and evaluation metrics reveal previously overlooked issues in sound source localization studies. Our proposed novel method, with enhanced cross-modal alignment, shows superior sound source localization performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using new and standard evaluation metrics.
Forward citations
Cited by 3 Pith papers
-
Learning from Silence and Noise for Visual Sound Source Localization
Adding silence and Gaussian noise as negative training pairs improves self-supervised visual sound source localization, and the authors provide IS3+ and a separability metric.
-
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
A new video-level benchmark and a temporally aware model show that tracking sound sources over time is necessary for robust audio-visual localization.
-
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual ...
Discussion (0). Sign in to comment.