Pith. sign in

REVIEW 2 cited by

Multi-scale Multi-instance Visual Sound Localization and Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.00486 v1 pith:STC2THKZ submitted 2024-08-31 cs.CV cs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.LGcs.MMcs.SDeess.AS
keywords multi-scalevisualsoundlocalizationfeaturesimagecorrespondingm2vsl
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual sound localization is a typical and challenging problem that predicts the location of objects corresponding to the sound source in a video. Previous methods mainly used the audio-visual association between global audio and one-scale visual features to localize sounding objects in each image. Despite their promising performance, they omitted multi-scale visual features of the corresponding image, and they cannot learn discriminative regions compared to ground truths. To address this issue, we propose a novel multi-scale multi-instance visual sound localization framework, namely M2VSL, that can directly learn multi-scale semantic features associated with sound sources from the input image to localize sounding objects. Specifically, our M2VSL leverages learnable multi-scale visual features to align audio-visual representations at multi-level locations of the corresponding image. We also introduce a novel multi-scale multi-instance transformer to dynamically aggregate multi-scale cross-modal representations for visual sound localization. We conduct extensive experiments on VGGSound-Instruments, VGG-Sound Sources, and AVSBench benchmarks. The results demonstrate that the proposed M2VSL can achieve state-of-the-art performance on sounding object localization and segmentation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

    cs.MM 2026-08 conditional novelty 6.0 of 10

    Selective convergence in contrastive audio-visual learning is exploited in a two-stage framework to localize dual sound sources without labels, with a new segmentation-mask evaluation benchmark.

  2. Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?

    cs.SD 2025-02 conditional novelty 6.0 of 10

    Audio-visual segmentation models are shown to rely on visual salience rather than audio; a new robustness benchmark and a balanced-training method largely correct this behavior under negative audio conditions.

Pith tools