Pith. sign in

REVIEW 2 cited by

AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.01836 v1 pith:YNKD7XE5 submitted 2023-05-03 cs.CV cs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.LGcs.MMcs.SDeess.AS
keywords segmentationaudio-visualav-samlocalizationanythingfeaturesmodelsegment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localization and segmentation. In this work, we propose a simple yet effective audio-visual localization and segmentation framework based on the Segment Anything Model, namely AV-SAM, that can generate sounding object masks corresponding to the audio. Specifically, our AV-SAM simply leverages pixel-wise audio-visual fusion across audio features and visual features from the pre-trained image encoder in SAM to aggregate cross-modal representations. Then, the aggregated cross-modal features are fed into the prompt encoder and mask decoder to generate the final audio-visual segmentation masks. We conduct extensive experiments on Flickr-SoundNet and AVSBench datasets. The results demonstrate that the proposed AV-SAM can achieve competitive performance on sounding object localization and segmentation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unlocking Spatial Grounding in Large Audio-Visual Retrieval models

    cs.IR 2026-06 conditional novelty 6.0 of 10

    An audio-guided spatial pooling module inserted into the frozen PE-AV retrieval model yields sound-source localization from intermediate visual tokens, nearly doubling prior AVATAR performance.

  2. How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A user-changeable material mask lets an encoder-decoder generate a room's impulse response from a single audio-visual observation, trained and evaluated on the new Acoustic Wonderland Dataset.

Pith tools