REVIEW 4 cited by
AV-SAM: Segment Anything Model Meets Audio-Visual Localization and Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localization and segmentation. In this work, we propose a simple yet effective audio-visual localization and segmentation framework based on the Segment Anything Model, namely AV-SAM, that can generate sounding object masks corresponding to the audio. Specifically, our AV-SAM simply leverages pixel-wise audio-visual fusion across audio features and visual features from the pre-trained image encoder in SAM to aggregate cross-modal representations. Then, the aggregated cross-modal features are fed into the prompt encoder and mask decoder to generate the final audio-visual segmentation masks. We conduct extensive experiments on Flickr-SoundNet and AVSBench datasets. The results demonstrate that the proposed AV-SAM can achieve competitive performance on sounding object localization and segmentation.
Forward citations
Cited by 4 Pith papers
-
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
An audio-guided spatial pooling module inserted into the frozen PE-AV retrieval model yields sound-source localization from intermediate visual tokens, nearly doubling prior AVATAR performance.
-
How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
A user-changeable material mask lets an encoder-decoder generate a room's impulse response from a single audio-visual observation, trained and evaluated on the new Acoustic Wonderland Dataset.
-
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
Audio-visual segmentation models are shown to rely on visual salience rather than audio; a new robustness benchmark and a balanced-training method largely correct this behavior under negative audio conditions.
-
INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation
INT improves task-generic promptable segmentation by progressively mining negative candidates, using VLM output differences after masking to select and refine instance-specific prompts.
Discussion (0). Continue with ORCID to comment.