REVIEW 7 cited by
Audio-Visual Segmentation with Semantics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark, i.e., AVSBench, providing pixel-wise annotations for sounding objects in audible videos. It contains three subsets: AVSBench-object (Single-source subset, Multi-sources subset) and AVSBench-semantic (Semantic-labels subset). Accordingly, three settings are studied: 1) semi-supervised audio-visual segmentation with a single sound source; 2) fully-supervised audio-visual segmentation with multiple sound sources, and 3) fully-supervised audio-visual semantic segmentation. The first two settings need to generate binary masks of sounding objects indicating pixels corresponding to the audio, while the third setting further requires generating semantic maps indicating the object category. To deal with these problems, we propose a new baseline method that uses a temporal pixel-wise audio-visual interaction module to inject audio semantics as guidance for the visual segmentation process. We also design a regularization loss to encourage audio-visual mapping during training. Quantitative and qualitative experiments on AVSBench compare our approach to several existing methods for related tasks, demonstrating that the proposed method is promising for building a bridge between the audio and pixel-wise visual semantics. Code is available at https://github.com/OpenNLPLab/AVSBench. Online benchmark is available at http://www.avlbench.opennlplab.cn.
Forward citations
Cited by 7 Pith papers
-
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Selective convergence in contrastive audio-visual learning is exploited in a two-stage framework to localize dual sound sources without labels, with a new segmentation-mask evaluation benchmark.
-
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
A new video-level benchmark and a temporally aware model show that tracking sound sources over time is necessary for robust audio-visual localization.
-
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
A SAM2-based framework that uses a fused text-audio-visual token to prompt video segmentation achieves 58.5 J&F on Ref-AVS, outperforming the previous state of the art by 8.5 points.
-
AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
AVS-Mamba applies Mamba with temporal and cross-modal scanning to audio-visual segmentation, reporting top scores on AVSBench-object but not on AVSBench-semantic with the stronger backbone.
-
Collaborative Hybrid Propagator for Temporal Misalignment in Audio-Visual Segmentation
Co-Prop uses LLM-generated audio control points to split videos into consistent sound segments and propagates keyframe masks frame-by-frame with audio inserted, improving audio-visual segmentation scores.
-
Towards Open-Vocabulary Audio-Visual Event Localization
An ImageBind-based fine-tuned model outperforms a training-free zero-shot baseline on the new OV-AVEBench, which spans 67 event classes with 21 unseen at test time.
-
Video-Guided Foley Sound Generation with Multimodal Controls
A video-guided diffusion model generates synchronized foley sound from text, audio, and video controls, using joint training on noisy internet videos and professional sound-effect libraries to reach 48kHz output.
Discussion (0). Continue with ORCID to comment.