Pith. sign in

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segment objects within the visual domain based on expressions containing multimodal cues. Such expressions are articulated in natural language forms but are enriched with multimodal cues, including audio and visual descriptions. To facilitate this research, we construct the first Ref-AVS benchmark, which provides pixel-level annotations for objects described in corresponding multimodal-cue expressions. To tackle the Ref-AVS task, we propose a new method that adequately utilizes multimodal cues to offer precise segmentation guidance. Finally, we conduct quantitative and qualitative experiments on three test subsets to compare our approach with existing methods from related tasks. The results demonstrate the effectiveness of our method, highlighting its capability to precisely segment objects using multimodal-cue expressions. Dataset is available at \href{https://gewu-lab.github.io/Ref-AVS}{https://gewu-lab.github.io/Ref-AVS}.

citation-role summary

background 1

citation-polarity summary

fields

cs.SD 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? cs.SD · 2025-02-01 · conditional · none · ref 49 · internal anchor

    Audio-visual segmentation models are shown to rely on visual salience rather than audio; a new robustness benchmark and a balanced-training method largely correct this behavior under negative audio conditions.