A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual segmentation on three datasets.
Many approaches leverage cross-modal attention [1, 2, 3] combined with contrastive learning [4, 5] to align audio and visual features
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual segmentation on three datasets.