A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual segmentation on three datasets.
Models Below, we elaborate on the final models built for each approach
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
A training-free pipeline that converts audio into a text query via classification, captioning, or textual inversion and feeds it to a referring image segmentation model achieves state-of-the-art zero-shot audiovisual segmentation on three datasets.