A reasoning-based agent (multimodal LLM, then Grounding-DINO, then SAM2) achieves state-of-the-art Referring Audio-Visual Segmentation without pixel-level supervision.
The sounding object near the woman.\
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.MM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation
A reasoning-based agent (multimodal LLM, then Grounding-DINO, then SAM2) achieves state-of-the-art Referring Audio-Visual Segmentation without pixel-level supervision.