A SAM2-based framework that uses a fused text-audio-visual token to prompt video segmentation achieves 58.5 J&F on Ref-AVS, outperforming the previous state of the art by 8.5 points.
Cnn archi- tectures for large-scale audio classification
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
A SAM2-based framework that uses a fused text-audio-visual token to prompt video segmentation achieves 58.5 J&F on Ref-AVS, outperforming the previous state of the art by 8.5 points.