A multi-band semantic-spatial alignment network (AV-SSAN) localizes a target sound source using a cross-instance visual prompt, achieving 16.59 degrees mean error and 71.29% accuracy on the new VGGSound-SSL benchmark.
IEEE Journal of Selected Topics in Signal Processing, 13(1): 34–48
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
A multi-band semantic-spatial alignment network (AV-SSAN) localizes a target sound source using a cross-instance visual prompt, achieving 16.59 degrees mean error and 71.29% accuracy on the new VGGSound-SSL benchmark.