A multi-band semantic-spatial alignment network (AV-SSAN) localizes a target sound source using a cross-instance visual prompt, achieving 16.59 degrees mean error and 71.29% accuracy on the new VGGSound-SSL benchmark.
InIEEE International Conference on Acous- tics, Speech and Signal Processing, 711–715
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
A multi-band semantic-spatial alignment network (AV-SSAN) localizes a target sound source using a cross-instance visual prompt, achieving 16.59 degrees mean error and 71.29% accuracy on the new VGGSound-SSL benchmark.