A masked-autoencoder SSL model with sparse cross-attention and pretrained audio embeddings reports the best localization error on LuViRA music3 and speech3 while also estimating faulty microphone positions.
FN-SSL: Full-Band and Narrow-Band Fusion for Sound Source Localization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Extracting direct-path spatial features is critical for sound source localization in adverse acoustic environments. This paper proposes a full-band and narrow-band fusion network for estimating direct-path inter-channel phase difference (DP-IPD) from microphone signals. The alternating full-band and narrow-band layers are responsible for learning the full-band correlation and narrow-band extraction of DP-IPD, respectively. Experiments show that the proposed network noticeably outperforms other advanced methods on both simulated and real-world data.
citation-role summary
citation-polarity summary
fields
cs.SD 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
A masked-autoencoder SSL model with sparse cross-attention and pretrained audio embeddings reports the best localization error on LuViRA music3 and speech3 while also estimating faulty microphone positions.