FusionVAD shows that simple addition or concatenation of MFCC and pre-trained model features outperforms cross-attention fusion for voice activity detection, with the best model beating Pyannote by 2.04 average DER.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Attention Is Not Always the Answer: Optimizing Voice Activity Detection with Simple Feature Fusion
FusionVAD shows that simple addition or concatenation of MFCC and pre-trained model features outperforms cross-attention fusion for voice activity detection, with the best model beating Pyannote by 2.04 average DER.