Contrastive training of a transformer encoder for acoustic word embeddings improves keyword spotting over contrastive RNNs and DTW in Luganda and Bambara radio broadcasts.
Pre-trained models For meanpooling and subsampling as described in Section 2.2, we consider four pre-trained self-supervised models from the wav2vec2.0 and HuBERT families [25, 26]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Low-resource keyword spotting using contrastively trained transformer acoustic word embeddings
Contrastive training of a transformer encoder for acoustic word embeddings improves keyword spotting over contrastive RNNs and DTW in Luganda and Bambara radio broadcasts.