BEST-STD trains a bidirectional Mamba encoder with contrastive learning and vector quantization to make speaker-agnostic speech tokens for fast spoken term detection.
and Goto, M., 2009
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
BEST-STD trains a bidirectional Mamba encoder with contrastive learning and vector quantization to make speaker-agnostic speech tokens for fast spoken term detection.