A cross-attention fusion of two discrete speech-unit streams, including cheap self-augmented variants, cuts character error rates over single-stream baselines in English and multilingual ASR.
HuBERT: Self-supervised speech rep- resentation learning by masked prediction of hidden units,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
A cross-attention fusion of two discrete speech-unit streams, including cheap self-augmented variants, cuts character error rates over single-stream baselines in English and multilingual ASR.