Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.