Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.
wav2vec: Unsu- pervised pre-training for speech recognition,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
Zipformer-based self-supervised pretraining is up to 3.5x faster than HuBERT and yields lower ASR word error rates on LibriSpeech.