A Mamba-based state-space ASR model and four fine-tuned self-supervised models achieve claimed state-of-the-art word error rates on whispered and normal speech across three English dialects, including near-perfect results on wTIMIT and CHAINS.
PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
PaddleSpeech is an open-source all-in-one speech toolkit. It aims at facilitating the development and research of speech processing technologies by providing an easy-to-use command-line interface and a simple code structure. This paper describes the design philosophy and core architecture of PaddleSpeech to support several essential speech-to-text and text-to-speech tasks. PaddleSpeech achieves competitive or state-of-the-art performance on various speech datasets and implements the most popular methods. It also provides recipes and pretrained models to quickly reproduce the experimental results in this paper. PaddleSpeech is publicly avaiable at https://github.com/PaddlePaddle/PaddleSpeech.
citation-role summary
citation-polarity summary
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition
A Mamba-based state-space ASR model and four fine-tuned self-supervised models achieve claimed state-of-the-art word error rates on whispered and normal speech across three English dialects, including near-perfect results on wTIMIT and CHAINS.