Pith. sign in

State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Whispered speech recognition presents significant challenges for conventional automatic speech recognition systems, particularly when combined with dialect variation. However, utilizing an efficient method to solve this problem using a low-range dataset and processing load is beneficial. This paper proposes a solution using a Mamba-based state-space model and four fine-tuned self-supervised models consisting of Wav2Vec2, WavLM, HuBERT, and Whisper to address the dual challenges of whispered speech and dialect diversity. Based on our knowledge, this represents the best performance reported on the wTIMIT and CHAINS datasets for whispered speech recognition. We trained the models using whispered and normal speech data across Singaporean, US, and Irish dialects. The findings demonstrated that utilizing the proposed Mamba-based model could work as a highly efficient model trained with low amounts of whispered data to simultaneously work on whispered and normal speech recognition. The code for this work is freely available.

citation-role summary

background 1

citation-polarity summary

fields

eess.AS 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • State-Space Models in Efficient Whispered and Multi-dialect Speech Recognition eess.AS · 2025-06-20 · conditional · none · ref 1 · internal anchor

    A Mamba-based state-space ASR model and four fine-tuned self-supervised models achieve claimed state-of-the-art word error rates on whispered and normal speech across three English dialects, including near-perfect results on wTIMIT and CHAINS.