Pith. sign in

REVIEW 2 cited by

Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18832 v1 pith:UDOCJB4E submitted 2024-12-25 eess.AS cs.SD

classification eess.AScs.SD
keywords speechadaptationadaptersdysarthricsfmsspeakerspeakersdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data-intensive fine-tuning of speech foundation models (SFMs) to scarce and diverse dysarthric and elderly speech leads to data bias and poor generalization to unseen speakers. This paper proposes novel structured speaker-deficiency adaptation approaches for SSL pre-trained SFMs on such data. Speaker and speech deficiency invariant SFMs were constructed in their supervised adaptive fine-tuning stage to reduce undue bias to training data speakers, and serves as a more neutral and robust starting point for test time unsupervised adaptation. Speech variability attributed to speaker identity and speech impairment severity, or aging induced neurocognitive decline, are modelled using separate adapters that can be combined together to model any seen or unseen speaker. Experiments on the UASpeech dysarthric and DementiaBank Pitt elderly speech corpora suggest structured speaker-deficiency adaptation of HuBERT and Wav2vec2-conformer models consistently outperforms baseline SFMs using either: a) no adapters; b) global adapters shared among all speakers; or c) single attribute adapters modelling speaker or deficiency labels alone by statistically significant WER reductions up to 3.01% and 1.50% absolute (10.86% and 6.94% relative) on the two tasks respectively. The lowest published WER of 19.45% (49.34% on very low intelligibility, 33.17% on unseen words) is obtained on the UASpeech test set of 16 dysarthric speakers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition

    eess.AS 2025-05 conditional novelty 5.0 of 10

    MOPSA uses K-means clustered speaker prompts with a trained router to provide zero-shot, real-time Whisper adaptation for elderly speech, achieving WER/CER reductions on DementiaBank Pitt and JCCOCC MoCA.

  2. On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An on-the-fly router that predicts speaker-specific adapter weights lets a speech foundation model adapt to dysarthric speakers with zero-shot, real-time processing, achieving the lowest reported word error rate on UASpeech.

Pith tools