Pith. sign in

REVIEW 2 cited by

Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06310 v1 pith:Q7ZQ4UID submitted 2024-07-08 cs.SD cs.AIcs.HCcs.LGeess.AS

classification cs.SDcs.AIcs.HCcs.LGeess.AS
keywords adaptationfeaturesspeakerspeaker-leveldysarthricelderlyspeechon-the-fly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The application of data-intensive automatic speech recognition (ASR) technologies to dysarthric and elderly adult speech is confronted by their mismatch against healthy and nonaged voices, data scarcity and large speaker-level variability. To this end, this paper proposes two novel data-efficient methods to learn homogeneous dysarthric and elderly speaker-level features for rapid, on-the-fly test-time adaptation of DNN/TDNN and Conformer ASR models. These include: 1) speaker-level variance-regularized spectral basis embedding (VR-SBE) features that exploit a special regularization term to enforce homogeneity of speaker features in adaptation; and 2) feature-based learning hidden unit contributions (f-LHUC) transforms that are conditioned on VR-SBE features. Experiments are conducted on four tasks across two languages: the English UASpeech and TORGO dysarthric speech datasets, the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech corpora. The proposed on-the-fly speaker adaptation techniques consistently outperform baseline iVector and xVector adaptation by statistically significant word or character error rate reductions up to 5.32% absolute (18.57% relative) and batch-mode LHUC speaker adaptation by 2.24% absolute (9.20% relative), while operating with real-time factors speeding up to 33.6 times against xVectors during adaptation. The efficacy of the proposed adaptation techniques is demonstrated in a comparison against current ASR technologies including SSL pre-trained systems on UASpeech, where our best system produces a state-of-the-art WER of 23.33%. Analyses show VR-SBE features and f-LHUC transforms are insensitive to speaker-level data quantity in testtime adaptation. T-SNE visualization reveals they have stronger speaker-level homogeneity than baseline iVectors, xVectors and batch-mode LHUC transforms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOPSA: Mixture of Prompt-Experts Based Speaker Adaptation for Elderly Speech Recognition

    eess.AS 2025-05 conditional novelty 5.0 of 10

    MOPSA uses K-means clustered speaker prompts with a trained router to provide zero-shot, real-time Whisper adaptation for elderly speech, achieving WER/CER reductions on DementiaBank Pitt and JCCOCC MoCA.

  2. On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition

    cs.SD 2025-05 conditional novelty 5.0 of 10

    An on-the-fly router that predicts speaker-specific adapter weights lets a speech foundation model adapt to dysarthric speakers with zero-shot, real-time processing, achieving the lowest reported word error rate on UASpeech.

Pith tools