Pith. sign in

REVIEW 1 cited by

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.04379 v1 pith:H2ECDZMD submitted 2025-01-08 cs.SD eess.AS

classification cs.SDeess.AS
keywords discretefeaturesk-meansspeechtokensdysarthricphone-puritytoken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Discrete tokens extracted provide efficient and domain adaptable speech features. Their application to disordered speech that exhibits articulation imprecision and large mismatch against normal voice remains unexplored. To improve their phonetic discrimination that is weakened during unsupervised K-means or vector quantization of continuous features, this paper proposes novel phone-purity guided (PPG) discrete tokens for dysarthric speech recognition. Phonetic label supervision is used to regularize maximum likelihood and reconstruction error costs used in standard K-means and VAE-VQ based discrete token extraction. Experiments conducted on the UASpeech corpus suggest that the proposed PPG discrete token features extracted from HuBERT consistently outperform hybrid TDNN and End-to-End (E2E) Conformer systems using non-PPG based K-means or VAE-VQ tokens across varying codebook sizes by statistically significant word error rate (WER) reductions up to 0.99\% and 1.77\% absolute (3.21\% and 4.82\% relative) respectively on the UASpeech test set of 16 dysarthric speakers. The lowest WER of 23.25\% was obtained by combining systems using different token features. Consistent improvements on the phone purity metric were also achieved. T-SNE visualization further demonstrates sharper decision boundaries were produced between K-means/VAE-VQ clusters after introducing phone-purity guidance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition

    eess.AS 2025-06 conditional novelty 4.0 of 10

    Regularized federated learning (parameter, embedding, and KL-loss based) consistently outperforms FedAvg for dysarthric and elderly speech recognition by up to 0.55% absolute WER, and per-batch communication approache...

Pith tools