Pith. sign in

REVIEW 1 cited by

Zero-Shot Automatic Pronunciation Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.19563 v1 pith:YAYVVMU4 submitted 2023-05-31 cs.SD cs.CLcs.LGeess.AS

classification cs.SDcs.CLcs.LGeess.AS
keywords automaticmethodassessmentbaselinesdatamaskingmodelsmodule
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic Pronunciation Assessment (APA) is vital for computer-assisted language learning. Prior methods rely on annotated speech-text data to train Automatic Speech Recognition (ASR) models or speech-score data to train regression models. In this work, we propose a novel zero-shot APA method based on the pre-trained acoustic model, HuBERT. Our method involves encoding speech input and corrupting them via a masking module. We then employ the Transformer encoder and apply k-means clustering to obtain token sequences. Finally, a scoring module is designed to measure the number of wrongly recovered tokens. Experimental results on speechocean762 demonstrate that the proposed method achieves comparable performance to supervised regression baselines and outperforms non-regression baselines in terms of Pearson Correlation Coefficient (PCC). Additionally, we analyze how masking strategies affect the performance of APA.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    A framework using native-only trained discrete token surprisal and DTW alignment features improves pronunciation assessment PCC to 0.66 on SpeechOcean762, approaching supervised performance.

Pith tools