Pith. sign in

REVIEW 1 cited by

Context-aware Goodness of Pronunciation for Computer-Assisted Pronunciation Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.08647 v1 pith:IMXEQ4TU submitted 2020-08-19 eess.AS cs.SD

classification eess.AScs.SD
keywords scoringpronunciationfactormodeldetectiondurationgoodnessmispronunciation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Mispronunciation detection is an essential component of the Computer-Assisted Pronunciation Training (CAPT) systems. State-of-the-art mispronunciation detection models use Deep Neural Networks (DNN) for acoustic modeling, and a Goodness of Pronunciation (GOP) based algorithm for pronunciation scoring. However, GOP based scoring models have two major limitations: i.e., (i) They depend on forced alignment which splits the speech into phonetic segments and independently use them for scoring, which neglects the transitions between phonemes within the segment; (ii) They only focus on phonetic segments, which fails to consider the context effects across phonemes (such as liaison, omission, incomplete plosive sound, etc.). In this work, we propose the Context-aware Goodness of Pronunciation (CaGOP) scoring model. Particularly, two factors namely the transition factor and the duration factor are injected into CaGOP scoring. The transition factor identifies the transitions between phonemes and applies them to weight the frame-wise GOP. Moreover, a self-attention based phonetic duration modeling is proposed to introduce the duration factor into the scoring model. The proposed scoring model significantly outperforms baselines, achieving 20% and 12% relative improvement over the GOP model on the phoneme-level and sentence-level mispronunciation detection respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JCAPT: A Joint Modeling Approach for CAPT

    cs.CL 2025-06 conditional novelty 4.0 of 10

    JCAPT, a Mamba-based joint APA and MDD model with phonological features and think tokens, improves mispronunciation detection and several scoring aspects on speechocean762 over JAM.

Pith tools