Pith. sign in

REVIEW 1 cited by

SpeechBlender: Speech Augmentation Framework for Mispronunciation Data Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.00923 v3 pith:US6ZZVUX submitted 2022-11-02 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords datamispronunciationspeechspeechblenderaugmentationcompareddetectiongenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronunciation errors to overcome such data scarcity. The SpeechBlender utilizes varieties of masks to target different regions of phonetic units, and use the mixing factors to linearly interpolate raw speech signals while augmenting pronunciation. The masks facilitate smooth blending of the signals, generating more effective samples than the `Cut/Paste' method. Our proposed technique achieves state-of-the-art results, with Speechocean762, on ASR dependent mispronunciation detection models at phoneme level, with a 2.0% gain in Pearson Correlation Coefficient (PCC) compared to the previous state-of-the-art [1]. Additionally, we demonstrate a 5.0% improvement at the phoneme level compared to our baseline. We also observed a 4.6% increase in F1-score with Arabic AraVoiceL2 testset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A new public benchmark for Arabic mispronunciation detection using Quranic recitation, with baseline models reaching F1 scores below 30%.

Pith tools