A sequence-to-sequence voice conversion model trained on a native rater's shadowing utterances can spot unintelligible segments in L2 speech, beating an ASR baseline on the native rater but not on all listeners.
Dataset We utilized the shadowing dataset described in [8, 9]
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
A sequence-to-sequence voice conversion model trained on a native rater's shadowing utterances can spot unintelligible segments in L2 speech, beating an ASR baseline on the native rater but not on all listeners.