A wav2vec 2.0 model trained only on native Finland Swedish speech, plus temperature scaling and top-k normalization, detects L2 mispronunciations without any L2 training data.
The predicted labels’ proba- bilities, P , are obtained by softmax functionσ: P = σ(z/T )
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
other 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
other 1polarities
unclear 1representative citing papers
citing papers explorer
-
Mispronunciation Detection Without L2 Pronunciation Dataset in Low-Resource Setting: A Case Study in Finland Swedish
A wav2vec 2.0 model trained only on native Finland Swedish speech, plus temperature scaling and top-k normalization, detects L2 mispronunciations without any L2 training data.