Proposes a classifier framework that audits TTS output for phonological faithfulness using human speech benchmarks, revealing realization biases in Meta's MMS TTS for Assamese ATR harmony.
Towards a Phonology-Informed Evaluation of Multilingual TTS
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We propose a classifier-based framework that audits TTS output against language-specific phonological patterns using human speech as a benchmark. Testing Assamese advanced tongue root (ATR) vowel harmony with Meta's MMS TTS, we show that a classifier trained on human speech transfers to synthesized speech with minimal loss. The faithfulness audit reveals that [+ATR] mid vowels are realized as [-ATR] in 1/3 tokens despite an underlying [+ATR] specification, a bias absent in human speech. At the word level, predicted ATR labels classify harmony more accurately than transcription labels, indicating a gap between intended and produced phonology. The framework offers task-specific diagnostics and generalizes to other phonological contrasts with measurable acoustic cues.
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Towards a Phonology-Informed Evaluation of Multilingual TTS
Proposes a classifier framework that audits TTS output for phonological faithfulness using human speech benchmarks, revealing realization biases in Meta's MMS TTS for Assamese ATR harmony.