A new Chinese TTS benchmark and protocol, Audio Turing Test, shows top LLM-based TTS models achieve only about 0.4 out of 1.0 on human-likeness, far below real human speech.
Why we should report the details in subjective evaluation of tts more rigorously
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
A new Chinese TTS benchmark and protocol, Audio Turing Test, shows top LLM-based TTS models achieve only about 0.4 out of 1.0 on human-likeness, far below real human speech.