REVIEW 1 major objections 20 references
CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation
T0 review · 1 major / 0 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read A new benchmark measures how accurately Chinese news TTS systems pronounce tricky written forms like scores and abbreviations from raw text alone.
desk verdict The paper releases a usable open benchmark for Chinese TTS pronunciation on raw news text with tricky written forms, but the ASR-derived ground truth has no shown human validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The target-level automatic scorer that matches TTS output against the ASR-ensemble transcripts for 992 specific written forms in news text.
What would settle it
A human listening test or independent transcription that systematically differs from the three-ASR ensemble transcripts on a large fraction of the targets would falsify the ground truth used by the scorer.
Extended reading notes
Core claim
CN-NewsTTS Bench v0.1 supplies a 200-record development set, an 800-record public test set, 992 public targets drawn from Chinese news, fixed transcripts from a three-ASR ensemble, an automatic target scorer, and initial results for seven product TTS systems that demonstrate varying pronunciation accuracy on raw text without user-side interventions.
Load-bearing premise
The fixed transcripts from the three-ASR ensemble accurately capture the intended pronunciations for the 992 targets.
Editorial extensions
If this is right
- TTS systems can now be compared automatically on pronunciation of written forms such as percentages, English abbreviations, and mixed names without manual edits or SSML.
- Performance gaps exist among current products, with top accuracy at 0.879 and several systems below 0.60 on the same targets.
- Category-level breakdowns and ASR-route diagnostics allow identification of specific weakness patterns in different systems.
- The public test set and scorer enable repeated evaluation as new TTS versions are released.
Reading between the lines
- Developers could use the benchmark to prioritize improvements in handling hyphenated model names and unit symbols without relying on LLM rewriting.
- The same target-level approach might transfer to news TTS evaluation in other languages that mix scripts or symbols.
- If ASR errors cluster on particular target categories, future versions could incorporate human-verified subsets for those categories.
- Widespread adoption might shift industry focus from general naturalness metrics toward precise pronunciation of frequent written forms.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims to introduce CN-NewsTTS Bench v0.1, an open target-level benchmark for evaluating Chinese news TTS products on pronouncing complex written forms such as scores, model names, English abbreviations from raw text without additional hints. It includes development and test sets, 992 targets with fixed three-ASR ensemble transcripts as ground truth, an automatic scorer, and initial results for seven systems with the best achieving 0.879 strict accuracy. The work also provides ASR diagnostics, ablations, category results, and confidence intervals.
Significance. Should the ASR transcripts prove to be a reliable proxy for intended pronunciations, the benchmark would provide a useful automatic evaluation tool for a common challenge in Chinese TTS for news content. The open release, ablations, and confidence intervals are positive features that support reproducibility and allow for nuanced analysis of system performance across categories.
major comments (1)
- [Abstract] The reported system accuracies (0.879 best, several below 0.60) are measured using transcripts from a three-ASR ensemble as ground truth for the 992 targets. The manuscript does not provide human validation, agreement rates, or error analysis for these transcripts on ambiguous targets, which is central to validating the benchmark's automatic scorer as a measure of correct pronunciation rather than agreement with ASR.
Simulated Author's Rebuttal
We thank the referee for highlighting the importance of validating the ASR-derived ground truth. We address the single major comment below.
read point-by-point responses
-
Referee: [Abstract] The reported system accuracies (0.879 best, several below 0.60) are measured using transcripts from a three-ASR ensemble as ground truth for the 992 targets. The manuscript does not provide human validation, agreement rates, or error analysis for these transcripts on ambiguous targets, which is central to validating the benchmark's automatic scorer as a measure of correct pronunciation rather than agreement with ASR.
Authors: We agree that the absence of human validation and agreement rates for the ensemble transcripts on ambiguous targets is a limitation. The three-ASR ensemble was chosen to reduce single-system errors via majority voting, and the manuscript already includes ASR-route diagnostics, subset ablations, and category results to characterize consistency. However, these do not substitute for direct human assessment. We will add (i) inter-ASR agreement rates across the 992 targets and (ii) a human validation study on a stratified sample of 150 targets (including ambiguous cases) with reported agreement to the ensemble, to be included in the revised manuscript. revision: yes
Circularity Check
No circularity: benchmark accuracies measured against external ASR transcripts
full rationale
The paper introduces CN-NewsTTS Bench with 992 targets and fixed transcripts from a three-ASR ensemble as ground truth, then reports TTS system accuracies (e.g., 0.879 strict accuracy) against those transcripts. This evaluation chain does not reduce any result to a quantity defined by the benchmark itself, nor does it involve self-definitional equations, fitted inputs renamed as predictions, or load-bearing self-citations. The derivation remains self-contained against the external ASR references, consistent with a standard benchmark release rather than a closed derivation loop.
Assumptions & free parameters
assumptions (1)
- domain assumption The three-ASR ensemble produces fixed transcripts that serve as accurate ground truth for target-level scoring.
Cite this review
Pith. "Pith review of CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation." pith.science (2026). https://pith.science/paper/MEMIZGUX
@misc{pith2026260624714,
author = {Pith},
title = {Pith review of: CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation},
year = {2026},
howpublished = {\url{https://pith.science/paper/MEMIZGUX}},
note = {Machine review of arXiv:2606.24714}
}
read the original abstract
Chinese news text contains dense written forms such as scores, hyphenated model names, ranges, unit symbols, percentages, English abbreviations, and mixed Chinese-Latin-digit names. These forms are frequent in real listening workflows, and a text-to-speech (TTS) system can preserve the written string while changing the spoken meaning. We introduce CN-NewsTTS Bench v0.1, an open target-level benchmark for evaluating whether Chinese news TTS products pronounce such targets correctly from raw text, without user-side rules, LLM rewriting, SSML hints, or manual edits. The release contains a 200-record development set, an 800-record public test set, 992 public auto-evaluable targets, fixed transcripts from a three-ASR ensemble, an automatic target scorer, and initial results for seven product TTS systems. We additionally report ASR-route diagnostics, ASR-subset ablations, category-level results, confidence intervals, and provider configuration metadata. The best system reaches 0.879 strict accuracy, while several systems remain below 0.60.
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Modern Chinese news TTS systems are used in settings where lis- teners expect symbols and compact written forms to be read accord- ing to news conventions. A short article may contain a sports score such as 96-91, an aircraft model such as ්-27, a year range such as 2028-2030୍, a torque unit such as 620N ·m, or a generation label such as 80ުT...
-
[2]
CN-NewsTTS Bench: a target-level automatic benchmark for raw-input Chinese news TTS pronunciation
BENCHMARK DESIGN 2.1. Raw Input Product Track The main track measures product-facing behavior under a fixed raw- input condition. Each system receives the same original Chinese news-style text. Provider-internal text normalization and frontend behavior are allowed, but external rule frontends, LLM rewrites, SSML pronunciation hints, and manual benchmark-t...
work page Pith review arXiv 2026
-
[3]
Three-ASR Protocol The v0.1 public leaderboard uses three heterogeneous ASR routes: MiMo API ASR, SenseVoiceSmall, and Paraformer-zh
AUTOMATIC EV ALUATION 3.1. Three-ASR Protocol The v0.1 public leaderboard uses three heterogeneous ASR routes: MiMo API ASR, SenseVoiceSmall, and Paraformer-zh. SenseVoic- eSmall and Paraformer-zh are open-source local recognizers from the FunAudioLLM/FunASR ecosystem [ 11, 12]; MiMo API ASR pro- vides an API-based route. The public release fixes all thre...
-
[4]
Each system generates one audio file per public test record, using one fixed voice and provider- level default normalization
INITIAL PUBLIC RESULTS We evaluate seven product TTS systems: Volcano/Doubao TTS, Azure Speech TTS, Google Cloud TTS, MiniMax TTS, Aliyun CosyVoice, MiMo TTS, and A WS Polly. Each system generates one audio file per public test record, using one fixed voice and provider- level default normalization. Generated audio is not redistributed in the GitHub repos...
-
[5]
The release includes data, schema, scoring code, fixed ASR transcripts, leaderboard files, the dashboard, audit, and checksums
REPRODUCIBILITY AND RELEASE Repository: https://github.com/Jayden-X-L/ cn-news-tts-bench Release: https://github.com/Jayden-X-L/ cn-news-tts-bench/releases/tag/v0.1 Commit: f94a679fc7fc. The release includes data, schema, scoring code, fixed ASR transcripts, leaderboard files, the dashboard, audit, and checksums. Table 8 lists compact model and voice labe...
2026
-
[6]
ASR text can hide pronunciation errors, especially same- character tones, and unknown rates vary by category
LIMITATIONS The benchmark is automatic and should not replace human listen- ing tests. ASR text can hide pronunciation errors, especially same- character tones, and unknown rates vary by category. The public split supports reproducibility but may later require a hidden test split. Commercial TTS audio is omitted because provider terms vary; the release fi...
-
[7]
CONCLUSION CN-NewsTTS Bench v0.1 provides an open, target-level benchmark for raw-input Chinese news TTS pronunciation accuracy. By fixing the data split, three-ASR transcripts, scorer, category diagnostics, and initial seven-system leaderboard, it makes a common production failure mode measurable and reproducible. The results indicate that compact Chines...
-
[8]
No human-subject audio, personal speech data, or pri- vate user data is released
COMPLIANCE WITH ETHICAL STANDARDS This work uses synthetic news-style text and released benchmark transcripts. No human-subject audio, personal speech data, or pri- vate user data is released
Show all 20 references
-
[9]
ITU-T Recom- mendation P .800: Methods for subjective determination of transmission quality,
International Telecommunication Union, “ITU-T Recom- mendation P .800: Methods for subjective determination of transmission quality,” https://www.itu.int/rec/T- REC-P.800, 1996
1996
-
[10]
MOSNet: Deep learning based objective assessment for voice conver- sion,
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Ju- nichi Y amagishi, Yu Tsao, and Hsin-Min Wang, “MOSNet: Deep learning based objective assessment for voice conver- sion,” in Proc. Interspeech, 2019, pp. 1541–1545
2019
-
[11]
Paul Taylor, Text-to-Speech Synthesis, Cambridge University Press, 2009
2009
-
[12]
Normalization of non-standard words,
Richard Sproat, Alan W. Black, Stanley Chen, Shankar Kumar, Mari Ostendorf, and Christopher Richards, “Normalization of non-standard words,” Computer Speech & Language , vol. 15, no. 3, pp. 287–333, 2001
2001
-
[13]
The Kestrel TTS text nor- malization system,
Peter Ebden and Richard Sproat, “The Kestrel TTS text nor- malization system,” Natural Language Engineering , vol. 21, no. 3, pp. 333–353, 2015
2015
-
[14]
RNN approaches to text normalization: A challenge,
Richard Sproat and Navdeep Jaitly, “RNN approaches to text normalization: A challenge,” arXiv preprint arXiv:1611.00068, 2016
2016 arXiv
-
[15]
A three-stage text normalization strategy for Mandarin text-to- speech systems,
Tao Zhou, Yuan Dong, Dezhi Huang, Wu Liu, and Haila Wang, “A three-stage text normalization strategy for Mandarin text-to- speech systems,” in Proc. International Symposium on Chinese Spoken Language Processing, 2008, pp. 125–128
2008
-
[16]
An end-to-end Chinese text normalization model based on rule-guided flat-lattice trans- former,
Wenlin Dai, Changhe Song, Xiang Li, Zhiyong Wu, Huashan Pan, Xiulin Li, and Helen Meng, “An end-to-end Chinese text normalization model based on rule-guided flat-lattice trans- former,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing, 2022
2022
-
[17]
g2pM: A neural grapheme- to-phoneme conversion package for Mandarin Chinese based on a new open benchmark dataset,
Kyubyong Park and Seanie Lee, “g2pM: A neural grapheme- to-phoneme conversion package for Mandarin Chinese based on a new open benchmark dataset,” in Proc. Interspeech, 2020
2020
-
[18]
CVTE-Poly: A new bench- mark for Chinese polyphone disambiguation,
Siheng Zhang, Xingjun Tan, Y anqiang Lei, Xianxiang Wang, Zhizhong Zhang, and Yuan Xie, “CVTE-Poly: A new bench- mark for Chinese polyphone disambiguation,” in Proc. Inter- speech, 2023, pp. 5526–5530
2023
-
[19]
SenseVoice: Multilingual speech understanding model,
FunAudioLLM Contributors, “SenseVoice: Multilingual speech understanding model,” https://github.com/ FunAudioLLM/SenseVoice, 2024
2024
-
[20]
FunASR: A fundamental end-to- end speech recognition toolkit,
FunASR Contributors, “FunASR: A fundamental end-to- end speech recognition toolkit,” https://github.com/ modelscope/FunASR, 2024
2024
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.