{"id":"c4f34602-c416-4052-ba0b-44fcd775d394","arxiv_id":"2606.24714","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces CN-NewsTTS Bench v0.1 with 1000 records and 992 targets for raw-input Chinese news TTS evaluation, reporting strict accuracies from 0.879 to below 0.60 across seven product systems.","lead":"The paper releases CN-NewsTTS Bench v0.1, a public dataset and automatic scorer for testing whether Chinese TTS systems correctly pronounce dense written forms like numbers, abbreviations, and mixed names in news text from raw input. A smart generalist might read it to see measurable gaps in current commercial voice systems for real-world Chinese media content.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"ASR ensemble transcripts as ground truth for 992 targets lack independent human validation on ambiguous cases","rationale":"The reader's weakest_assumption directly identifies the load-bearing condition for the central claim. Full-text details on diagnostics would be needed to weaken it, but the provided abstract leaves the assumption unverified, preserving the CONDITIONAL verdict.","tokens_in":1724,"tokens_out":282,"duration_ms":14600,"concrete_test":"Select 100 random targets from the public set; obtain independent native-speaker phonetic transcriptions (with consensus adjudication) and compute exact match rate against the published ASR ensemble transcripts; if agreement falls below 95%, recompute all system scores using the human transcripts and report the delta in strict accuracy.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The benchmark's automatic scorer depends on the three-ASR ensemble producing fixed transcripts that match intended pronunciations for targets such as English abbreviations, model names, and mixed forms. If the ensemble errs on these (precisely the cases where ASR is weakest), the reported accuracies (0.879 best, <0.60 for others) measure agreement with potentially flawed references rather than true pronunciation correctness. The abstract mentions ASR-route diagnostics and ablations but supplies no per-target error rates or human agreement figures, leaving the proxy assumption untested at the scale of the 992 targets.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims to introduce CN-NewsTTS Bench v0.1, an open target-level benchmark for evaluating Chinese news TTS products on pronouncing complex written forms such as scores, model names, English abbreviations from raw text without additional hints. It includes development and test sets, 992 targets with fixed three-ASR ensemble transcripts as ground truth, an automatic scorer, and initial results for seven systems with the best achieving 0.879 strict accuracy. The work also provides ASR diagnostics, ablations, category results, and confidence intervals.","tokens_in":1820,"tokens_out":340,"duration_ms":33833,"significance":"Should the ASR transcripts prove to be a reliable proxy for intended pronunciations, the benchmark would provide a useful automatic evaluation tool for a common challenge in Chinese TTS for news content. The open release, ablations, and confidence intervals are positive features that support reproducibility and allow for nuanced analysis of system performance across categories.","major_comments":[{"comment":"The reported system accuracies (0.879 best, several below 0.60) are measured using transcripts from a three-ASR ensemble as ground truth for the 992 targets. The manuscript does not provide human validation, agreement rates, or error analysis for these transcripts on ambiguous targets, which is central to validating the benchmark's automatic scorer as a measure of correct pronunciation rather than agreement with ASR.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":"The primary concern is the unvalidated ground truth; if authors can provide even a small human-annotated subset showing high agreement, it would strengthen the submission significantly."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for highlighting the importance of validating the ASR-derived ground truth. We address the single major comment below.","responses":[{"response":"We agree that the absence of human validation and agreement rates for the ensemble transcripts on ambiguous targets is a limitation. The three-ASR ensemble was chosen to reduce single-system errors via majority voting, and the manuscript already includes ASR-route diagnostics, subset ablations, and category results to characterize consistency. However, these do not substitute for direct human assessment. We will add (i) inter-ASR agreement rates across the 992 targets and (ii) a human validation study on a stratified sample of 150 targets (including ambiguous cases) with reported agreement to the ensemble, to be included in the revised manuscript.","revision_made":"yes","referee_comment":"[Abstract] The reported system accuracies (0.879 best, several below 0.60) are measured using transcripts from a three-ASR ensemble as ground truth for the 992 targets. The manuscript does not provide human validation, agreement rates, or error analysis for these transcripts on ambiguous targets, which is central to validating the benchmark's automatic scorer as a measure of correct pronunciation rather than agreement with ASR."}],"tokens_in":1293,"tokens_out":272,"duration_ms":14891,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The colleague should know this paper puts out CN-NewsTTS Bench v0.1 with 992 public targets drawn from Chinese news text and reports system accuracies from under 0.60 up to 0.879 on an automatic scorer.\n\nWhat the work actually does is collect dense written forms like model names, abbreviations, ranges, and mixed scripts, then score whether TTS systems produce the right spoken form from raw input alone. They release a 200-record dev set, 800-record test set, the targets themselves, three-ASR ensemble transcripts, category breakdowns, ASR ablations, confidence intervals, and metadata on the seven tested systems. That package is the concrete new thing.\n\nThe construction pipeline and reporting look solid enough on the surface for a practical resource. Releasing the targets and scorer openly lets others run their own checks.\n\nThe soft spot is the ground truth. The scores rest on the ASR ensemble transcripts matching intended pronunciations, yet the abstract gives no human agreement numbers or per-target error rates on the ambiguous cases that matter most. If the ensemble slips on abbreviations or model names, the reported accuracies measure agreement with the references rather than actual correctness. The stress-test concern holds up from what is shown.\n\nThis is for TTS developers and evaluators who work on Chinese news content and need a repeatable test for raw-input handling. A reader in that area gets a ready-made set of targets and some baseline numbers.\n\nIt deserves peer review. The benchmark fills a narrow but real gap even if the validation step needs strengthening.","headline":"The paper releases a usable open benchmark for Chinese TTS pronunciation on raw news text with tricky written forms, but the ASR-derived ground truth has no shown human validation.","tokens_in":2300,"tokens_out":391,"would_cite":false,"duration_ms":26228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A new benchmark measures how accurately Chinese news TTS systems pronounce tricky written forms like scores and abbreviations from raw text alone.","keywords":["Chinese TTS","news text pronunciation","automatic benchmark","target-level evaluation","ASR ensemble","raw text input","pronunciation accuracy","written form handling"],"falsifier":"A human listening test or independent transcription that systematically differs from the three-ASR ensemble transcripts on a large fraction of the targets would falsify the ground truth used by the scorer.","tokens_in":2601,"feed_emoji":"🔊","tokens_out":690,"duration_ms":24054,"temperature":0.7,"pith_summary":"The paper introduces CN-NewsTTS Bench v0.1, an open target-level benchmark designed to test whether TTS products correctly pronounce dense written forms in Chinese news text when given only raw input and no extra rules or hints. It supplies 992 auto-evaluable targets, fixed transcripts generated by a three-ASR ensemble, and an automatic scorer, along with results across seven product systems. The evaluation shows clear performance differences, with the strongest system reaching 0.879 strict accuracy while others fall below 0.60. This setup matters because such written forms occur often in news and can change the spoken meaning if the TTS output deviates from the intended pronunciation.","feed_headline":"Benchmark finds Chinese news TTS accuracy ranges from below 0.6 to 0.879","feed_subtitle":"Target-level test of 992 written forms shows some systems mispronounce scores, abbreviations and model names from raw text.","key_machinery":"The target-level automatic scorer that matches TTS output against the ASR-ensemble transcripts for 992 specific written forms in news text.","core_discovery":"CN-NewsTTS Bench v0.1 supplies a 200-record development set, an 800-record public test set, 992 public targets drawn from Chinese news, fixed transcripts from a three-ASR ensemble, an automatic target scorer, and initial results for seven product TTS systems that demonstrate varying pronunciation accuracy on raw text without user-side interventions.","pith_inferences":["Developers could use the benchmark to prioritize improvements in handling hyphenated model names and unit symbols without relying on LLM rewriting.","The same target-level approach might transfer to news TTS evaluation in other languages that mix scripts or symbols.","If ASR errors cluster on particular target categories, future versions could incorporate human-verified subsets for those categories.","Widespread adoption might shift industry focus from general naturalness metrics toward precise pronunciation of frequent written forms."],"forward_implications":["TTS systems can now be compared automatically on pronunciation of written forms such as percentages, English abbreviations, and mixed names without manual edits or SSML.","Performance gaps exist among current products, with top accuracy at 0.879 and several systems below 0.60 on the same targets.","Category-level breakdowns and ASR-route diagnostics allow identification of specific weakness patterns in different systems.","The public test set and scorer enable repeated evaluation as new TTS versions are released."],"fun_headline_variants":["CN-NewsTTS Bench tests TTS on 992 raw Chinese news targets","Chinese news TTS accuracy spans below 0.6 to 0.879","Benchmark evaluates seven TTS systems on tricky news forms","992 targets measure TTS pronunciation without manual edits"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The fixed transcripts from the three-ASR ensemble accurately capture the intended pronunciations for the 992 targets.","fun_headline_variants_meta":{"raw":{"variants":["CN-NewsTTS Bench tests TTS on 992 raw Chinese news targets","Chinese news TTS accuracy spans below 0.6 to 0.879","Benchmark evaluates seven TTS systems on tricky news forms","992 targets measure TTS pronunciation without manual edits"]},"model":"grok-4.3","cost_usd":0.004807,"raw_usage":{"total_tokens":2350,"prompt_tokens":639,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":48074500,"prompt_tokens_details":{"text_tokens":639,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1644,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":639,"tokens_out":67,"duration_ms":13519,"temperature":1.0,"reasoning_tokens":1644,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T00:01:06.267986+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A human listening test or independent transcription that systematically differs from the three-ASR ensemble transcripts on a large fraction of the targets would falsify the ground truth used by the scorer.","supporting_citations":[],"review_version":1}