{"id":"c6a6d9c9-0265-4de7-80be-4ad2ba66eda1","arxiv_id":"2505.11200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new Chinese TTS benchmark and protocol, Audio Turing Test, shows top LLM-based TTS models achieve only about 0.4 out of 1.0 on human-likeness, far below real human speech.","lead":"This paper introduces the Audio Turing Test, a Chinese text-to-speech evaluation benchmark that asks listeners to judge whether audio clips sound human, plus an automatic scoring model. It finds the best tested LLM-based TTS system still scores well below human speech, suggesting current systems are more machine-like than MOS scores imply.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Auto-ATT's 'strong alignment' is unverified: Table 4 mixes training and held-out voices, and the reported Kendall tau of about 0.33 is only moderate agreement; held-out-voice performance is not reported.","rationale":"As a second-pass reviewer, I find the human-evaluation core of the paper credible: the ATT protocol's forced-choice design, attention-check traps, expert consistency review, and GLMM analysis collectively support the claim that ATT differentiates TTS systems (e.g., Seed-TTS 0.417 with 95% HDI [0.398,0.438] vs GPT-4o 0.138 [0.118,0.158]). The main unresolved risk is the automatic-judge contribution, Auto-ATT. The reader correctly flags a train/test voice overlap: Section 3.5 reserves one voice per family for testing, yet Section 4.2.2 evaluates Auto-ATT on the same audio data as the human evaluation, which includes the training voices. I agree this is the primary load-bearing concern. My addition is that even the leaked evaluation numbers are not as strong as claimed: Kendall tau of about 0.33 is a moderate association, and the table's 'lower is better' caption conflicts with the text's use of tau as a correlation, so the claim of 'strong alignment' is numerically overstated. A held-out-only recomputation is a concrete and decisive check. If it fails, the paper's Auto-ATT claim must be downgraded. The human benchmark verdict is unaffected, so the CONDITIONAL assessment remains appropriate, conditional primarily on revalidating Auto-ATT on held-out voices and clarifying the statistic.","tokens_in":15614,"tokens_out":8797,"duration_ms":84118,"concrete_test":"Recompute the Table 4 analysis on held-out voices only: for each of the five dimensions and for the 'All' ranking, compute Kendall's tau between Auto-ATT-predicted HLS and human HLS using only the one reserved voice per trained family (4 voices) plus the four GPT-4o voices (8 voices total), with permutation p-values; repeat for the 12 training voices. If the held-out tau is not significantly positive or is substantially lower than the in-sample tau, Auto-ATT's reported alignment reflects training-data fit rather than generalization. Also report explicitly whether the table entries are Kendall's tau correlation or normalized distance, since the caption and text currently conflict.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2.2 states Auto-ATT is run on the same audio data as the human evaluation to predict HLS, then ranks voices per dimension and compares with human rankings using Kendall's distance. Section 3.5 says three of four voices per family were used for training, with one voice reserved for testing. Therefore at least 12 of the 20 voices in the Section 4.2.2 ranking are training voices (three each for Cosyvoice, MiniMax, Seed-TTS, and Step-Audio); GPT-4o's four voices were not in Auto-ATT training but are not reported separately. The reported Kendall tau for 'All' is 0.3316 (p=0.0398); while statistically nonzero, this is a moderate effect, not 'strong alignment,' and the table caption's 'Lower tau is better' is inconsistent with interpreting tau as a correlation. The marginal advantage over the unfinetuned Qwen2-Audio model (0.3474) is small. Because the authors do not report the agreement restricted to held-out voices, the claim that Auto-ATT is a reliable judge for new voices is not supported by the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes the Audio Turing Test (ATT), a Chinese TTS human-likeness evaluation framework consisting of a five-dimension corpus (ATT-Corpus), a Turing-test-style human listening protocol with trap items and free-text justifications, and a Human-likeness Score (HLS) computed from ternary Human/Unclear/Machine labels. The authors collect judgments from 857 native Chinese listeners for 20 voices across five TTS systems, report model-level HLS rankings, and claim that even the best system (Seed-TTS) reaches only about 0.4 HLS, far below real human speech. They also fine-tune Qwen2-Audio-Instruct on human labels to create Auto-ATT, a model-as-a-judge automatic evaluator, and evaluate it on trap items and on agreement with human rankings. The main human benchmark protocol is carefully designed, but the Auto-ATT generalization claim and some central quantitative comparisons are not fully supported by the reported experiments.","tokens_in":15856,"tokens_out":5253,"duration_ms":54769,"significance":"If the claims hold, ATT would be a valuable contribution to TTS evaluation: it addresses genuine limitations of MOS by using a simple forced-choice human-likeness judgment, includes trap items for attention screening, applies expert consistency checks, and provides a reproducible white-box corpus. The multi-dimensional design and the finding that leading LLM-based TTS systems still fall substantially short of human speech on a human-likeness criterion are important and potentially impactful results. The machine-checkable strengths include the detailed protocol, the GLMM convergence diagnostics, the qualitative attribution coding, and the public release of the corpus and tools. However, the strongest auxiliary claim, that Auto-ATT is a fast and reliable automatic judge with strong alignment to human evaluations, is weakened by the absence of held-out-voice results and by the moderate reported agreement; the headline comparison with human speech also lacks a measured human-reference HLS. These issues are fixable with additional analyses rather than being fundamental to the ATT human benchmark itself.","major_comments":[{"comment":"The claim that Auto-ATT shows strong alignment with human evaluations and is reliable for new voices is not supported by the reported experiment. Section 3.5 states that within each of four model families, one voice was reserved for testing and the remaining three voices were used for training, but Section 4.2.2 evaluates Auto-ATT on the same audio data as the human evaluation, which includes all four voices per family. At least 12 of the 20 voices ranked in Table 4 are therefore training voices, and GPT-4o voices are not separated in the reported agreement. The Kendall tau of 0.3316 for the All dimension is moderate, not strong, and the held-out-voice agreement is not reported. Please report the ranking agreement computed only on the reserved held-out voices, with confidence intervals, and restrict the generalization claim to what that analysis supports.","section":"Section 3.5 and Section 4.2.2, Table 4"},{"comment":"The headline result that Seed-TTS reaches only about 0.4 HLS and is 'considerably lower than that of real human speech' lacks a measured human reference. Since HLS is defined from ternary labels, human recordings are not guaranteed to receive a score of 1.0; listeners may label genuine human clips as Unclear or Machine. The protocol includes human recordings as trap items, so the authors have the data to compute an HLS for real human speech. Please report the mean HLS and uncertainty for the human reference clips used in the evaluation, and use this measured value rather than the nominal value of 1.0 when claiming a large remaining gap between synthetic and human speech.","section":"Section 4.1.2"},{"comment":"The interpretation of Table 4 is internally inconsistent. The caption says 'Lower tau is better,' but Kendall's tau is a rank correlation coefficient for which higher values indicate better agreement, and the p-value column is consistent with testing a correlation. If the authors instead intend a Kendall distance or discordant-pair proportion, that quantity and its reference should be defined explicitly. In addition, the 'All' row reports exactly the same value as the 'Polyphonic Characters' row (0.3316, p=0.0398); an aggregate over all dimensions would not be expected to coincide exactly with one subset, so this appears to be a reporting error that should be corrected.","section":"Table 4 caption and rows"},{"comment":"The comparison with UTMOSv2 and DNSMOS Pro on trap items is not calibrated fairly. The figure normalizes all predictions to a 0-1 scale and applies the same 0.5 decision threshold to models that were not designed to output ternary human-likeness probabilities; MOS predictors are not trained to separate human speech from deliberately flawed synthetic speech on this scale, so an F1 score at an arbitrary threshold does not establish that Auto-ATT is intrinsically superior. Please report threshold-free measures such as AUROC or area under the precision-recall curve, and show the raw score distributions for all three models on the trap items.","section":"Section 4.2.1, Figure 3"}],"minor_comments":[{"comment":"The description of Auto-ATT training data is inconsistent with the in-distribution/out-of-distribution split in Table 4. Section 3.5 says training focuses on 'three capability subsets' but lists four phenomena and does not clearly state whether Special Characters and Numerals was included; Table 4 labels Special Characters and Numerals as an in-distribution dimension. Please reconcile the number of subsets and the dimension names.","section":"Section 3.5 and Table 4"},{"comment":"There are several typographical errors: 'shownshown' in Section 3.3, 'ofof' in the paragraph introducing Section 3, and 'choral quality' in the Related Works section, which should likely be 'vocal quality.'","section":"Section 3.3 and Section 3 text"},{"comment":"The name of the DNSMOS baseline is inconsistent: the text refers to 'DMSMOS Pro' in one place and 'DNSMOSPro' in another, while the figure uses 'DNSMOS Pro.' Please use a single consistent name matching the cited reference.","section":"Section 4.2.1"},{"comment":"Figure 2 shows point estimates without error bars, and Table 6 gives voice-level HLS values to four decimal places without uncertainty intervals. Since the text draws conclusions about differences between voice styles within a model (for example Seed-TTS 'Skye' at 0.47 versus lower-ranked voices), please provide confidence intervals or posterior intervals for these per-dimension and per-voice estimates.","section":"Figure 2 and Table 6"},{"comment":"The voice style names are inconsistent between Table 2 and Table 6: Table 2 lists 'Sky' for Seed-TTS while Table 6 uses 'skye,' and MiniMax names such as 'siyuan' and 'xinyue' in Table 6 do not match the platform-style names in Table 2. Please harmonize the naming and clarify whether these are the same voice styles.","section":"Table 6"},{"comment":"The text refers to 'Kendall's distance [1],' but the cited reference is titled 'The Kendall rank correlation coefficient.' Please use the correct terminology or cite a genuine distance measure, and ensure the statistic reported in Table 4 matches the definition used.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core ATT human-evaluation benchmark is a solid and useful contribution, and the issues I raise do not warrant rejection. The Auto-ATT generalization and comparison claims are the main weakness; they can be addressed with a held-out-voice evaluation, a measured human-reference HLS, and threshold-free trap-item metrics. I would also ask the editor to check the Table 4 duplicate 'All' row before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The human evaluation side of this paper is a solid, useful contribution. The ATT-Corpus is thoughtfully constructed for Chinese-specific challenges (code-switching, polyphones, classical prose, paralinguistic emotion), the trap-item protocol with ternary judgments is a real improvement over plain MOS, and the GLMM analysis with expert consistency checks is careful. The headline result—Seed-TTS reaches only about 0.4 HLS, far below real human speech—is likely robust and worth knowing if you work on TTS evaluation. Releasing the white-box corpus and Auto-ATT adds practical value.\n\nWhere the paper overreaches is the automatic evaluation. Auto-ATT is fine-tuned on human labels for three voices per model family, then evaluated on the same full audio set that includes those training voices. The Kendall tau values in Table 4 therefore partly measure in-sample fit, and the paper never reports held-out-voice agreement. On top of that, tau around 0.33 is at best moderate agreement—not the 'strong alignment' the abstract claims—and the table caption saying 'Lower tau is better' is simply wrong for a correlation. The marginal gain over the unfinetuned Qwen2-Audio baseline is also small. None of this undermines the human ATT benchmark itself, but the Auto-ATT story needs either held-out results or a much more modest claim.\n\nMinor issues: human judgment data are not released, which will slow independent verification, and the paper's own limitation section only mentions language specificity. There is also some sloppiness in table captions and terminology that a referee should clean up.\n\nOverall, this is a worthwhile evaluation paper. The human benchmark deserves to be in the literature, and the protocol is a good model for future TTS listening tests. The automatic judge part needs another revision round before it can be cited as reliable. If I were an editor, I would send this to peer review with a request to fix the Auto-ATT evaluation and the reporting of the correlation metric.","headline":"ATT is a genuinely useful Chinese TTS evaluation benchmark; trust the human-study results, but treat the Auto-ATT 'strong alignment' claim as unverified until it is shown on held-out voices.","tokens_in":16366,"tokens_out":1184,"would_cite":true,"duration_ms":13866,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Turing-style listening test shows that even the strongest LLM-based Chinese TTS system, Seed-TTS, scores only about 0.4 on human-likeness, well below real human speech.","keywords":["text-to-speech evaluation","Turing test","human-likeness score","Chinese speech synthesis","mean opinion score","automatic evaluation","Qwen2-Audio","code-switching"],"falsifier":"Two checks would settle the central claims. First, compute the HLS of the genuine human recordings already used as trap items: if their score is not clearly above Seed-TTS's 0.4, the 'considerably lower than real human speech' claim fails. Second, retrain Auto-ATT on three voices per model and evaluate only on the held-out fourth voice; if the Kendall-tau agreement with human rankings drops to chance, Auto-ATT is not a general judge of new voices.","tokens_in":15405,"feed_emoji":"🎧","tokens_out":7347,"duration_ms":70212,"temperature":0.7,"pith_summary":"This paper proposes the Audio Turing Test (ATT), a Chinese-language benchmark that replaces the five-point Mean Opinion Score with a simple question: does this voice sound human? It contributes a multi-dimensional corpus (ATT-Corpus) covering code-switching, polyphonic characters, paralinguistic emotion, classical poetry, and numerals, plus a protocol in which listeners label clips Human, Unclear, or Machine and must pass hidden trap items. The central result is that ATT cleanly separates five state-of-the-art LLM-based TTS systems, and that the best system, Seed-TTS, reaches a Human-likeness Score of only about 0.4, far below what the paper expects of real human speech. The paper also trains Auto-ATT, an automatic judge built on Qwen2-Audio-Instruct, and reports that it agrees with human rankings and detects flawed synthetic clips far better than existing MOS predictors. If right, ATT gives the field a more sensitive, interpretable yardstick for the remaining gap between synthetic and human voices.","feed_headline":"Best AI voice scores only 0.4 on human-likeness test","feed_subtitle":"A new Audio Turing Test ranks five LLM text-to-speech systems and shows the gap to real speech remains large.","key_machinery":"The load-bearing object is the Human-likeness Score (HLS), defined as the average over clips of $s_i = \\mathbf{1}(\\text{Label}=\\text{Human}) + 0.5\\,\\mathbf{1}(\\text{Label}=\\text{Unclear})$, with Machine scored 0. The measurement apparatus is ATT-Corpus, a semi-automatically built corpus spanning five Chinese linguistic difficulty dimensions, combined with a protocol that inserts one flawed synthetic clip and two genuine human recordings into every ten clips and discards batches that miss them. Auto-ATT carries the automatic-evaluation half: Qwen2-Audio-Instruct is adapted by LoRA, and its output logits for the tokens Human, Unclear, and Machine are softmaxed and converted into a weighted score, trained with a combination of Bradley-Terry and mean-squared-error losses.","core_discovery":"The paper's claim is that human-likeness in TTS can be measured directly by asking listeners whether synthesized speech is human, and that this simpler ternary judgment yields sharper distinctions than MOS. Under ATT, a clip earns 1 point for a Human label and 0.5 for Unclear, and a system's Human-likeness Score is the average across clips. Evaluated on 20 voice styles from five model families, Seed-TTS ranks first at roughly 0.4, MiniMax-Speech follows near 0.39, Step-Audio and CosyVoice sit around 0.22-0.29, and GPT-4o lags at 0.13; each model retains its rank in a held-out black-box split. The paper argues this ordering, and the finding that no system comes close to a score of 1, shows ATT exposes a large gap that MOS-style scores hide. In addition, the fine-tuned Auto-ATT judge reproduces the human voice ranking (Kendall distances around 0.27-0.34 across dimensions) and scores trap items with F1 of 0.92, while UTMOSv2 scores 0.14 and DNSMOS Pro 0.00 at the same threshold.","pith_inferences":["An implicit consequence is that published MOS claims that modern TTS is nearly indistinguishable from human speech may be systematically over-optimistic; the paper's 0.4 ceiling suggests the metric rather than the technology deserves scrutiny.","A testable extension would be to report HLS for genuine human recordings under the same protocol as a calibration anchor; the paper uses human clips as trap items but does not publish their HLS, leaving the size of that gap to be quantified.","The Auto-ATT results would generalize more convincingly if the Kendall-tau comparison were rerun on voices never seen in training; the paper reserves one voice per family for testing but evaluates on the same audio as the human study, so part of the agreement may reflect familiar voices.","If the ATT protocol transfers to other languages, the same ternary-judgment design with trap items could become a common currency for cross-lingual human-likeness benchmarks, with Auto-ATT's zero-shot cross-lingual transfer as a natural next test."],"forward_implications":["TTS developers can use ATT-Corpus and the HLS protocol to compare systems on specific weaknesses: Seed-TTS is strongest at code-switching and numerals but falls behind MiniMax-Speech on classical Chinese prose.","Auto-ATT provides a fast proxy that ranks voices in nearly the same order as human listeners, so model iteration no longer has to wait for crowdsourced listening tests.","Because the black-box and white-box splits give the same model ordering, published white-box results can be read as a fair preview of blind evaluation outcomes.","The qualitative justifications collected by the protocol locate the common failure modes, such as abrupt prosody, missing micro-pauses, flattened emotion, and artifacts like foreign accent and hiss, giving system designers concrete targets."],"supporting_citations":[{"why":"Supplies the Turing-test framing that ATT adapts to audio.","marker":"[11]"},{"why":"Documents the reporting and bias problems in subjective TTS evaluation that motivate trap items and justification checks.","marker":"[5]"},{"why":"Argues MOS loses resolution at high quality, the ceiling effect ATT is designed to overcome.","marker":"[23]"},{"why":"Is the strongest evaluated system; its 0.4 Human-likeness Score anchors the paper's headline result.","marker":"[2]"},{"why":"Provides the base audio-language model that Auto-ATT fine-tunes.","marker":"[6]"},{"why":"Is the UTMOSv2 baseline that Auto-ATT outperforms on trap items.","marker":"[3]"},{"why":"Is the DNSMOS Pro baseline that Auto-ATT outperforms on trap items.","marker":"[7]"},{"why":"Is the low-rank adaptation method used to train Auto-ATT.","marker":"[14]"},{"why":"Provides the generalized linear mixed model used to test whether human evaluation scores differ significantly across systems.","marker":"[4]"}],"fun_headline_variants":["Best AI voice scores 0.4 on Audio Turing Test","Audio Turing Test: top AI voice hits 0.4, far from human","New Audio Turing Test: best TTS scores 0.4, still not human","AI voices flunk human-likeness test: best score 0.4"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that Auto-ATT reliably predicts human judgments assumes that the voices it is scored on are not the voices it was trained on; the paper trains Auto-ATT on three voices per model family but reports agreement on audio that includes all four voices, so the agreement numbers could partly reflect in-sample familiarity.","fun_headline_variants_meta":{"raw":{"variants":["Best AI voice scores 0.4 on Audio Turing Test","Audio Turing Test: top AI voice hits 0.4, far from human","New Audio Turing Test: best TTS scores 0.4, still not human","AI voices flunk human-likeness test: best score 0.4"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000592,"raw_usage":{"total_tokens":2849,"prompt_tokens":1091,"completion_tokens":1758,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":707,"completion_tokens_details":{"reasoning_tokens":1674}},"tokens_in":707,"tokens_out":1758,"duration_ms":11477,"temperature":1.0,"reasoning_tokens":1674,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:56:01.803502+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Two checks would settle the central claims. First, compute the HLS of the genuine human recordings already used as trap items: if their score is not clearly above Seed-TTS's 0.4, the 'considerably lower than real human speech' claim fails. Second, retrain Auto-ATT on three voices per model and evaluate only on the held-out fourth voice; if the Kendall-tau agreement with human rankings drops to chance, Auto-ATT is not a general judge of new voices.","supporting_citations":[{"cited_title":"The turing test: the first 50 years","cited_arxiv_id":null,"evidence_quote":"Supplies the Turing-test framing that ATT adapts to audio."},{"cited_title":"Why we should report the details in subjective evaluation of tts more rigorously","cited_arxiv_id":null,"evidence_quote":"Documents the reporting and bias problems in subjective TTS evaluation that motivate trap items and justification checks."},{"cited_title":"The limits of the mean opinion score for speech synthesis evaluation","cited_arxiv_id":null,"evidence_quote":"Argues MOS loses resolution at high quality, the ceiling effect ATT is designed to overcome."},{"cited_title":"The t05 system for the voicemos challenge 2024: Transfer learning from deep image classifier to naturalness mos prediction of high-quality synthetic speech","cited_arxiv_id":null,"evidence_quote":"Is the UTMOSv2 baseline that Auto-ATT outperforms on trap items."},{"cited_title":"Lora: Low-rank adaptation of large language models","cited_arxiv_id":null,"evidence_quote":"Is the low-rank adaptation method used to train Auto-ATT."},{"cited_title":"Generalized linear mixed models: a practical guide for ecology and evolution","cited_arxiv_id":null,"evidence_quote":"Provides the generalized linear mixed model used to test whether human evaluation scores differ significantly across systems."}],"review_version":1}