{"id":"d9202eb3-faf5-4d6c-b915-1e37dff65e0c","arxiv_id":"2608.09930","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Automated TTS evaluators, both MOS predictors and Audio-LLM judges, systematically fail to detect most linguistically grounded speech errors, with MOS predictors collapsing onto signal-level artifacts.","lead":"Researchers built a speech-quality benchmark that breaks 'naturalness' into 10 separate listening dimensions, such as pronunciation, stress, and emotion, and had trained linguists label 860 synthetic utterances. They found that both neural quality-scoring models and audio-capable language models miss most of these dimensions, so automated TTS evaluation is less reliable than its users assume.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sec. 3.2's 'controlled upper-bound' assumption is not supported by the paper's own miss-rate (up to 48.4%) and cross-dimensional contamination data, so model 'blind spots' may reflect co-occurring artifacts rather than the named dimensions, and do not yet transfer to natural TTS failures.","rationale":"The paper makes a valuable contribution: a 10-dimension linguistically grounded annotation schema, 860 human-annotated utterances, and a systematic audit of four MOS predictors and four Audio-LLM judges under four prompting conditions. The reported pattern — MOS predictors detected on acoustic-manipulation dimensions and null on text/IPA-based dimensions, while Audio-LLMs show sparse, prompt-dependent sensitivity — is internally consistent and carefully tabulated. My stress-test pass converged on the same load-bearing assumption the reader identified: the controlled-perturbation design (Sec. 3.2) is the bridge from 'models fail on these stimuli' to 'models cannot serve as dimension-level diagnostic for TTS errors.' That bridge requires the perturbations to be pure, representative instances of the named dimensions. The paper's own numbers weaken that requirement: a third of intended-error samples were not perceived as errors (Table 9), inter-annotator agreement is low for two acoustic-manipulation dimensions (Table 2), and Figure 2 shows strong cross-dimensional co-occurrence. The hedge in Sec. 4.4 about Expressiveness is an internal admission that acoustic-manipulation stimuli may be detected through shared low-level artifacts rather than the intended dimension. Since the detected dimensions (paralinguistic/acoustic) and the null dimensions (word-level/text-based) are confounded with generation pathway, the headline 'signal-level vs. linguistic' split is not yet a clean dimension-level result. This does not invalidate the benchmark as a resource, nor the descriptive claim about these stimuli; it does mean the Conclusion's general statement about 'current automated evaluators' requires the proposed reanalysis or an external validation on natural TTS errors. The reader's CONDITIONAL verdict already captures this uncertainty, so I recommend no change.","tokens_in":27170,"tokens_out":9023,"duration_ms":81373,"concrete_test":"Recompute the headline Kendall-tau values in Table 3 on the subset of samples where the target dimension is the ONLY majority-labeled failure (all other nine dimensions majority-clean), using the released per-dimension human labels. If significant correlations (e.g., MOS predictors on Speech Rate/Human Plausibility/Expressiveness; Gemini on Phonetic Accuracy) collapse to null on this subset, the reported dimension-level sensitivity is an artifact of co-occurring signal degradation; if they persist, the dimension attribution is supported. If the per-dimension labels are not yet available, this reanalysis should be a release condition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim generalizes from benchmark stimuli to 'current automated evaluators' (Sec. 5). The bridge is the Sec. 3.2 assertion that a judge failing on a deliberately introduced perturbation also fails on subtler natural deviations. This upper-bound inference requires each perturbation to be a pure, representative instance of its target dimension. The paper's own data contradict that purity: Table 9 shows 32.3% of intended-error samples majority-labeled clean, with miss rates up to 48.4% (Speaker Identity) and 46.2% (Emotional Appropriateness); Table 2 reports low inter-annotator agreement for Intonation (alpha=0.469) and Expressiveness (alpha=0.460). Figure 2 further shows substantial cross-dimensional failure co-occurrence, and Sec. 4.4 concedes that MOS sensitivity to Expressiveness may 'reflect ... shared low-level acoustic properties' rather than the dimension. The same confound applies to the acoustic-manipulation dimensions (Intonation, Speech Rate, Speaker Identity, Human Plausibility): Praat PSOLA and DSP edits introduce audible artifacts that can drive human 'failure' labels and model scores simultaneously. Because word-level and text-based prosodic dimensions were generated via IPA/text modification while the detected paralinguistic dimensions were generated via audio manipulation, the 'MOS predictors track signal-level artifacts' finding is partly a statement about generation pathway, not about perceptual dimensions. Consequently, the benchmark's null results (and its positive detections) are not yet established as dimension-level findings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a meta-evaluation benchmark for automated TTS evaluators, decomposing 'naturalness' into ten linguistically motivated perceptual dimensions across word, prosodic, and paralinguistic levels. A dataset of 860 synthesized utterances is annotated by three trained linguists on all dimensions, with error stimuli generated via LLM-modified IPA transcription, plain-text modification, API-enforced emotion tags, and Praat/DSP acoustic manipulation. The authors benchmark four MOS predictors and four Audio-LLM judges under four prompting conditions with and without reference transcripts, reporting Kendall's tau and significance tests against human ground truth. Their main conclusions are that MOS predictors track signal-level artifacts but miss word-level and prosodic dimensions, while Audio-LLM judges show selective, prompt-dependent sensitivity that never generalizes across the full dimension set, and that no current automated judge is reliable as a dimension-level diagnostic. The dataset, annotation schema, and evaluation code are publicly released.","tokens_in":27383,"tokens_out":6184,"duration_ms":50838,"significance":"If the central claim holds, this is a valuable diagnostic contribution: it is, to my knowledge, the first benchmark to audit automated TTS evaluators at the level of linguistically grounded perceptual dimensions with expert-annotated ground truth, and it demonstrates that holistic naturalness scores obscure heterogeneous failure types. The paper's strengths include the carefully constructed annotation rubric, the use of trained linguist raters, the reporting of inter-annotator reliability with bootstrap confidence intervals, the systematic variation of prompting conditions, and the public release of data and code. The observation that schema guidance can recover word-level sensitivity in some Audio-LLM judges is an interesting and actionable finding. The paper's negative conclusions, however, rest on a stimulus-validity assumption that the authors' own data substantially weaken; addressing this issue would materially strengthen the contribution.","major_comments":[{"comment":"The 'controlled upper-bound' assumption is not supported by the paper's own miss rates. The statement that 'a judge that cannot detect a deliberately introduced perturbation cannot detect subtler naturally occurring deviations of the same type' presupposes that the perturbations are valid, perceptible instances of the intended dimension. However, Table 9 reports that 32.3% of intended-error samples were majority-labeled clean, with miss rates as high as 48.4% for Speaker Identity Consistency and 46.2% for Emotional Appropriateness. If a large fraction of the deliberately introduced errors are not even perceived by the human raters who define ground truth, then the model null results on those dimensions do not establish that the models cannot detect the target errors; they may simply be failing on stimuli that do not instantiate the named dimension. This weakens the central claim in Section 5 that current automated evaluators exhibit systematic blind spots for these dimensions.","section":"3.2, Table 9"},{"comment":"The generation-pathway confound undermines the interpretation of the MOS-predictor results. The word-level dimensions (Phonetic Accuracy, Lexical Stress) and two prosodic dimensions (Prosodic Stress, Prosodic Boundary Placement) were created through linguistic modifications (IPA or text), whereas Speech Rate, Expressiveness, Speaker Identity, and Human Plausibility were created through acoustic manipulation. The finding that MOS predictors 'align strongly with signal-level artifacts' is therefore confounded with the generation pathway: the predictors may be responding to low-level acoustic degradation rather than to the perceptual dimension itself. The paper acknowledges this possibility for Expressiveness in Section 4.4 ('shared low-level acoustic properties') but not for Speech Rate, Speaker Identity, or Human Plausibility. Without a control condition that applies analogous acoustic perturbations without the target dimension, or a validation on naturally occurring TTS failures, the claim that MOS predictors track signal-level artifacts rather than generation artifacts is not established.","section":"3.2, 4.4"},{"comment":"The ground-truth reliability is insufficient to support the strong null claims on several dimensions. Krippendorff's alpha is only 0.469 for Intonation and 0.460 for Expressiveness, both below common acceptability thresholds, and Emotional Appropriateness has only 46 samples with alpha=0.596 and a 95% CI extending to 0.767. With such noisy labels, the absence of significant model correlations on these dimensions (Tables 3 and 4) cannot be interpreted as evidence of model blind spots. Furthermore, the paper does not report a power analysis; given the small per-dimension sample sizes (n=46–128), the non-significant results may simply reflect low power. The conclusions in Section 5 should be restricted to dimensions with acceptable reliability and adequate sample size, or accompanied by a sensitivity analysis showing the results hold when unreliable dimensions are excluded.","section":"Table 2"},{"comment":"The significance testing framework does not account for multiple comparisons. With four MOS predictors, four Audio-LLM judges, ten dimensions, four prompting conditions, and two transcript settings, the paper conducts a very large number of significance tests at the alpha=0.05 level, yet no multiple-comparison correction is reported. The specific pattern of 'selective, prompt-dependent detection' may therefore include false positives, and the negative correlations in Table 3 could reflect noise. Correcting for multiple comparisons, or pre-registering the primary comparisons, would strengthen the interpretation of which dimensions are genuinely detected by each model.","section":"4.3, Tables 3–4"}],"minor_comments":[{"comment":"There is a typo: 'acosutic' should be 'acoustic'.","section":"3.2"},{"comment":"The caption 'AUDIO-LLMSjudge' is missing a space; also, the table uses the symbols '†' and '△' that are explained only in the caption, which may be confusing in the main text.","section":"Table 3 caption"},{"comment":"The main text states that all dimensions are rated as binary (0/1), but Table 6 shows ternary scales for several dimensions; the collapse from ternary to binary is described only in Section 6.2. Please make the collapse explicit in Section 3.3.","section":"3.3, 6.2"},{"comment":"The phrase 'perceptually salient errors on higher-level dimensions' in the discussion of miss rates is inconsistent with the fact that Lexical Stress (a word-level dimension) has a 38.4% miss rate; consider revising.","section":"6.5"},{"comment":"The paper states that Kendall's tau_b is the unified comparison metric, but the significance flags in Tables 3 and 4 come from Mann-Whitney U and McNemar tests, not from tau; please clarify how the reported p-values relate to the tau values.","section":"4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a substantial contribution to TTS evaluation benchmarking, and the public release of the dataset and code is commendable. My main concern is that the framing of the conclusions goes beyond what the stimulus design supports; however, this is fixable by qualifying the claims and adding the recommended analyses. The use of a single TTS architecture (Cartesia Sonic) for IPA-driven samples is acknowledged as a limitation but may still affect the generality of the word-level findings. I would encourage the editor to seek a revision that addresses the three major concerns about stimulus validity, generation-pathway confounding, and reliability/power before considering publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful benchmark paper—a 10-dimension linguistically grounded schema, expert annotations on 860 items, and a clean audit of four MOS predictors and four Audio-LLMs across four prompting conditions. That combination is new and worth having. But the headline claim about 'systematic blind spots' is stronger than the evidence supports for the prosodic and paralinguistic dimensions, because the controlled-error stimuli often fail to instantiate the dimensions they target.\n\nWhat's good: the schema is well thought out, with clear operational definitions and real effort to anchor dimensions in linguistic structure (Crystal's distinction, acoustic channel overlap, the annotation rubric). The per-dimension inter-annotator agreement on word-level dimensions is solid (alpha 0.767–0.821), and the null result there—MOS predictors at chance on phonetic accuracy and lexical stress—is likely robust. That's a real finding. The prompt-condition comparison for Audio-LLMs is also useful; output collapse under per-dimension scoring is a practical caution. The paper is transparent about its limitations, which counts for something.\n\nThe soft spots are in the bridge from stimuli to claims. Table 9 shows 32.3% of intended-error samples were majority-labeled clean, with miss rates up to 48.4% for Speaker Identity and 46.2% for Emotional Appropriateness. Low alpha on Intonation (0.469) and Expressiveness (0.460) further undermine ground truth on exactly the dimensions where the 'blind spot' narrative is strongest. The stress-test note is about right: the Sec 3.2 upper-bound assumption—a judge that fails on a deliberate perturbation cannot detect subtler natural deviations—doesn't follow when the perturbations are often not even perceived as errors, and when cross-dimensional contamination (Fig 2) and shared low-level acoustic properties (Praat PSOLA artifacts) are present. The generation pathway confound is real: word-level dimensions were made via IPA/text modification, while Expressiveness, Speech Rate, Speaker Identity, and Human Plausibility were made via audio manipulation. So 'MOS predictors track signal-level artifacts' is partly a statement about how the stimuli were constructed, not just about what the models perceive. The paper half-acknowledges this in Sec 4.4 and the Limitations, but the Conclusion doesn't carry the caveat.\n\nMinor issues: no multiple-testing correction across the many Kendall tau comparisons, Emotional Appropriateness n=46, and the claimed data/code release has no link in the v1 text. None are fatal; they temper the strength of the claims.\n\nThis deserves a serious referee. The core word-level finding holds and the resource is valuable. The authors should be pushed to either soften the 'dimension-level diagnostic' framing or validate that their acoustic manipulations instantiate the intended dimensions—ideally with the expert-purity ratings they partially already have. Recommend acceptance after major revision.","headline":"A valuable new benchmark and audit of TTS evaluators, but its central 'dimension-level blind spots' claim is overstated because the controlled-error stimuli don't cleanly instantiate the dimensions they target.","tokens_in":28025,"tokens_out":2591,"would_cite":true,"duration_ms":23402,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that naturalness in TTS evaluation cannot be a single scalar: using a ten-dimension, linguist-annotated benchmark of 860 utterances, it shows that MOS predictors collapse onto signal-level artifacts while audio-LLM…","keywords":["text-to-speech evaluation","mean opinion score predictors","audio large language models","naturalness","prosody","perceptual dimensions","meta-evaluation benchmark","speech quality assessment"],"falsifier":"A reader could settle this by collecting naturally occurring TTS failures from deployed systems, labeling them with the same ten-dimension rubric, and running the same eight judges: if any judge reliably detects natural word-level or prosodic errors that the benchmark declared undetectable—or if human listeners cannot hear the benchmark's own engineered perturbations—the upper-bound transfer claim would be falsified.","tokens_in":26905,"feed_emoji":"🔊","tokens_out":7546,"duration_ms":61598,"temperature":0.7,"pith_summary":"The paper breaks 'naturalness' into ten listener-perceivable dimensions and asks whether automated TTS evaluators—four MOS predictors and four audio-LLM judges—can detect failures on each one. It builds a benchmark of 860 synthetic utterances, each rated by three trained linguists on all ten dimensions, and finds that MOS predictors concentrate on acoustic signal quality and ignore word-level and most prosodic failures, while audio-LLM judges pick up only a few dimensions depending on the prompt. The authors argue that naturalness cannot be treated as a single scalar construct and that no current automated judge can serve as a dimension-level diagnostic for TTS errors. The upshot is that robust TTS evaluation will require dimension-aware benchmarks rather than holistic scores.","feed_headline":"TTS evaluators miss the errors humans actually hear","feed_subtitle":"A new 10-dimension, linguist-annotated benchmark shows MOS predictors track glitches, not word or prosody errors.","key_machinery":"The central mechanism is the annotation schema itself: a rubric that decomposes naturalness into ten binary-rated dimensions across three linguistic tiers—word level (phonetic accuracy, lexical stress), prosodic level (intonation, prosodic stress, boundary placement, speech rate), and paralinguistic level (emotional appropriateness, expressiveness, speaker identity consistency, human plausibility). Each dimension is targeted by a distinct error-generation pathway: LLM-altered IPA for phoneme and stress errors, LLM-altered plain text for prosodic stress and boundaries, API-enforced emotion tags for emotion mismatches, and Praat-based acoustic manipulation for intonation, rate, expressiveness, identity, and plausibility. This design lets a judge's per-dimension correlation with the linguist majority labels measure detection of that specific failure type. The load-bearing inference is the 'controlled upper-bound' assumption: a judge that cannot detect the deliberately exaggerated perturbation cannot detect subtler naturally occurring deviations of the same type.","core_discovery":"The paper's central claim is that current automated TTS evaluators have systematic but distinct blind spots. Four MOS predictors (UTMOSv2, DNSMOS-Pro, NISQA, Audiobox-Aesthetics) show strong, significant alignment with human labels only on dimensions produced by acoustic signal manipulation—Human Plausibility and Speech Rate in particular—while showing near-chance or negative correlation on Phonetic Accuracy, Lexical Stress, and Prosodic Boundary Placement. Four audio-LLM judges (Gemini 3 Flash, Gemini 3.5 Flash, Qwen3-Omni, Step-Audio-2-mini) under four prompting conditions never generalise across all ten dimensions: attention is selective, prompt-dependent, sometimes negative, and collapses entirely when the model is asked to score all dimensions jointly, with two of four models giving every sample the top score on most dimensions. Prompting with the annotation schema recovers word-level sensitivity in some Gemini models, but no condition yields reliable detection across all of the word, prosodic, and paralinguistic tiers. The authors conclude that naturalness is not a single scalar construct and that robust evaluation requires dimension-aware benchmarks.","pith_inferences":["The high human miss rates in the annotation data—one in three intended errors was not majority-labeled as a failure, up to 48% for speaker identity—mean the benchmark's 'failed to detect' results rest partly on stimuli that are ambiguous; a judge scoring low on a dimension could be failing on invalid stimuli rather than on the dimension itself.","The upper-bound logic cuts both ways: it justifies reading a null result as evidence against subtler natural errors, but it also means the benchmark cannot certify that a judge will detect natural errors unless the perturbations adequately represent those errors, a question that needs real deployment data to settle.","The negative correlations between MOS predictors and several dimensions suggest those models are not merely insensitive but systematically biased, effectively rewarding the artifacts they were trained on; investigating which training artifacts drive this inversion could be a direct follow-up.","Extending the schema to ordinal severity levels, multiple TTS architectures, and naturalistic stimuli from deployed systems would turn the benchmark from an upper-bound audit into a predictive test of real-world evaluator performance."],"forward_implications":["A holistic naturalness score from any current automated judge cannot be trusted as a diagnostic: two failures that sound completely different to humans get lumped together or missed entirely.","TTS evaluations that report only system-level MOS correlation may hide complete blindness to word-level and prosodic errors.","Audio-LLM judges' dimension sensitivity is governed by prompt structure; joint-schema scoring causes output collapse, so evaluation protocols that ask for per-dimension scores must guard against constant-output degeneracy.","Schema-guided prompting can elicit latent word-level capabilities that an underspecified naturalness prompt does not surface, but transcript availability does not consistently help."],"supporting_citations":[{"why":"Supplies EmergentTTS-Eval sentences with complex syntax and questions used as source text, and the model-as-a-judge benchmark the paper extends by adding attribute-level human labels.","marker":"Manku et al., 2025"},{"why":"InstructTTS-Eval is the complementary dimension-category benchmark that lacks human attribute-level ground truth; the paper's design is contrasted against it.","marker":"Huang et al., 2025"},{"why":"Provides the linguistic-versus-paralinguistic distinction that structures the three-tier, ten-dimension annotation schema.","marker":"Crystal, 1969"},{"why":"Defines UTMOSv2, one of the four MOS predictors whose per-dimension sensitivity is audited.","marker":"Baba et al., 2024"},{"why":"Defines DNSMOS-Pro, one of the four MOS predictors evaluated.","marker":"Cumlin et al., 2024"},{"why":"Defines NISQA, one of the four MOS predictors evaluated.","marker":"Mittag et al., 2021"},{"why":"Defines Audiobox-Aesthetics, the fourth MOS predictor evaluated.","marker":"Tjandra et al., 2025"},{"why":"Praat and Parselmouth are the tools used for the F0, duration, and loudness manipulations that generate the acoustic-error stimuli.","marker":"Boersma, 2001; Jadoul et al., 2018"},{"why":"Supplies the inter-annotator agreement measure used to report reliability of the linguist labels.","marker":"Krippendorff, 2011"},{"why":"Cartesia Sonic-3 is the TTS engine through which IPA- and text-modified samples are synthesized, and the single architecture used for all IPA-driven samples.","marker":"Cartesia AI, 2025"}],"fun_headline_variants":["TTS evaluators miss word and prosody errors","MOS and LLM judges fail on linguistic errors","New benchmark exposes TTS evaluator blind spots","10-dimension test shows TTS eval gaps","TTS scores don't match human ears"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole negative audit rests on the assumption that a judge which fails to detect a deliberately exaggerated, obviously engineered error on a dimension would also fail on subtler natural errors of the same type; that transfer holds only if the generated stimuli are valid, representative instances of the dimension, which the paper's own human miss rates of up to 48% on intended errors leave open to question.","fun_headline_variants_meta":{"raw":{"variants":["TTS evaluators miss word and prosody errors","MOS and LLM judges fail on linguistic errors","New benchmark exposes TTS evaluator blind spots","10-dimension test shows TTS eval gaps","TTS scores don't match human ears"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000829,"raw_usage":{"total_tokens":3621,"prompt_tokens":945,"completion_tokens":2676,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":2605}},"tokens_in":561,"tokens_out":2676,"duration_ms":16444,"temperature":1.0,"reasoning_tokens":2605,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:12:33.428237+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could settle this by collecting naturally occurring TTS failures from deployed systems, labeling them with the same ten-dimension rubric, and running the same eight judges: if any judge reliably detects natural word-level or prosodic errors that the benchmark declared undetectable—or if human listeners cannot hear the benchmark's own engineered perturbations—the upper-bound transfer claim would be falsified.","supporting_citations":[],"review_version":1}