{"id":"41218e77-0efd-4b3f-b40f-51e68962c40f","arxiv_id":"2608.02235","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"No single Urdu TTS system wins across all metrics, and the system listeners liked most, Google Gemini, was the furthest from reference audio on objective acoustic scores.","lead":"This study tested four Urdu text-to-speech systems on four speaking styles and found that the friendliest-sounding system is not the most faithful to the original recording, and that emotional speech is the hardest for all of them. It's a reusable evaluation blueprint, with code released, for checking low-resource language voices before deployment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-speaker objective references confound both system and domain rankings; normalization test needed","rationale":"The reader's weakest assumption identifies the same load-bearing concern: objective metrics (MCD, F0 RMSE) are computed against cross-speaker references without normalization, so system rankings and possibly domain rankings conflate synthesis fidelity with accidental speaker/register match. This is indeed the most serious threat to the central claim. The paper's own Section 5.3.3 concedes the register-offset issue, and Section 5.3.2 does not even acknowledge the analogous cross-speaker confound for MCD. Since the headline finding includes a claimed divergence between subjective preference (Gemini wins MUSHRA) and objective fidelity (Edge wins MCD/F0 RMSE), the validity of that divergence is load-bearing. If the objective rankings are driven by speaker match, the 'no single system dominates' conclusion becomes far weaker and the interpretive framework in Section 7.6.2 loses its empirical foundation. The domain-level claim about emotional speech being hardest also depends on cross-domain reference comparability, which is threatened by dataset differences (studio read speech versus drama recordings). These concerns are concrete and fixable, but they are not merely cosmetic; they bear directly on the interpretation of the benchmark. I therefore agree with the reader that the paper should be CONDITIONAL pending normalization or explicit modeling of speaker/register offsets, and recommend adding a natural-natural baseline to calibrate the metric values. The paper's qualitative conclusions may well survive such a test, but they are not yet established.","tokens_in":24109,"tokens_out":6002,"duration_ms":56714,"concrete_test":"Recompute MCD (Table 10) and F0 RMSE (Table 12) after per-reference-speaker normalization: subtract each reference utterance's mean log-F0 from both reference and synthetic F0 contours (removing the register offset), and apply cepstral-mean subtraction to the mel-cepstra before DTW/MCD calculation (a crude spectral speaker-normalization). If Edge TTS's overall objective lead over Gemini TTS shrinks or reverses, the claimed subjective–objective split is an artifact of cross-speaker speaker/register match rather than synthesis fidelity. Additionally, compute a natural–natural cross-speaker baseline per domain by running the same MCD/F0 pipeline on pairs of reference utterances from different speakers; if the emotional domain's baseline already approaches 12 dB / 888 cents, then 'emotional speech is hardest' reflects reference variability, not synthesis difficulty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'no single system dominates because Edge TTS wins the objective acoustic metrics' and that 'emotional speech is the hardest domain' rests on MCD and F0 RMSE values computed between each synthetic utterance and a natural reference from a different speaker (Section 5.3.2, 5.3.3). Section 5.3.3 explicitly acknowledges that F0 RMSE 'reflects both pitch-contour deviation and a speaker-dependent register offset,' yet no normalization is applied before Tables 10 and 12 are used for system- and domain-level comparisons. The same cross-speaker confound applies to MCD: mel-cepstral coefficients encode vocal-tract and spectral characteristics of the particular reference speaker, so a system whose fixed voice happens to match the reference speakers' average spectral envelope or pitch register will appear artificially better. Moreover, domain-level comparisons are additionally confounded by the use of different source datasets per domain (FLEURS and UrduSpeech for Formal, UrduSpeech for Conversational/Literary, UrSEC drama recordings for Emotional), which likely differ in recording conditions, microphone quality, and acoustic variability. If the emotional references are noisier or more variable, all systems will show inflated MCD/F0 RMSE regardless of synthesis quality. Thus the objective strand of the central claim—both the system-level subjective/objective divergence and the domain-level difficulty ranking—is not secure until the speaker/register offset is modeled or controlled.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a domain-stratified, multi-metric benchmarking framework for evaluating TTS systems in a low-resource language, demonstrated on Urdu. Four systems (MMS-TTS, Indic-Parler-TTS, Microsoft Edge TTS, Google Gemini TTS) are compared across Formal, Conversational, Literary/Storytelling, and Emotional domains using MUSHRA-style listening tests, ABX discrimination, Resemblyzer speaker similarity, MCD, and F0 RMSE over 960 utterance pairs. The main empirical claims are that emotional speech is consistently the most difficult synthesis domain and that no single system dominates: Gemini TTS wins listener preference, Edge TTS wins the objective acoustic metrics, and Indic-Parler-TTS often leads on speaker similarity. Evaluation scripts, tables, and Colab notebooks are publicly released.","tokens_in":24454,"tokens_out":6665,"duration_ms":57633,"significance":"If the findings are secure, this is a useful contribution: it is one of the first domain-stratified TTS benchmarks for a South Asian low-resource language, and the explicit discussion of subjective-objective divergence is a valuable caution against single-metric leaderboards. The public release of scripts and result tables is a genuine strength, as is the paper's transparency about its own methodological limitations. However, the load-bearing objective comparisons are computed against cross-speaker natural references, and the ABX results contain internal inconsistencies. The central domain-difficulty and system-ranking claims therefore need additional analysis before the empirical conclusions can be accepted.","major_comments":[{"comment":"MCD and F0 RMSE compare every synthetic utterance to a natural reference from a different speaker. Eq. (1) uses mel-cepstral coefficients that carry speaker-specific spectral-envelope information, and Eq. (2) includes a speaker-dependent register offset in log-F0. §5.3.3 explicitly acknowledges the F0 offset, but no normalization or correction is applied before Tables 10 and 12 are used to conclude that Edge TTS is the best objective system and that Emotional speech is the hardest domain. The domain rankings are additionally confounded because the four domains come from three different corpora with different speaker pools and recording conditions (§4.2–4.4). The emotional-difficulty claim is well supported by MUSHRA and Resemblyzer, but the MCD/F0 RMSE leg needs same-speaker references or explicit normalization (for example, per-speaker mean log-F0 removal, cepstral-mean subtraction, or","section":"§5.3.2–5.3.3, Tables 10/12"},{"comment":"The subjective test is labeled MUSHRA-style but deviates from ITU-R BS.1534 in fundamental ways: the hidden reference and anchor are not rated, the response scale is discrete 1–5 rather than continuous 0–100, and only 10 listeners per domain are used. The paper acknowledges these deviations, yet the abstract and Section 7.6 continue to treat the resulting scores as a MUSHRA paradigm and count them as one of the five converging measures. The findings would be better presented as reference-anchored MOS-style preference ratings, with the convergence argument restated accordingly. This is not merely a naming issue: without a rated hidden reference, there is no check on whether listeners actually used the reference consistently.","section":"§5.2.1, §7.6"},{"comment":"The ABX results are internally inconsistent. Table 5 reports only Literary/Storytelling and Formal domains, with n=54 trials per domain and 9 listeners, but §5.2.2 says the ABX panel was the same as the MUSHRA panel (10 listeners per domain pair). Section 8.3 instead cites 87.5–88.9% discrimination for \"Emotional and Literary/Storytelling\" trials, and Section 7.2 says accuracy was 90.7% in both domains tested. These numbers cannot all be correct. The ABX conclusion that top systems remain distinguishable from natural speech should be repaired with the actual domain counts and listener numbers, and the conflicting statements reconciled.","section":"§7.2, Table 5, §8.3"}],"minor_comments":[{"comment":"Section 3.2 states that \"MMS-TTS is the only fully open-weight, locally executable system in this evaluation,\" but Indic-Parler-TTS is also open-weight and was executed locally (Section 6.1). This is a factual inconsistency that should be corrected.","section":"§3.2, Table 1"},{"comment":"Typo: \"Participants were were native Urdu speakers\" should be \"Participants were native Urdu speakers.\"","section":"§5.2.1"},{"comment":"The phrase \"listeners correctly identified the reference-matching system\" is confusing. In an ABX task, listeners judge whether X is closer to A or B; they do not identify a system. Please rephrase to describe the discrimination rate directly.","section":"§7.2"},{"comment":"The note says SDs are \"in parentheses, per domain-system cell,\" but the table does not display parentheses. Format the table so the note matches the presentation.","section":"Table 4"},{"comment":"The mean Emotional-domain F0 RMSE is reported as 889 cents in the abstract and 888.54 cents in Section 7.5.1. Use a consistent rounding convention.","section":"§7.5.1 vs. Abstract"},{"comment":"The Formal (FLEURS) row reports identical F0 RMSE values (470.66) for Indic-Parler-TTS and MMS-TTS. If this is a genuine tie, state it explicitly; if it is a copy/paste artifact, correct it.","section":"Table 12"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and represents useful work, but the cross-speaker objective-reference problem is central rather than peripheral. I would suggest the editor ask the authors to add a normalization or sensitivity analysis for MCD and F0 RMSE, and to fix the ABX data inconsistencies, before considering acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a real empirical benchmark—four TTS systems, four Urdu domains, 960 pairs, five metrics, with scripts and notebooks released. That is genuinely new for Urdu and useful as a template for low-resource TTS evaluation. The paper's main qualitative findings—emotional speech is the hardest domain, and no one system wins across all metrics—are plausible and mostly well-documented. The subjective data alone support the emotional-hardest claim.\n\nThe soft spot is the one the authors half-admit in Section 5.3.3: the MCD and F0 RMSE compare every synthetic clip to a natural reference from a different male speaker. That is not a minor detail. Different speakers have different average vocal-tract spectra and pitch registers, so a system whose fixed voice happens to sit near the reference speakers will look artificially better, and a system like Gemini that listeners prefer may look artificially worse. The paper's central system-level story—Gemini wins MUSHRA, Edge wins acoustic metrics—may be partly a story about voice choice rather than synthesis quality. The same confound affects domain ranking: emotional references come from UrSEC drama recordings, not the same conditions as FLEURS/UrduSpeech, so inflated MCD/F0 RMSE there may reflect reference variability as much as synthesis difficulty. The paper notes the register offset but never corrects it, and the 'subjective-objective inversion' interpretation is used as evidence for a naturalness/fidelity trade-off. I would not treat that inversion as established until the comparison is rerun with matched speakers or an explicit speaker-offset model.\n\nOther issues are minor: the subjective protocol is MUSHRA-inspired but uses a 1–5 discrete scale with unrated anchor, 10 listeners per domain pair, no inferential statistics. That is documented and appropriately softened, but the absence of significance tests limits confidence. The ABX section has a small inconsistency (9 vs 10 listeners) and one typo where MMS is named in place of Edge.\n\nAlso worth flagging: an in-text citation to 'Anonymous, 2025' for a Pashto benchmark looks like a leftover from a double-blind submission. That should be cleaned up before publication.\n\nOverall: solid, honest engineering, a useful resource for anyone building low-resource TTS benchmarks, and the framework is reproducible. The central empirical claims need a same-speaker or offset-corrected analysis before they carry weight. I would send it to peer review, expecting revision. Worth citing if you work on TTS evaluation, but I'd cite the framework and dataset, not the numerical comparisons.","headline":"Useful first Urdu multi-domain TTS benchmark, but the headline subjective-objective inversion may be an artifact of comparing every system to a different-speaker reference.","tokens_in":24942,"tokens_out":3429,"would_cite":true,"duration_ms":30944,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Emotional Urdu speech is the hardest test for text-to-speech, and no single system wins on all measures.","keywords":["Urdu TTS","text-to-speech evaluation","emotional speech synthesis","low-resource languages","comparative listening test","speaker similarity","mel-cepstral distortion","F0 RMSE"],"falsifier":"Synthesize the same Urdu utterances using each system with a voice cloned from the exact reference speaker (or, minimally, subtract each reference speaker's mean F0 offset before computing F0 RMSE), then re-rank systems and domains. If the Emotional domain no longer shows the largest errors and the Gemini-vs-Edge inversion disappears, the paper's central claims rest on speaker mismatch rather than synthesis difficulty.","tokens_in":23997,"feed_emoji":"🎙️","tokens_out":4580,"duration_ms":38048,"temperature":0.7,"pith_summary":"The paper sets out to measure Urdu text-to-speech systems the way they would actually be used: across formal, conversational, literary, and emotional speech, and with perceptual, acoustic, and speaker-identity tests at once. It finds that emotional speech is the hardest condition by every measure, and that listener preference and objective reference fidelity point to different winners. The result matters because single-metric, single-domain evaluations—the common practice—would mislead anyone choosing a TTS system for a real application. The paper also releases its scripts and result tables so the benchmark can be rerun and extended.","feed_headline":"Emotional speech defeats every Urdu TTS system","feed_subtitle":"All four synthesis systems fail hardest on expressive speech; the winner depends on the metric.","key_machinery":"The framework is a domain-stratified, multi-metric pipeline: 60 utterances per domain, each synthesized by four systems, then scored by (i) a comparative listening test with a hidden reference and an anchor, (ii) a forced-choice discrimination test, (iii) speaker-embedding cosine similarity, (iv) mel-cepstral distortion after dynamic time warping, and (v) fundamental-frequency RMSE in cents. The analytic core is the cross-metric comparison: where the five measures agree (the Emotional domain is hardest) and where they diverge (Gemini preferred but deviant on acoustic fidelity), the divergences are used to identify what each metric actually captures.","core_discovery":"Across 960 reference-synthetic pairs, four systems, four domains, and five evaluation methods, the paper shows convergent evidence that emotional speech is the bottleneck for Urdu synthesis: lowest listening scores for every system, highest spectral distortion (mean MCD 12.03 dB), highest pitch error (mean F0 RMSE 888.54 cents), and lowest speaker-embedding similarity (0.5437). At the system level, no winner emerges: Google Gemini TTS receives the best listener ratings overall while Edge TTS is closest to the reference on objective acoustic metrics; the two metrics disagree so strongly that the most preferred system is the worst reference-matcher on several objective measures. The paper inte","pith_inferences":["A direct test of the paper's 'reference fidelity vs perceived quality' claim would be to re-run the acoustic metrics with F0 normalized per speaker (removing register offset) and with synthetic voices cloned to match the reference speaker; if the inversion persists, it is about prosody, not speaker mismatch.","The framework could be transferred to other low-resource languages with available single-speaker corpora; the domain labels (formal/conversational/emotional/literary) are generic and the scripts are released.","The listening test used a 1-5 scale and 10 listeners per domain, so the subjective ranking may compress variance; a 0-100 continuous scale with more listeners could sharpen or shift the Gemini-Edge gap.","The boredom result (lowest speaker similarity despite low arousal) suggests emotion-specific embedding instability; a follow-up with per-emotion prosody analysis could separate synthesis error from reference variability."],"forward_implications":["Emotional speech should be a required test condition in future TTS evaluations; it exposes failures that neutral read speech hides.","Single-number leaderboards for Urdu TTS are an artifact of metric choice; deployment decisions need a per-domain, per-metric profile.","A system can be rated as natural yet be reliably detected as synthetic (listeners identified the synthetic sample in 90.7% of ABX trials), so quality and authenticity should be tracked separately.","Systems optimized to minimize reference distance may not maximize listener preference, and vice versa; the two objectives need separate optimization targets.","Domain labels are not acoustically homogeneous: the Formal domain's two source subsets produced system rankings as different as across-domain differences, so subdomain analysis is needed."],"fun_headline_variants":["Emotional Urdu speech stumps all four TTS systems","Which Urdu TTS wins? Depends on the metric","No best Urdu TTS: metrics clash with listeners"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The objective acoustic comparisons assume that differences between synthetic speech and a natural reference are due to synthesis quality, but each reference comes from a different male speaker and no speaker-register normalization is applied, so part of the measured MCD and F0 RMSE reflects speaker mismatch rather than synthesis error.","fun_headline_variants_meta":{"raw":{"variants":["Emotional Urdu speech stumps all four TTS systems","Which Urdu TTS wins? Depends on the metric","No best Urdu TTS: metrics clash with listeners"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000345,"raw_usage":{"total_tokens":1754,"prompt_tokens":793,"completion_tokens":961,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":917}},"tokens_in":537,"tokens_out":961,"duration_ms":7141,"temperature":1.0,"reasoning_tokens":917,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T10:55:33.921151+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Synthesize the same Urdu utterances using each system with a voice cloned from the exact reference speaker (or, minimally, subtract each reference speaker's mean F0 offset before computing F0 RMSE), then re-rank systems and domains. If the Emotional domain no longer shows the largest errors and the Gemini-vs-Edge inversion disappears, the paper's central claims rest on speaker mismatch rather than synthesis difficulty.","supporting_citations":[],"review_version":1}