{"id":"b63085e4-3e24-420e-8273-2e26e1e63287","arxiv_id":"2508.04179","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Human Fooling Rate (HFR) measures how often listeners mistake TTS output for human speech; a large-scale evaluation shows commercial zero-shot TTS approaches human deception while open-source models lag.","lead":"This paper introduces a new metric, Human Fooling Rate, that measures how often people mistake computer-generated speech for human speech. It reports that commercial voice models can fool listeners in zero-shot settings, while open-source models still fall short, especially in conversational speech.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HFR's human-baseline construction and listener protocol are undisclosed; the headline model-vs-human comparisons may be artifacts of low-bar human references.","rationale":"The reader's weakest_assumption correctly identifies that the HFR findings' validity depends on the listener panel, speech samples, and experimental protocol. My stress-test agrees: this is the single most load-bearing concern because every headline claim (i)-(iv) is a comparative statement about HFR values, and HFR is not a physical quantity but a protocol-dependent measurement. The abstract explicitly warns that low-quality human references set a low bar for TTS systems, yet provides no evidence that the human reference clips used in this study are high-quality or that listener instructions are neutral. This is not an accusation of misconduct; it is a request for evidence that the paper itself suggests is necessary. Because the full text is unavailable, the correct verdict remains UNVERDICTED, and my concern does not change that. The concrete test I propose would settle the concern by directly measuring protocol sensitivity; if the test passes, the headline claims would be substantially strengthened. I found no need to adjust the reader's verdict.","tokens_in":678,"tokens_out":2047,"duration_ms":26406,"concrete_test":"The authors should release the full HFR protocol and recompute the reported HFR comparisons under two conditions: (A) the original human reference clips, and (B) a new set of human clips deliberately selected to be high-expressive and prosodically varied. If the rank order of model HFRs or the machine-vs-human gap changes by more than the reported confidence interval (e.g., >5 percentage points), the headline claims are protocol-dependent. A second independent check: run the same listener evaluation with two different instruction phrasings—'Is this a real human recording?' versus 'Does this sound natural?'—and compare HFR values; if they diverge significantly, the metric conflates human-likeness with naturalness.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that CMOS-based parity claims fail under HFR and that commercial models approach human deception—depends entirely on the HFR values assigned to human reference clips. The abstract does not state how those clips were selected, what instructions listeners received, or how many listeners participated. If the human reference set is monotonous or low-expressive, the paper's own warning about 'low bar' baselines applies to its own evaluation: machine HFR would be artificially inflated, making 'approach human' look more impressive than it is. Conversely, unusually expressive human references would suppress machine HFR and overstate the gap. The four findings (i)-(iv) all rely on the HFR scale being stable across models and human references; without the protocol and a stability analysis, the specific numeric claims are unverifiable. This is not an internal inconsistency but a load-bearing missing support: the paper flags low-bar baselines as a pitfall for other datasets, and the same pitfall could undermine its own comparisons. The abstract-only format prevents checking this directly.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Human Fooling Rate (HFR), a deception-based metric for TTS, and reports a 'large-scale evaluation' of open-source and commercial models. Four findings are claimed: (i) CMOS-based claims of human parity often fail under HFR; (ii) TTS benchmarks should use datasets where human speech achieves high HFR; (iii) commercial models approach human deception in zero-shot settings while open-source systems struggle with conversational speech; (iv) fine-tuning on high-quality data improves realism but does not fully close the gap. The abstract asserts that HFR is a more realistic human-centric complement to existing subjective tests. No methodological details (listener counts, sample selection, statistical testing) are provided, and the full text is not available for this review.","tokens_in":913,"tokens_out":3831,"duration_ms":43249,"significance":"If the findings hold, HFR would offer a valuable complementary metric for TTS evaluation, and the reported disconnect between CMOS scores and deception rates would be an important caution for the field. The paper's call to calibrate benchmarks by human HFR is a reasonable proposal. However, significance cannot be evaluated from the abstract alone: the metric's operational definition, the scale of the evaluation, and the supporting statistical evidence are all unspecified. As it stands, the contribution is a proposal plus a set of unreviewable empirical claims.","major_comments":[{"comment":"The core metric is not defined precisely. 'How often machine-generated speech is mistaken for human' leaves open the experimental paradigm (e.g., single-interval yes/no vs. two-alternative forced choice), the number and provenance of listeners, the acoustic conditions, and the scoring rule. Without this, none of the quantitative comparisons (i)-(iv) can be reproduced or assessed. The authors must provide the full protocol, including sample sizes and listener demographics, before the claims are testable.","section":"Abstract (HFR operationalization)"},{"comment":"The paper itself warns in finding (ii) that low-quality reference samples lower the bar, yet the abstract does not report the HFR for the human reference clips used in the model comparisons. If humans are also frequently misclassified as machines, then 'approaching human' carries little meaning. This is a load-bearing missing support: the authors must report the human-reference HFR and show that it is high enough to make the comparison meaningful.","section":"Abstract (human baseline)"},{"comment":"Claim (i) states that 'CMOS-based claims of human parity often fail under deception testing.' The abstract does not say which CMOS evaluations were re-examined, whether the scores were recomputed on the same stimuli, or how the deception test relates to the original CMOS test. Without a side-by-side description, this claim is ambiguous and cannot be evaluated.","section":"Abstract (CMOS comparison)"},{"comment":"No confidence intervals, error bars, or significance tests are mentioned. Listener-based metrics are noisy; 'approach human deception' and 'still struggle' are relative statements that need quantitative uncertainty quantification. The paper must report inter-listener agreement and the precision of HFR estimates to support its rank ordering of models.","section":"Abstract (statistical analysis)"}],"minor_comments":[{"comment":"The phrase 'Turing-like evaluation' is vague; specify whether this is a true Turing test or a more limited listening test.","section":"Abstract"},{"comment":"The term 'human deception' is used both for the task (machines fooling humans) and the human baseline (humans fooling other humans). This dual use is confusing; consider 'human-reference HFR' for the latter.","section":"Abstract"},{"comment":"The contribution would be strengthened by stating whether the evaluation uses a pre-registered protocol or includes publicly available code/data.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"Given the abstract-only submission, I cannot assess the manuscript's validity. The methodological omissions are substantial. I advise the editor to obtain the full text before any decision. If the full paper does not supply the protocol details, statistical analyses, and human-baseline HFR, it would require major revision at best."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about 2508.04179. I can only see the abstract, so anything I say is provisional. The core idea is straightforward and worth taking seriously: instead of asking raters which sample they prefer, measure how often machines fool listeners into thinking they are human. That's a meaningful complement to CMOS, and the abstract's point (ii) is the sharpest one — deception rates are only interpretable relative to how deceptive human references are. That is exactly the right methodological instinct, and it's a genuine contribution to the TTS evaluation debate.\n\nWhat I can't verify from the abstract is the actual experiment. No listener counts, no sample selection protocol, no human baseline numbers, no statistical analysis. The stress-test concern is real: if the human reference clips are easy to identify as human, machine HFR is inflated, and the 'commercial models approach human deception' finding could be an artifact. But the authors flag the same pitfall themselves, so they may well have handled it. I'm not going to manufacture a flaw out of an abstract that is honestly silent on details.\n\nThe other findings have a similar shape: plausible but unverifiable. Fine-tuning helps but doesn't close the gap; open-source lags on conversational naturalness. These are empirically sharp, testable claims. If the full text ships the protocol and data, this is a useful paper for the TTS community and for anyone designing human-centric eval metrics.\n\nMy recommendation: send it to peer review. Not because the abstract proves a thing, but because the research question is timely and the authors have clearly thought about baseline pitfalls. A serious referee who sees the full protocol can tell whether the central claim holds. I'd bring it to reading group only if the full text is available.","headline":"HFR is a plausible and timely metric, but this abstract alone can't support the load-bearing claims; the authors' own low-bar warning is a good sign.","tokens_in":1360,"tokens_out":1489,"would_cite":false,"duration_ms":17564,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Human-parity TTS claims fail a Turing-style deception test.","keywords":["TTS","Human Fooling Rate","deception test","CMOS","zero-shot TTS","open-source TTS","speech evaluation","Turing test"],"falsifier":"Run the same HFR protocol with a different listener panel and more expressive human reference clips. If human HFR does not, on those clips, clearly exceed machine HFR, or if the ordering of commercial versus open-source models flips, the paper's central comparison would be called into question.","tokens_in":629,"feed_emoji":"🗣️","tokens_out":3501,"duration_ms":37461,"temperature":0.7,"pith_summary":"Most TTS evaluation relies on subjective CMOS comparisons, and recent scores suggest machines have reached human parity. This paper argues those parity claims do not survive a Turing-like deception test in which listeners must decide whether a clip is human or machine. It introduces Human Fooling Rate (HFR), the fraction of machine-generated clips that listeners mistake for human speech, and runs a large-scale comparison of open-source and commercial systems. The finding: commercial zero-shot voices approach human deception rates, open-source systems still lag on conversational speech, and fine-tuning on high-quality data narrows but does not close the gap. The paper's central proposal is that progress should be benchmarked on datasets where real human speech itself scores high HFR, so the bar is not artificially low.","feed_headline":"Human-parity TTS claims fail a Turing-style deception test","feed_subtitle":"New Human Fooling Rate metric shows commercial voice clones near human, open-source still lags.","key_machinery":"Human Fooling Rate (HFR): the rate at which machine-generated speech is mistaken for human speech by listeners in a Turing-like test. It is the central metric of the paper, used to benchmark TTS models and to argue that existing CMOS-based parity claims are overstated.","core_discovery":"The paper claims that CMOS-based claims of human parity often fail under deception testing. By measuring HFR across models, it finds that commercial models approach human deception in zero-shot settings, while open-source systems still struggle with natural conversational speech, and that fine-tuning on high-quality data improves realism but does not fully bridge the gap. The paper also argues that evaluations should be run on datasets where human speech achieves high HFR, since monotonous references set a low bar.","pith_inferences":["A likely consequence is that TTS vendors will optimize specifically for deception, which could change what 'quality' means in the field; the paper does not discuss this incentive.","Since the paper's protocol details are not described, the ranking of models may depend on listener demographics; testing HFR with different listener panels would show whether the results are stable.","HFR could be applied to other audio domains such as deepfake detection, where the same metric could serve as an adversarial measure.","The paper's critique of CMOS-based parity suggests older published claims of human parity may need re-testing under deception protocols."],"forward_implications":["If HFR becomes a standard companion to CMOS, published claims of human parity will need to be re-validated under deception protocols.","Commercial zero-shot TTS already fool listeners at rates near human, so realistic synthetic speech is available in practice.","Open-source TTS has not reached the same conversational realism, pointing to a specific gap for that community.","Fine-tuning on high-quality data is a partial fix, so data alone is not enough to close the human-machine gap.","Benchmark design matters: low-expressiveness reference speech sets an easy bar, so future TTS evaluations should use expressive human references."],"supporting_citations":[],"fun_headline_variants":["TTS fooling rates: commercial near human, open-source lags","Voice clones win Turing-style test, but open-source lags","Human parity in TTS? Not so fast, new metric shows","New test: commercial TTS fools humans, open-source doesn't","CMOS parity claims fail deception testing in TTS"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The HFR results assume the listener panel, the speech samples, the instructions, and the selected human reference clips produce stable, representative deception rates; in particular, if the human reference clips are not expressive, the human HFR baseline is low and machines look better than they are.","fun_headline_variants_meta":{"raw":{"variants":["TTS fooling rates: commercial near human, open-source lags","Voice clones win Turing-style test, but open-source lags","Human parity in TTS? Not so fast, new metric shows","New test: commercial TTS fools humans, open-source doesn't","CMOS parity claims fail deception testing in TTS"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000139,"raw_usage":{"total_tokens":937,"prompt_tokens":630,"completion_tokens":307,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":374,"completion_tokens_details":{"reasoning_tokens":219}},"tokens_in":374,"tokens_out":307,"duration_ms":3703,"temperature":1.0,"reasoning_tokens":219,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:47:50.072033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same HFR protocol with a different listener panel and more expressive human reference clips. If human HFR does not, on those clips, clearly exceed machine HFR, or if the ordering of commercial versus open-source models flips, the paper's central comparison would be called into question.","supporting_citations":[],"review_version":1}