Pith. sign in

REVIEW 4 major objections 6 minor 33 references

Voice AI quality cannot be captured by one number: performance is dimension-specific and should be reported as a profile.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 00:53 UTC pith:XLP2AMWV

load-bearing objection A large, well-resourced benchmark with a plausible but not fully proven profile-based evaluation claim; worth reviewing, but needs data release and statistical tightening. the 4 major comments →

arxiv 2607.14846 v1 pith:XLP2AMWV submitted 2026-07-16 cs.SD cs.AI

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

classification cs.SD cs.AI
keywords voice AI benchmarkparalinguistic evaluationtext-to-speechspeech-to-speechspeech understandingASR robustnessdimension-specific performancebenchmark optimization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces RW-Voice-EQ, a benchmark that measures voice AI systems on four domains: generating speech, conversing through speech, understanding paralinguistic cues in audio, and transcribing under real-world acoustic variation. The core claim is that performance across these capabilities is largely independent: a system can be natural but not expressive, expressive but not identity-stable, competent on transcripts but deaf to vocal tone, or strong on clean speech but brittle under accent, emotion, noise, and conversation. The authors support this with roughly 830,000 human ratings plus objective scoring, and with ablations that compare audio-input against transcript-only input. If correct, the field should adopt per-dimension profiles with withheld evaluation sets rather than one-number leaderboards.

Core claim

RW-Voice-EQ evaluates text-to-speech along seven dimensions (acting/role-fit, expressiveness, voice identity, language stability, reliability, long-form stability, acoustic quality), speech-to-speech along five (emotion understanding, emotion alignment, expressivity robustness, voice naturalness under stress, problem redirection), speech understanding along three (emotion recognition, speaker verification, synthetic-speech detection), and ASR across four condition sets (accent, emotion, background audio, conversation). The central finding is that systems rank differently in every dimension: no system leads all dimensions, and even within a single system, task behavior and vocal quality can d

What carries the argument

The central mechanism is the dimension-factor scoring pipeline: individual evaluations are grouped into latent factors, each evaluation's provider-level scores are Spearman-correlated against every factor leaderboard, and groupings are retained or regrouped when the best-fit correlation falls below 0.30, with manual verification. For speech-to-speech, the load-bearing design is the audio-only versus transcript-only ablation, which isolates whether access to the acoustic signal changes agent behavior. For ASR, the benchmark uses four human-curated private datasets with consensus human-verified reference transcripts. Speech-language-model judges are treated as auxiliary evaluators, validated a

Load-bearing premise

The independence conclusion rests on the assumption that each evaluation dimension measures the construct it names, and that the observed dimension separability is not an artifact of the grouping procedure, since the factor validation reuses the same data that generate the dimension scores.

What would settle it

A re-run of the factor grouping on an independent item set that does not reproduce the same dimension structure (evaluations' best-fit Spearman correlations shift across factor leaderboards or fall below the 0.30 threshold) would undermine the claim that performance is dimension-specific. Similarly, if the audio-only versus transcript-only contrasts in speech-to-speech do not survive paired significance testing with larger samples, the 'transcript-driven' conclusion fails.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Single-number leaderboards should be replaced by per-dimension profiles; model selection becomes use-case dependent on which capability matters most for deployment.
  • Clean-speech ASR benchmarks are insufficient for production decisions; systems should be reported with condition-level robustness profiles covering accent, emotion, background audio, and conversation.
  • Access to audio in speech-to-speech agents does not guarantee use of vocal affect; evaluation should measure the audio-versus-transcript difference explicitly.
  • Speech-language-model judges are verification-dependent: they agree well with humans on target-given correctness tasks but degrade on open-ended perceptual judgments such as voice identity and acting role-fit.
  • Withheld, private evaluation sets are necessary because public benchmark optimization is already measurable in state-of-the-art ASR systems.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If dimension-specific profiles become standard, downstream applications will select systems per deployment context (e.g., long-form narration vs. identity-critical assistive voice), and aggregate rankings will lose predictive authority entirely.
  • The benchmark's diagnostic probes (masked-word tracking, orthographic-switch detection, synthetic-voice matching) could be generalized into a standard contamination audit for any speech leaderboard.
  • The verification-dependence finding for speech-language-model judges suggests a routing rule: use SLMs for unambiguous, answer-graded tasks and reserve human raters for open-ended perceptual constructs, rather than treating SLM scores as uniformly valid.
  • The consistent gap between positive high-arousal speech and negative high-arousal speech in ASR points to a concrete training-data target: adding expressive positive speech should reduce that gap if the paper's underrepresentation explanation is correct.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces RW-Voice-EQ Bench, a large multidimensional benchmark for voice AI covering TTS, speech-to-speech (STS), speech understanding (SU), and ASR robustness. TTS and STS are scored by human raters on task-specific rubrics (over 830k ratings), while SU and ASR are scored against human-labeled references, transcripts, and pairwise targets. The benchmark includes 31 TTS configurations, 15 STS configurations, 18–19 SU systems, and about 40 ASR models, plus an analysis of speech-language-model judges and evidence of ASR benchmark optimization. The central claim is that voice-AI performance is dimension-specific: strong performance in one capability does not predict strong performance in another, so systems should be reported as capability profiles rather than single aggregate scores. The paper also reports that some STS agents remain largely transcript-driven, that relative emotion comparison is much easier than absolute emotion identification, that dedicated speaker-verification models outperform general audio-language judges, and that ASR robustness failures are not captured by clean-speech benchmarks.

Significance. If the central claim holds, the benchmark would support a meaningful shift in how voice AI is evaluated, away from single-number leaderboards and toward per-dimension diagnostic profiles. The paper has genuine strengths: the evaluation is unusually broad, with over one million human ratings collected from an external rater pool; SU and ASR results are scored against human-verified labels and transcripts rather than model outputs; the benchmark and leaderboards are publicly released; and the analysis of SLM-judge agreement is a useful methodological contribution. The ASR robustness dataset with human-corrected references is a valuable resource. The main unresolved issue is statistical: the dimension-separability conclusion is asserted from descriptive top-five patterns and a factor-grouping procedure that is validated on the same data it defines. The paper itself acknowledges the lack of paired statistical testing and the non-inferential nature of reported standard deviations, yet the abstract and discussion draw strong dimension-specific conclusions from those same results.

major comments (4)
  1. [§2.2, factor grouping and dimension validation] The dimensionality claim is validated circularly. Each eval's provider vector is Spearman-correlated against every factor leaderboard constructed from those same evals, and evals are assigned to their best-fitting factor if the correlation exceeds 0.30. With 31 TTS providers and seven candidate factors, this procedure can create apparent separation by chance. The paper does not report between-factor correlations, confidence intervals for factor assignments, or a stability analysis such as drop-one-eval or split-half replication. Since the abstract's 'largely independent evaluation dimensions' is load-bearing for the profile-vs-aggregate conclusion, please add cross-validated or out-of-sample factor replication and report the between-factor correlation matrix.
  2. [§5.2, audio-only vs transcript-only ablation] The STS conclusion 'access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven' rests on descriptive direction counts: four systems scored higher under AO, one unchanged, ten lower on the naturalistic set, and ten higher on the scripted set. No significance tests, confidence intervals, or effect sizes are reported. The Figure 6 caption explicitly states that the displayed standard deviations 'are not estimates of rater-level uncertainty or confidence intervals,' and §8 defers 'paired statistical testing' to future work. Please add per-system paired tests or bootstrap CIs across scenarios, and report the AO–TO difference with uncertainty for each of the 15 systems, especially for GPT-Realtime-2's +0.17 difference.
  3. [§4.2, Figure 5 and TTS dimension independence] The claim that naturalness, expressiveness, identity stability, and reliability are 'largely independent evaluation dimensions' is supported only by top-five membership patterns. No correlation matrix among the seven dimension scores is provided, no null model for expected top-five overlap is given, and there are no confidence intervals around dimension means. With only 31 systems and top-five truncation, the observed absence of a system in all seven top-five lists is weak evidence of independence. Please report the full inter-dimension Spearman correlation matrix over all 31 providers, with uncertainty, and test whether the observed overlap differs from what chance would produce.
  4. [§7.2, ASR clean-benchmark comparison] The claim that 'real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks' requires a direct quantitative comparison between the evaluated models' clean-benchmark WER and their per-condition WER. The paper shows that rankings reorder across the four tracks, but it does not report a correlation between clean leaderboard WER and the robustness WERs for the same models. Please add a scatterplot or correlation table (e.g., LibriSpeech or VoxPopuli WER vs accent/emotion/noise/conversational WER) to substantiate the claim that clean performance is not predictive.
minor comments (6)
  1. [Abstract] Typo: 'Real World Voice EQ Bench a' should be 'Real World Voice EQ Bench'.
  2. [§6.1 / Table 6] The text says 18 speech-understanding systems were evaluated, but Table 6 lists 19 rows (including gemma-3n). Please reconcile the count and the table.
  3. [§7.1 / Table 7] The number of ASR systems is inconsistent: §7.1 says 40, §7.2 says 41, and Table 7 appears to contain 42 rows. Please correct the counts consistently.
  4. [Table 1, Acoustic Quality row] The Acoustic Quality row has no prompt/generation/rating counts, and the text says fidelity ratings were collected 'within the expression evaluations.' Please state explicitly how the Acoustic Quality dimension score in Figure 5 was computed and from which evaluations.
  5. [§4.1 vs §2.1] The TTS dimension is called 'Multilingual Code-Switching' in §4.1 and Table 1's header, but 'Language Stability' elsewhere. Use one consistent name.
  6. [Figure 3] The caption says 'top 3 models across all categories' but the text describes seven models; please clarify whether the right panel shows the top three judges per dimension group.

Circularity Check

1 steps flagged

Dimension separability is partly self-validated by the same grouping procedure that defines the dimensions; central benchmark results otherwise rest on external human ratings and human-verified references.

specific steps
  1. self definitional [Section 2.2, 'Eval Dimensions and Factor Scoring' (used to construct the dimension leaderboards in Figures 5–6 and the dimension-specific conclusion in Sections 8–9)]
    "A factor score is the equal-weighted mean of its constituent evals’ rater-controlled provider means. To validate the grouping empirically, each eval’s provider vector is Spearman-correlated against every factor leaderboard (eval dimension); an eval is treated as fitting the factor with which it correlates most strongly, and any eval whose best-fit correlation falls below 0.30 is flagged as an orphan candidate for regrouping."

    The validation criterion is not independent of the object being validated: factor leaderboards are equal-weighted means of the same eval-level provider vectors that are then correlated against those leaderboards. An eval’s correlation with the factor it belongs to is mechanically inflated by its own contribution. Assigning each eval to the factor with which it correlates most strongly, then reporting that the resulting factors are distinct, is a self-consistency check rather than an out-of-sample demonstration that TTS naturalness, expressiveness, identity stability, and reliability are 'largely independent evaluation dimensions.' The paper does not report between-factor correlations or leave-one-eval-out/cross-validated factor replication, and it defers paired statistical testing to futur

full rationale

RW-Voice-EQ is largely self-contained against external evidence: TTS and STS scores are human ratings from an external rater pool (785,679 and 48,053 ratings, respectively); SU and ASR are scored against corpus labels, human-verified reference transcripts, and pairwise targets, not against the outputs of the very models whose capabilities are being claimed. The 'dimension-specific' conclusions are therefore empirical observations from direct measurement rather than predictions generated by a fitted model. The one place where the paper's own methodology becomes circular is the factor-validation step in §2.2: factor scores are defined as equal-weighted means of constituent evals, and then the same eval provider vectors are correlated against those factor leaderboards to 'validate' the grouping. Because each eval's own vector contributes to the factor it is being tested against, this validation is a self-consistency check, not an independent test. The paper also acknowledges in §8 that future work should 'expand sample sizes, incorporate paired statistical testing, refine partial-credit scoring for emotion labels, further disentangle acoustic perception from downstream response behavior,' which limits the statistical strength of the independence claims. Nevertheless, the different top-five orderings in Figures 5–6 and the audio-only versus transcript-only ablations in §5.2 are independent empirical content, and there is no significant self-citation chain or fitted-input-called-prediction issue. Overall circularity burden is low.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 1 invented entities

The paper is an empirical benchmark rather than a derivation, so the ledger records measurement choices the conclusions depend on. Two hand-set thresholds (0.30 factor-fit floor; synthetic-speech 'real' cutoff at rating 3) condition reported results. The key domain assumptions are that Likert human ratings, WER vs human-verified transcripts, and the AO/TO matched design measure what they claim, and that the four auto-labeled-then-human-verified ASR sets represent real-world conditions. The only newly posited constructs are the Evaluation Dimensions themselves, validated internally via same-data correlation-based regrouping; no physical entities or fitted natural constants are introduced.

free parameters (2)
  • Factor-orphan correlation threshold = 0.30 (Spearman)
    In §2.2, an eval is treated as fitting the factor with which it Spearman-correlates most strongly; evals whose best-fit correlation falls below 0.30 are flagged as orphan candidates and regrouped. This hand-set threshold shapes which evals compose each Evaluation Dimension and thus shapes the reported dimension-specificity.
  • Synthetic-speech 'real' threshold = rating ≥ 3 (1–5 human-likeness scale)
    Illustrative threshold-derived accuracy for synthetic-speech detection treats scores above 3 as 'real' and a rating of exactly 3 as incorrect (§6.1). The paper notes the mean real-vs-synthetic score gap is the native metric, so this threshold is an analysis choice that conditions the reported SU accuracy numbers.
axioms (5)
  • domain assumption Aggregated 5-point Likert human ratings (≥3 raters per clip) are a valid measure of naturalness, expressiveness, identity, and response quality
    The entire TTS and STS evidence base is mean human ratings collected under the Hume Study Runner protocol (§2.2, §4.1); no independent listening study validates the rubrics' behavior.
  • domain assumption WER against human-verified reference transcripts, using HF Open ASR Leaderboard normalization, is the correct ASR error measure
    ASR conclusions (noise vs music gap, positive vs negative emotion gap, native vs non-native gap) all reduce to corpus-level WER (§7.1, §7.2).
  • domain assumption The four curated ASR sets are representative of real-world production conditions
    The claim that clean benchmarks miss real failures presupposes the curated sets (auto-labeled by Hume's proprietary tagger and Gemini, stratified, then human-verified) are representative; §7.1 describes curation but no external representativeness check.
  • domain assumption Z-scored mixed-effects composites remove systematic rater bias and per-item variance
    Factor scoring assumes the rater-controlled provider means from the mixed-effects model are unbiased capability estimates (§2.2); the model specification is not given.
  • domain assumption The AO and TO conditions in the STS ablation differ only in the presence of audio, with matched lexical content
    The 'transcript-driven agents' conclusion interprets AO–TO differences as audio sensitivity (§5.1–5.2); confounds such as prompt formatting or rater expectations would invalidate that interpretation.
invented entities (1)
  • RW-Voice-EQ Evaluation Dimensions (latent factors: 7 TTS, 5 STS, 3 SU, 4 ASR condition groups) no independent evidence
    purpose: Organize ~50 constituent evaluations into reportable capability constructs and justify per-dimension leaderboards
    The dimensions are posited latent constructs validated only internally — Spearman-correlating each eval against factor leaderboards on the same data, plus manual review (§2.2). No external criterion or cross-benchmark convergent validation is given, so the factor structure lacks an independently falsifiable handle.

pith-pipeline@v1.3.0-alltime-deepseek · 28125 in / 20787 out tokens · 179546 ms · 2026-08-02T00:53:50.393348+00:00 · methodology

0 comments
read the original abstract

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

Figures

Figures reproduced from arXiv: 2607.14846 by Alice Baird, David Ayllon, Franc Camps-Febrer, Georg Streich, Hoon Shin, Jakub Piotr C{\l}apa, Jeffrey Brooks, Jens Madsen, Olya Ossipova, Panagiotis Tzirakis, Rashish Tandon, Sharath Rao, Theo Lebryk, Tigran Soghbatyan.

Figure 1
Figure 1. Figure 1: Overview of the RW-Voice-EQ benchmark. RW-Voice-EQ evaluates voice AI along four domains: (1) Text-to-Speech (TTS), (2) Speech-to-Speech (STS), (3) Speech Understanding (SU), (4) ASR Robustness. Together the four domains characterize a voice system as a capability profile rather than a single aggregate score. 2.1 Benchmark Design RW-Voice-EQ is designed to evaluate whether voice AI systems remain effective… view at source ↗
Figure 2
Figure 2. Figure 2: The RW-Voice-EQ evaluation methodology. Human evaluation (TTS, STS) generates audio per model-item, presents it to multiple raters against a task-specific rubric, and aggregates into per-evaluation scores. Label-based scoring (SU, ASR) compares model predictions against ground-truth labels, pairwise targets, or reference transcripts using accuracy or word error rate. The two paradigms respectively cover pe… view at source ↗
Figure 3
Figure 3. Figure 3: shows the mean Spearman’s correlation between human and SLM ratings across TTS evaluation dimensions, averaged across seven publicly available models (left) and the top 3 models across all categories (right). These modes include proprietary frontier models and open-weight models: Gemini 3.1 Pro (preview), Gemini 2.5 Pro, Gemini 3.1 Flash Lite, Gemini 2.5 Flash, GPT Audio 1.5, Kimi Audio-7b Instruct, Nemotr… view at source ↗
Figure 4
Figure 4. Figure 4: Cross-model audit on VoxPopuli-English. WER (%) from the June 2026 Open ASR Leaderboard [30]. Bad-reference acceptance: the rate at which each model returned reference transcripts verbatim on cases where the reference was identified as requiring a correction These preliminary results establish benchmark optimization as a measurable phenomenon across modern ASR systems and motivate the more detailed analyse… view at source ↗
Figure 5
Figure 5. Figure 5: Top-five TTS system configurations by capability group. Cells report the capability-level mean ± standard deviation. Shading indicates within-capability rank, with darker cells denoting higher rank. Systems are ordered by the number of top-five appearances; the horizontal rule separates systems appearing in multiple top-five lists from those appearing in only one. For Likert-scale evaluations, ratings were… view at source ↗
Figure 6
Figure 6. Figure 6: Top-performing speech-to-speech systems by evaluation dimension. Each column shows the five highest-scoring systems within an evaluation dimension, with cells reporting mean human rating ± standard deviation. Darker coral indicates a higher within-dimension rank (rank 1 darkest; rank 5 lightest), while blank cells indicate that the system did not place in the top five. Emotion Understanding maps to a singl… view at source ↗
Figure 7
Figure 7. Figure 7: Top-performing speech-understanding systems across three evaluation dimensions. Panels show the five highest-scoring systems for speech-emotion recognition (chance-adjusted skill), speaker verification (accuracy), and synthetic-speech detection (threshold-derived accuracy). Colors distinguish general speech-language models from dedicated speaker-verification models. The speech-understanding results are org… view at source ↗
Figure 8
Figure 8. Figure 8: Performance of the top 10 ASR systems by average word error rate (WER) across four human-curated evaluation datasets. Colors distinguish proprietary and open-source models. Parakeet TDT-CTC 110M) to large multi-billion-parameter systems, and includes both dedicated ASR architectures and general-purpose multimodal models applied to transcription. We report Word Error Rate (WER) computed using the text norma… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

33 extracted references · 5 canonical work pages

  1. [1]

    Cambridge University Press, 2018

    Elizabeth Couper-Kuhlen and Margret Selting.Interactional Linguistics: Studying Language in Social Interaction. Cambridge University Press, 2018. doi: 10.1017/9781139507318. 25

  2. [2]

    John Benjamins, 2010

    Dagmar Barth-Weingarten, Elisabeth Reber, and Margret Selting, editors.Prosody in Interaction, volume 23 of Studies in Discourse and Grammar. John Benjamins, 2010. doi: 10.1075/sidag.23

  3. [3]

    Oxford University Press, 2021

    Carlos Gussenhoven and Aoju Chen, editors.The Oxford Handbook of Language Prosody. Oxford University Press, 2021. doi: 10.1093/oxfordhb/9780198832232.001.0001

  4. [4]

    Cambridge University Press, 1980

    JohnLaver.The Phonetic Description of Voice Quality,volume31ofCambridge Studies in Linguistics. Cambridge University Press, 1980

  5. [5]

    Scherer and Howard Giles, editors.Social Markers in Speech

    Klaus R. Scherer and Howard Giles, editors.Social Markers in Speech. Cambridge University Press, 1979

  6. [6]

    Asimplestsystematicsfortheorganizationofturn-taking for conversation.Language, 50(4):696–735, 1974

    HarveySacks,EmanuelA.Schegloff,andGailJefferson. Asimplestsystematicsfortheorganizationofturn-taking for conversation.Language, 50(4):696–735, 1974. URLhttps://www.jstor.org/stable/412243

  7. [7]

    doi: 10.1162/tacl.a.628

    YimingChen,XianghuYue,ChenZhang,XiaoxueGao,RobbyT.Tan,andHaizhouLi.VoiceBench: Benchmarking LLM-based voice assistants.Transactions of the Association for Computational Linguistics, 14:378–398, 2026. doi: 10.1162/tacl.a.628. URLhttps://aclanthology.org/2026.tacl-1.18/

  8. [9]

    Shah, David Solans Noguero, Mikko A

    Muhammad A. Shah, David Solans Noguero, Mikko A. Heikkilä, Bhiksha Raj, and Nicolas Kourtellis. Speech robust bench: A robustness benchmark for speech recognition.arXiv preprint arXiv:2403.07937, 2024. URL https://arxiv.org/abs/2403.07937

  9. [10]

    SpeechParaling-Bench: A comprehensive benchmark for paralinguistic-aware speech generation

    Ruohan Liu, Shukang Yin, Tao Wang, Dong Zhang, Weiji Zhuang, Shuhuai Ren, Ran He, Caifeng Shan, and Chaoyou Fu. SpeechParaling-Bench: A comprehensive benchmark for paralinguistic-aware speech generation. arXiv preprint arXiv:2604.20842, 2026. URLhttps://arxiv.org/abs/2604.20842

  10. [11]

    ITU-T Recommendation P.800: Methods for subjective determination of transmission quality

    International Telecommunication Union. ITU-T Recommendation P.800: Methods for subjective determination of transmission quality. Technical Report P.800 (08/96), International Telecommunication Union, 1996. URL https://www.itu.int/rec/T-REC-P.800-199608-I/en

  11. [12]

    Technical Report BS.1534-3, International Telecommunication Union, 2015

    InternationalTelecommunicationUnion.ITU-RRecommendationBS.1534-3: Methodforthesubjectiveassessment of intermediate quality level of audio systems. Technical Report BS.1534-3, International Telecommunication Union, 2015. URLhttps://www.itu.int/rec/R-REC-BS.1534-3-201510-I/en

  12. [13]

    Black and Keiichi Tokuda

    Alan W. Black and Keiichi Tokuda. The blizzard challenge – 2005: Evaluating corpus-based speech synthesis on common datasets. InProceedings of Interspeech 2005, pages 77–80, 2005. doi: 10.21437/Interspeech.2005-72

  13. [14]

    The VoiceMOS challenge2022

    Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. The VoiceMOS challenge2022. InProceedings of Interspeech 2022,pages4536–4540,2022. doi: 10.21437/Interspeech.2022-970

  14. [15]

    Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao

    Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E. Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao. The VoiceMOS challenge 2024: Beyond speech quality prediction.arXiv preprint arXiv:2409.07001, 2024. URLhttps://arxiv.org/abs/2409.07001

  15. [16]

    InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems.arXiv preprint arXiv:2506.16381, 2025

    Kexin Huang, Qian Tu, Liwei Fan, Chenchen Yang, Dong Zhang, Shimin Li, Zhaoye Fei, Qinyuan Cheng, and Xipeng Qiu. InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems.arXiv preprint arXiv:2506.16381, 2025. URLhttps://arxiv.org/abs/2506.16381

  16. [17]

    EmergentTTS-Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge.arXiv preprint arXiv:2505.23009, 2025

    Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li, and Alex Smola. EmergentTTS-Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge.arXiv preprint arXiv:2505.23009, 2025. URLhttps://arxiv.org/abs/2505.23009

  17. [18]

    Walker, Diane J

    Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. PARADISE: A framework for evaluatingspokendialogueagents. InProceedings of the 35th Annual Meeting of the Association for Computational Linguistics and the 8th Conference of the European Chapter of the Association for Computational Linguistics, pages 271–280, 1997. doi: 10.3115/976909...

  18. [19]

    SD-Eval: A benchmark dataset for spoken dialogue understanding beyond words

    Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu. SD-Eval: A benchmark dataset for spoken dialogue understanding beyond words. InAdvances in Neural Information Processing Systems, volume 37, 2024. doi: 10.52202/079017-1813

  19. [20]

    URO-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models

    Ruiqi Yan, Xiquan Li, Wenxi Chen, Zhikang Niu, Chen Yang, Ziyang Ma, Kai Yu, and Xie Chen. URO-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models. InFindings of the Association for Computational Linguistics: EMNLP 2025,pages17211–17242,2025. doi: 10.18653/v1/2025.findings-emnlp.933. URLhttps://aclanthology.org/2025.findings-emnlp.933/

  20. [21]

    S2S-Arena: Evaluating paralinguistic instruction following in speech-to-speech models.arXiv preprint arXiv:2503.05085, 2026

    Feng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue, Fan Bu, Yuhao Du, Xiangying Chen, Benyou Wang, and Haizhou Li. S2S-Arena: Evaluating paralinguistic instruction following in speech-to-speech models.arXiv preprint arXiv:2503.05085, 2026. URLhttps://arxiv.org/abs/2503.05085. Version 2

  21. [22]

    AIR-bench: Benchmarking large audio-language models via generative comprehension

    Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, and Jingren Zhou. AIR-bench: Benchmarking large audio-language models via generative comprehension. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pages 1979–1998, 2024. doi: 10.18653/v1/2024.ac...

  22. [23]

    Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, and Nancy F. Chen. AudioBench: A universal benchmark for audio large language models. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4297–4316, 2025. ...

  23. [24]

    AHELM: A holistic evaluation of audio-language models.arXiv preprint arXiv:2508.21376, 2025

    Tony Lee, Haoqin Tu, Chi Heem Wong, Zijun Wang, Siwei Yang, Yifan Mai, Yuyin Zhou, Cihang Xie, and Percy Liang. AHELM: A holistic evaluation of audio-language models.arXiv preprint arXiv:2508.21376, 2025. URL https://arxiv.org/abs/2508.21376

  24. [25]

    Common voice: A massively-multilingual speech corpus

    Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. InProceedings of the Twelfth Language Resources and Evaluation Conference, pages 4218–4222, 2020. URL https://aclanthology.org/2020.lrec-1.520/

  25. [26]

    FLEURS: Few-shot learning evaluation of universal representations of speech

    AlexisConneau, MinMa, SimranKhanuja, YuZhang, VeraAxelrod, SiddharthDalmia, JasonRiesa, ClaraRivera, and Ankur Bapna. FLEURS: Few-shot learning evaluation of universal representations of speech. In2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805, 2023. doi: 10.1109/SLT54892.2023.10023141

  26. [27]

    CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings

    Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent, Ashish Arora, Xuankai Chang, Sanjeev Khudanpur, Vimal Manohar, Daniel Povey, Desh Raj, David Snyder, Aswin Shanmugam Subramanian, Jan Trmal, Bar Ben Yair, Christoph Boeddeker, Zhaoheng Ni, Yusuke Fujita, Shota Horiguchi, Naoyuki Kanda, Takuya Yoshioka, and Neville Ryant. CHiME-6 challenge: Tac...

  27. [28]

    Librispeech: An asr corpus based on public domain audio books

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An asr corpus based on public domain audio books. In2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210. IEEE, 2015

  28. [29]

    Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation

    ChanghanWang,MorganeRiviere,AnnLee,AnneWu,ChaitanyaTalnikar,DanielHaziza,MaryWilliamson,Juan Pino, and Emmanuel Dupoux. Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat...

  29. [30]

    Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation, 2025

    Vaibhav Srivastav, Steven Zheng, Eric Bezzam, Eustache Le Bihan, Nithin Koluguri, Piotr Żelasko, Somshubra Majumdar, Adel Moumen, and Sanchit Gandhi. Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation, 2025. URLhttps://arxiv.org/abs/2510.06961. 27

  30. [31]

    Colleen Richey, Maria A. Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, Paul Gamble, Jeffrey Hetherly, Cory Stephenson, and Karl Ni. Voices Obscured in Complex Environmental Settings (VOiCES) Corpus. InProc. Interspeech 2018, pages 1566–1570, 2018. doi: 10.214...

  31. [32]

    BERSting at the screams: A benchmark for distanced, emotional and shouted speech recognition.Computer Speech & Language, 2025

    Paige Tuttösí, Mantaj Dhillon, Luna Sang, Shane Eastwood, Poorvi Bhatia, Quang Minh Dinh, Avni Kapoor, Yewon Jin, and Angelica Lim. BERSting at the screams: A benchmark for distanced, emotional and shouted speech recognition.Computer Speech & Language, 2025. arXiv:2505.00059

  32. [33]

    A female speaker delivers a clear, expressive speech in a quiet, high- quality recording

    Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Maël Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. The AMI meeting corpus: A pre-announcement. InMachine Learning for Multimodal Inter...

  33. [2025]

    URLhttps://arxiv.org/abs/2505.15727