REVIEW 4 major objections 6 minor 33 references
Voice AI quality cannot be captured by one number: performance is dimension-specific and should be reported as a profile.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 00:53 UTC pith:XLP2AMWV
load-bearing objection A large, well-resourced benchmark with a plausible but not fully proven profile-based evaluation claim; worth reviewing, but needs data release and statistical tightening. the 4 major comments →
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
RW-Voice-EQ evaluates text-to-speech along seven dimensions (acting/role-fit, expressiveness, voice identity, language stability, reliability, long-form stability, acoustic quality), speech-to-speech along five (emotion understanding, emotion alignment, expressivity robustness, voice naturalness under stress, problem redirection), speech understanding along three (emotion recognition, speaker verification, synthetic-speech detection), and ASR across four condition sets (accent, emotion, background audio, conversation). The central finding is that systems rank differently in every dimension: no system leads all dimensions, and even within a single system, task behavior and vocal quality can d
What carries the argument
The central mechanism is the dimension-factor scoring pipeline: individual evaluations are grouped into latent factors, each evaluation's provider-level scores are Spearman-correlated against every factor leaderboard, and groupings are retained or regrouped when the best-fit correlation falls below 0.30, with manual verification. For speech-to-speech, the load-bearing design is the audio-only versus transcript-only ablation, which isolates whether access to the acoustic signal changes agent behavior. For ASR, the benchmark uses four human-curated private datasets with consensus human-verified reference transcripts. Speech-language-model judges are treated as auxiliary evaluators, validated a
Load-bearing premise
The independence conclusion rests on the assumption that each evaluation dimension measures the construct it names, and that the observed dimension separability is not an artifact of the grouping procedure, since the factor validation reuses the same data that generate the dimension scores.
What would settle it
A re-run of the factor grouping on an independent item set that does not reproduce the same dimension structure (evaluations' best-fit Spearman correlations shift across factor leaderboards or fall below the 0.30 threshold) would undermine the claim that performance is dimension-specific. Similarly, if the audio-only versus transcript-only contrasts in speech-to-speech do not survive paired significance testing with larger samples, the 'transcript-driven' conclusion fails.
If this is right
- Single-number leaderboards should be replaced by per-dimension profiles; model selection becomes use-case dependent on which capability matters most for deployment.
- Clean-speech ASR benchmarks are insufficient for production decisions; systems should be reported with condition-level robustness profiles covering accent, emotion, background audio, and conversation.
- Access to audio in speech-to-speech agents does not guarantee use of vocal affect; evaluation should measure the audio-versus-transcript difference explicitly.
- Speech-language-model judges are verification-dependent: they agree well with humans on target-given correctness tasks but degrade on open-ended perceptual judgments such as voice identity and acting role-fit.
- Withheld, private evaluation sets are necessary because public benchmark optimization is already measurable in state-of-the-art ASR systems.
Where Pith is reading between the lines
- If dimension-specific profiles become standard, downstream applications will select systems per deployment context (e.g., long-form narration vs. identity-critical assistive voice), and aggregate rankings will lose predictive authority entirely.
- The benchmark's diagnostic probes (masked-word tracking, orthographic-switch detection, synthetic-voice matching) could be generalized into a standard contamination audit for any speech leaderboard.
- The verification-dependence finding for speech-language-model judges suggests a routing rule: use SLMs for unambiguous, answer-graded tasks and reserve human raters for open-ended perceptual constructs, rather than treating SLM scores as uniformly valid.
- The consistent gap between positive high-arousal speech and negative high-arousal speech in ASR points to a concrete training-data target: adding expressive positive speech should reduce that gap if the paper's underrepresentation explanation is correct.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RW-Voice-EQ Bench, a large multidimensional benchmark for voice AI covering TTS, speech-to-speech (STS), speech understanding (SU), and ASR robustness. TTS and STS are scored by human raters on task-specific rubrics (over 830k ratings), while SU and ASR are scored against human-labeled references, transcripts, and pairwise targets. The benchmark includes 31 TTS configurations, 15 STS configurations, 18–19 SU systems, and about 40 ASR models, plus an analysis of speech-language-model judges and evidence of ASR benchmark optimization. The central claim is that voice-AI performance is dimension-specific: strong performance in one capability does not predict strong performance in another, so systems should be reported as capability profiles rather than single aggregate scores. The paper also reports that some STS agents remain largely transcript-driven, that relative emotion comparison is much easier than absolute emotion identification, that dedicated speaker-verification models outperform general audio-language judges, and that ASR robustness failures are not captured by clean-speech benchmarks.
Significance. If the central claim holds, the benchmark would support a meaningful shift in how voice AI is evaluated, away from single-number leaderboards and toward per-dimension diagnostic profiles. The paper has genuine strengths: the evaluation is unusually broad, with over one million human ratings collected from an external rater pool; SU and ASR results are scored against human-verified labels and transcripts rather than model outputs; the benchmark and leaderboards are publicly released; and the analysis of SLM-judge agreement is a useful methodological contribution. The ASR robustness dataset with human-corrected references is a valuable resource. The main unresolved issue is statistical: the dimension-separability conclusion is asserted from descriptive top-five patterns and a factor-grouping procedure that is validated on the same data it defines. The paper itself acknowledges the lack of paired statistical testing and the non-inferential nature of reported standard deviations, yet the abstract and discussion draw strong dimension-specific conclusions from those same results.
major comments (4)
- [§2.2, factor grouping and dimension validation] The dimensionality claim is validated circularly. Each eval's provider vector is Spearman-correlated against every factor leaderboard constructed from those same evals, and evals are assigned to their best-fitting factor if the correlation exceeds 0.30. With 31 TTS providers and seven candidate factors, this procedure can create apparent separation by chance. The paper does not report between-factor correlations, confidence intervals for factor assignments, or a stability analysis such as drop-one-eval or split-half replication. Since the abstract's 'largely independent evaluation dimensions' is load-bearing for the profile-vs-aggregate conclusion, please add cross-validated or out-of-sample factor replication and report the between-factor correlation matrix.
- [§5.2, audio-only vs transcript-only ablation] The STS conclusion 'access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven' rests on descriptive direction counts: four systems scored higher under AO, one unchanged, ten lower on the naturalistic set, and ten higher on the scripted set. No significance tests, confidence intervals, or effect sizes are reported. The Figure 6 caption explicitly states that the displayed standard deviations 'are not estimates of rater-level uncertainty or confidence intervals,' and §8 defers 'paired statistical testing' to future work. Please add per-system paired tests or bootstrap CIs across scenarios, and report the AO–TO difference with uncertainty for each of the 15 systems, especially for GPT-Realtime-2's +0.17 difference.
- [§4.2, Figure 5 and TTS dimension independence] The claim that naturalness, expressiveness, identity stability, and reliability are 'largely independent evaluation dimensions' is supported only by top-five membership patterns. No correlation matrix among the seven dimension scores is provided, no null model for expected top-five overlap is given, and there are no confidence intervals around dimension means. With only 31 systems and top-five truncation, the observed absence of a system in all seven top-five lists is weak evidence of independence. Please report the full inter-dimension Spearman correlation matrix over all 31 providers, with uncertainty, and test whether the observed overlap differs from what chance would produce.
- [§7.2, ASR clean-benchmark comparison] The claim that 'real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks' requires a direct quantitative comparison between the evaluated models' clean-benchmark WER and their per-condition WER. The paper shows that rankings reorder across the four tracks, but it does not report a correlation between clean leaderboard WER and the robustness WERs for the same models. Please add a scatterplot or correlation table (e.g., LibriSpeech or VoxPopuli WER vs accent/emotion/noise/conversational WER) to substantiate the claim that clean performance is not predictive.
minor comments (6)
- [Abstract] Typo: 'Real World Voice EQ Bench a' should be 'Real World Voice EQ Bench'.
- [§6.1 / Table 6] The text says 18 speech-understanding systems were evaluated, but Table 6 lists 19 rows (including gemma-3n). Please reconcile the count and the table.
- [§7.1 / Table 7] The number of ASR systems is inconsistent: §7.1 says 40, §7.2 says 41, and Table 7 appears to contain 42 rows. Please correct the counts consistently.
- [Table 1, Acoustic Quality row] The Acoustic Quality row has no prompt/generation/rating counts, and the text says fidelity ratings were collected 'within the expression evaluations.' Please state explicitly how the Acoustic Quality dimension score in Figure 5 was computed and from which evaluations.
- [§4.1 vs §2.1] The TTS dimension is called 'Multilingual Code-Switching' in §4.1 and Table 1's header, but 'Language Stability' elsewhere. Use one consistent name.
- [Figure 3] The caption says 'top 3 models across all categories' but the text describes seven models; please clarify whether the right panel shows the top three judges per dimension group.
Circularity Check
Dimension separability is partly self-validated by the same grouping procedure that defines the dimensions; central benchmark results otherwise rest on external human ratings and human-verified references.
specific steps
-
self definitional
[Section 2.2, 'Eval Dimensions and Factor Scoring' (used to construct the dimension leaderboards in Figures 5–6 and the dimension-specific conclusion in Sections 8–9)]
"A factor score is the equal-weighted mean of its constituent evals’ rater-controlled provider means. To validate the grouping empirically, each eval’s provider vector is Spearman-correlated against every factor leaderboard (eval dimension); an eval is treated as fitting the factor with which it correlates most strongly, and any eval whose best-fit correlation falls below 0.30 is flagged as an orphan candidate for regrouping."
The validation criterion is not independent of the object being validated: factor leaderboards are equal-weighted means of the same eval-level provider vectors that are then correlated against those leaderboards. An eval’s correlation with the factor it belongs to is mechanically inflated by its own contribution. Assigning each eval to the factor with which it correlates most strongly, then reporting that the resulting factors are distinct, is a self-consistency check rather than an out-of-sample demonstration that TTS naturalness, expressiveness, identity stability, and reliability are 'largely independent evaluation dimensions.' The paper does not report between-factor correlations or leave-one-eval-out/cross-validated factor replication, and it defers paired statistical testing to futur
full rationale
RW-Voice-EQ is largely self-contained against external evidence: TTS and STS scores are human ratings from an external rater pool (785,679 and 48,053 ratings, respectively); SU and ASR are scored against corpus labels, human-verified reference transcripts, and pairwise targets, not against the outputs of the very models whose capabilities are being claimed. The 'dimension-specific' conclusions are therefore empirical observations from direct measurement rather than predictions generated by a fitted model. The one place where the paper's own methodology becomes circular is the factor-validation step in §2.2: factor scores are defined as equal-weighted means of constituent evals, and then the same eval provider vectors are correlated against those factor leaderboards to 'validate' the grouping. Because each eval's own vector contributes to the factor it is being tested against, this validation is a self-consistency check, not an independent test. The paper also acknowledges in §8 that future work should 'expand sample sizes, incorporate paired statistical testing, refine partial-credit scoring for emotion labels, further disentangle acoustic perception from downstream response behavior,' which limits the statistical strength of the independence claims. Nevertheless, the different top-five orderings in Figures 5–6 and the audio-only versus transcript-only ablations in §5.2 are independent empirical content, and there is no significant self-citation chain or fitted-input-called-prediction issue. Overall circularity burden is low.
Axiom & Free-Parameter Ledger
free parameters (2)
- Factor-orphan correlation threshold =
0.30 (Spearman)
- Synthetic-speech 'real' threshold =
rating ≥ 3 (1–5 human-likeness scale)
axioms (5)
- domain assumption Aggregated 5-point Likert human ratings (≥3 raters per clip) are a valid measure of naturalness, expressiveness, identity, and response quality
- domain assumption WER against human-verified reference transcripts, using HF Open ASR Leaderboard normalization, is the correct ASR error measure
- domain assumption The four curated ASR sets are representative of real-world production conditions
- domain assumption Z-scored mixed-effects composites remove systematic rater bias and per-item variance
- domain assumption The AO and TO conditions in the STS ablation differ only in the presence of audio, with matched lexical content
invented entities (1)
-
RW-Voice-EQ Evaluation Dimensions (latent factors: 7 TTS, 5 STS, 3 SU, 4 ASR condition groups)
no independent evidence
read the original abstract
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.
Figures
Reference graph
Works this paper leans on
-
[1]
Cambridge University Press, 2018
Elizabeth Couper-Kuhlen and Margret Selting.Interactional Linguistics: Studying Language in Social Interaction. Cambridge University Press, 2018. doi: 10.1017/9781139507318. 25
-
[2]
Dagmar Barth-Weingarten, Elisabeth Reber, and Margret Selting, editors.Prosody in Interaction, volume 23 of Studies in Discourse and Grammar. John Benjamins, 2010. doi: 10.1075/sidag.23
-
[3]
Carlos Gussenhoven and Aoju Chen, editors.The Oxford Handbook of Language Prosody. Oxford University Press, 2021. doi: 10.1093/oxfordhb/9780198832232.001.0001
arXiv 2021
-
[4]
Cambridge University Press, 1980
JohnLaver.The Phonetic Description of Voice Quality,volume31ofCambridge Studies in Linguistics. Cambridge University Press, 1980
1980
-
[5]
Scherer and Howard Giles, editors.Social Markers in Speech
Klaus R. Scherer and Howard Giles, editors.Social Markers in Speech. Cambridge University Press, 1979
1979
-
[6]
Asimplestsystematicsfortheorganizationofturn-taking for conversation.Language, 50(4):696–735, 1974
HarveySacks,EmanuelA.Schegloff,andGailJefferson. Asimplestsystematicsfortheorganizationofturn-taking for conversation.Language, 50(4):696–735, 1974. URLhttps://www.jstor.org/stable/412243
1974
-
[7]
YimingChen,XianghuYue,ChenZhang,XiaoxueGao,RobbyT.Tan,andHaizhouLi.VoiceBench: Benchmarking LLM-based voice assistants.Transactions of the Association for Computational Linguistics, 14:378–398, 2026. doi: 10.1162/tacl.a.628. URLhttps://aclanthology.org/2026.tacl-1.18/
-
[9]
Shah, David Solans Noguero, Mikko A
Muhammad A. Shah, David Solans Noguero, Mikko A. Heikkilä, Bhiksha Raj, and Nicolas Kourtellis. Speech robust bench: A robustness benchmark for speech recognition.arXiv preprint arXiv:2403.07937, 2024. URL https://arxiv.org/abs/2403.07937
Pith/arXiv arXiv 2024
-
[10]
SpeechParaling-Bench: A comprehensive benchmark for paralinguistic-aware speech generation
Ruohan Liu, Shukang Yin, Tao Wang, Dong Zhang, Weiji Zhuang, Shuhuai Ren, Ran He, Caifeng Shan, and Chaoyou Fu. SpeechParaling-Bench: A comprehensive benchmark for paralinguistic-aware speech generation. arXiv preprint arXiv:2604.20842, 2026. URLhttps://arxiv.org/abs/2604.20842
Pith/arXiv arXiv 2026
-
[11]
ITU-T Recommendation P.800: Methods for subjective determination of transmission quality
International Telecommunication Union. ITU-T Recommendation P.800: Methods for subjective determination of transmission quality. Technical Report P.800 (08/96), International Telecommunication Union, 1996. URL https://www.itu.int/rec/T-REC-P.800-199608-I/en
1996
-
[12]
Technical Report BS.1534-3, International Telecommunication Union, 2015
InternationalTelecommunicationUnion.ITU-RRecommendationBS.1534-3: Methodforthesubjectiveassessment of intermediate quality level of audio systems. Technical Report BS.1534-3, International Telecommunication Union, 2015. URLhttps://www.itu.int/rec/R-REC-BS.1534-3-201510-I/en
2015
-
[13]
Alan W. Black and Keiichi Tokuda. The blizzard challenge – 2005: Evaluating corpus-based speech synthesis on common datasets. InProceedings of Interspeech 2005, pages 77–80, 2005. doi: 10.21437/Interspeech.2005-72
-
[14]
Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. The VoiceMOS challenge2022. InProceedings of Interspeech 2022,pages4536–4540,2022. doi: 10.21437/Interspeech.2022-970
-
[15]
Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao
Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E. Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao. The VoiceMOS challenge 2024: Beyond speech quality prediction.arXiv preprint arXiv:2409.07001, 2024. URLhttps://arxiv.org/abs/2409.07001
Pith/arXiv arXiv 2024
-
[16]
Kexin Huang, Qian Tu, Liwei Fan, Chenchen Yang, Dong Zhang, Shimin Li, Zhaoye Fei, Qinyuan Cheng, and Xipeng Qiu. InstructTTSEval: Benchmarking complex natural-language instruction following in text-to-speech systems.arXiv preprint arXiv:2506.16381, 2025. URLhttps://arxiv.org/abs/2506.16381
Pith/arXiv arXiv 2025
-
[17]
Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li, and Alex Smola. EmergentTTS-Eval: Evaluating TTS models on complex prosodic, expressiveness, and linguistic challenges using model-as-a-judge.arXiv preprint arXiv:2505.23009, 2025. URLhttps://arxiv.org/abs/2505.23009
Pith/arXiv arXiv 2025
-
[18]
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. PARADISE: A framework for evaluatingspokendialogueagents. InProceedings of the 35th Annual Meeting of the Association for Computational Linguistics and the 8th Conference of the European Chapter of the Association for Computational Linguistics, pages 271–280, 1997. doi: 10.3115/976909...
arXiv 1997
-
[19]
SD-Eval: A benchmark dataset for spoken dialogue understanding beyond words
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu. SD-Eval: A benchmark dataset for spoken dialogue understanding beyond words. InAdvances in Neural Information Processing Systems, volume 37, 2024. doi: 10.52202/079017-1813
-
[20]
URO-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models
Ruiqi Yan, Xiquan Li, Wenxi Chen, Zhikang Niu, Chen Yang, Ziyang Ma, Kai Yu, and Xie Chen. URO-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models. InFindings of the Association for Computational Linguistics: EMNLP 2025,pages17211–17242,2025. doi: 10.18653/v1/2025.findings-emnlp.933. URLhttps://aclanthology.org/2025.findings-emnlp.933/
-
[21]
Feng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue, Fan Bu, Yuhao Du, Xiangying Chen, Benyou Wang, and Haizhou Li. S2S-Arena: Evaluating paralinguistic instruction following in speech-to-speech models.arXiv preprint arXiv:2503.05085, 2026. URLhttps://arxiv.org/abs/2503.05085. Version 2
Pith/arXiv arXiv 2026
-
[22]
AIR-bench: Benchmarking large audio-language models via generative comprehension
Qian Yang, Jin Xu, Wenrui Liu, Yunfei Chu, Ziyue Jiang, Xiaohuan Zhou, Yichong Leng, Yuanjun Lv, Zhou Zhao, Chang Zhou, and Jingren Zhou. AIR-bench: Benchmarking large audio-language models via generative comprehension. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pages 1979–1998, 2024. doi: 10.18653/v1/2024.ac...
-
[23]
Bin Wang, Xunlong Zou, Geyu Lin, Shuo Sun, Zhuohan Liu, Wenyu Zhang, Zhengyuan Liu, AiTi Aw, and Nancy F. Chen. AudioBench: A universal benchmark for audio large language models. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4297–4316, 2025. ...
-
[24]
AHELM: A holistic evaluation of audio-language models.arXiv preprint arXiv:2508.21376, 2025
Tony Lee, Haoqin Tu, Chi Heem Wong, Zijun Wang, Siwei Yang, Yifan Mai, Yuyin Zhou, Cihang Xie, and Percy Liang. AHELM: A holistic evaluation of audio-language models.arXiv preprint arXiv:2508.21376, 2025. URL https://arxiv.org/abs/2508.21376
Pith/arXiv arXiv 2025
-
[25]
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber. Common voice: A massively-multilingual speech corpus. InProceedings of the Twelfth Language Resources and Evaluation Conference, pages 4218–4222, 2020. URL https://aclanthology.org/2020.lrec-1.520/
2020
-
[26]
FLEURS: Few-shot learning evaluation of universal representations of speech
AlexisConneau, MinMa, SimranKhanuja, YuZhang, VeraAxelrod, SiddharthDalmia, JasonRiesa, ClaraRivera, and Ankur Bapna. FLEURS: Few-shot learning evaluation of universal representations of speech. In2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805, 2023. doi: 10.1109/SLT54892.2023.10023141
arXiv 2023
-
[27]
CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings
Shinji Watanabe, Michael Mandel, Jon Barker, Emmanuel Vincent, Ashish Arora, Xuankai Chang, Sanjeev Khudanpur, Vimal Manohar, Daniel Povey, Desh Raj, David Snyder, Aswin Shanmugam Subramanian, Jan Trmal, Bar Ben Yair, Christoph Boeddeker, Zhaoheng Ni, Yusuke Fujita, Shota Horiguchi, Naoyuki Kanda, Takuya Yoshioka, and Neville Ryant. CHiME-6 challenge: Tac...
2020
-
[28]
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. Librispeech: An asr corpus based on public domain audio books. In2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5206–5210. IEEE, 2015
2015
-
[29]
Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
ChanghanWang,MorganeRiviere,AnnLee,AnneWu,ChaitanyaTalnikar,DanielHaziza,MaryWilliamson,Juan Pino, and Emmanuel Dupoux. Voxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat...
2021
-
[30]
Vaibhav Srivastav, Steven Zheng, Eric Bezzam, Eustache Le Bihan, Nithin Koluguri, Piotr Żelasko, Somshubra Majumdar, Adel Moumen, and Sanchit Gandhi. Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation, 2025. URLhttps://arxiv.org/abs/2510.06961. 27
arXiv 2025
-
[31]
Colleen Richey, Maria A. Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, Paul Gamble, Jeffrey Hetherly, Cory Stephenson, and Karl Ni. Voices Obscured in Complex Environmental Settings (VOiCES) Corpus. InProc. Interspeech 2018, pages 1566–1570, 2018. doi: 10.214...
-
[32]
Paige Tuttösí, Mantaj Dhillon, Luna Sang, Shane Eastwood, Poorvi Bhatia, Quang Minh Dinh, Avni Kapoor, Yewon Jin, and Angelica Lim. BERSting at the screams: A benchmark for distanced, emotional and shouted speech recognition.Computer Speech & Language, 2025. arXiv:2505.00059
Pith/arXiv arXiv 2025
-
[33]
A female speaker delivers a clear, expressive speech in a quiet, high- quality recording
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Maël Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre Wellner. The AMI meeting corpus: A pre-announcement. InMachine Learning for Multimodal Inter...
2005
-
[2025]
URLhttps://arxiv.org/abs/2505.15727
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.