REVIEW 4 major objections 6 minor 295 references
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A survey of 12 benchmarks finds LLM-based evaluators beat heuristic, embedding, and learning methods at matching human judgment across text, vision, and audio generation.
desk verdict Broad cross-modal survey with a genuinely useful taxonomy, but the central quantitative claim about LLM-based evaluation is not backed by the reported meta-evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a five-paradigm taxonomy—heuristic, embedding-based, learning-based, LLM-based, and benchmark-based evaluation—applied uniformly to text, visual, and audio generation. The comparative engine is meta-evaluation: measuring an automatic evaluator's agreement with human judgments, through Spearman correlation for single-wise protocols and preference accuracy for pair-wise protocols. The taxonomy organizes the field, and the meta-evaluation protocol produces the paper's central ranking of paradigms.
What would settle it
Re-run the same representative methods on an independent set of meta-evaluation benchmarks that includes open-ended creative generation, low-resource languages, and non-English tasks; if a non-LLM method such as COMET-22 or a fine-tuned reward model ties or beats the LLM-based judges on average across those benchmarks, the paper's 'significantly outperform' claim would fail to generalize.
Extended reading notes
Core claim
The paper's quantitative core is a comparison of representative methods from four paradigms (heuristic, embedding-based, learning-based, and LLM-based) on 12 meta-evaluation benchmarks, reported in Tables 9 and 10. On the seven NLG-specific benchmarks, LLM-based methods achieve the highest Spearman correlations, with only a slight exception on WMT-22 where the learning-based COMET-22 remains competitive. On five broader benchmarks spanning diverse domains, only LLM-based methods could be tested, and prompt-based and fine-tuned evaluators both perform strongly. The authors also report that reward models fine-tuned on human preferences can beat prompt-based methods on preference benchmarks like Auto-J and RewardBench, but underperform them on task-specific NLG benchmarks, and that reasoning-optimized models like DeepSeek-R1 are not universally better evaluators despite excelling on reasoning-heavy critique benchmarks.
Load-bearing premise
The comparative conclusion assumes the 12 chosen meta-evaluation benchmarks and the representative methods selected in Tables 9 and 10 give a fair, comprehensive basis for ranking the four paradigms.
Editorial extensions
If this is right
- If LLM-based evaluation is indeed superior, new evaluation tasks should default to LLM-as-a-judge protocols rather than bespoke n-gram or embedding metrics.
- Fine-tuned compact evaluators can replace expensive prompt-based judges at lower cost, making large-scale evaluation and RLHF reward modeling more practical.
- Reward models need further work on generalization before they can serve as general-purpose evaluators across diverse NLG tasks.
- Reasoning-optimized models may be the right choice for evaluating complex reasoning and critique tasks, but not for all generation domains.
- The same five-paradigm taxonomy, if it transfers to vision and audio, gives researchers a shared language for comparing evaluation methods across modalities.
Reading between the lines
- The paper's headline ranking is measured almost entirely on text-generation benchmarks; whether the same hierarchy holds for image and audio generation remains untested because meta-evaluation benchmarks in those modalities are scarce.
- Because the judges and the judged are often both LLMs, part of the apparent superiority could reflect shared model biases rather than genuine alignment with human taste; a human-blind re-test with non-LLM baselines would separate the two.
- The taxonomy suggests a testable extension: building cross-modal meta-evaluation benchmarks where the same generated concept is judged in text, image, and speech forms, to see whether LLM-based evaluation stays on top when the content is held constant.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys automatic evaluation methods for text, vision, and speech generation, organizing approaches into a proposed unified taxonomy of five paradigms: heuristic, embedding-based, learning-based, LLM-based, and benchmark-based evaluation. For text generation it provides detailed reviews of each paradigm and reports a quantitative meta-evaluation across twelve benchmarks (Tables 9 and 10), concluding that LLM-based methods currently outperform other approaches significantly. It then extends the taxonomy to image/video generation (Section 4) and audio/speech generation (Section 5), including comparative experiments for text-to-image alignment (Table 17) and MOS prediction (Tables 24 and 25), and closes with future directions for cross-modal evaluation.
Significance. If its central quantitative claim were well supported, this survey would be a valuable resource: it assembles a large and current literature, spans three modalities, and attempts the kind of systematic meta-evaluation that is often missing from surveys. The paper's strengths include its broad coverage of recent LLM/ALM-based evaluators, its effort to organize methods under a common taxonomy, and its explicit tabulation of meta-evaluation benchmarks. At the same time, the headline claim of significant LLM superiority rests on tables with missing entries, malformed values, and an underspecified protocol, and the paper's own audio results contain a counterexample. The taxonomy claim is also internally inconsistent (five vs. four paradigms). These issues are fixable within the paper's scope, but they currently block acceptance.
major comments (4)
- [§3.5.2, Tables 9 and 10] The central claim that 'LLM-based automated evaluation methods currently outperform other approaches significantly' is not supported by the evidence as reported. The protocol is underspecified: no prompt templates, model versions, inference settings, or reference-availability conditions are given, and no code is released. The comparison is also selective: by the authors' own statement, the last five benchmarks (CriticBench through RewardBench) were tested only with LLM-based methods, so Table 10 cannot inform a cross-paradigm ranking, and even in Table 9 several cells are empty (e.g., FED for BLEU, MoverScore, and COMET-22; WMT-22 for G-Eval and DeepSeek). Moreover, on WMT-22 the best non-LLM method, COMET-22 (56.4), beats the best LLM-based method, InstructScore (51.9), so the claim is at least task-dependent. The conclusion should be restricted to the seven complete benchmarks, or the missing protocol details and significance analysis should be supplied.
- [§3.5.2, Table 9] Table 9 contains malformed entries that make the quantitative comparison unreliable. The COMET-22 row appears as '33.8 11.6 -56.439.2 13.8 40.9', with two values run together; the TER row contains '21.95' and '6.20' where plausible values would be '21.9' and '6.2'; and the BLEU row reports '-1.17' for OpenMEVA without explanation. Since this table is the only quantitative basis for the paper's main superiority claim, these values must be corrected and the missing entries either filled or explicitly excluded from the conclusion.
- [Abstract, §3, §6] The paper is internally inconsistent about the number of paradigms. The abstract and Section 3 state five fundamental paradigms, including benchmark-based evaluation, but Section 3.5.2 compares only 'four automatic evaluation approaches' and dismisses benchmark-based methods with the sentence 'Benchmark-based evaluation methods do not require meta-evaluation.' The Conclusion then enumerates only four categories ('heuristic-based, embedding-based, learning-based and LLM/VLM/ALM-based evaluation') while claiming five. The authors should either include benchmark-based evaluation in the comparative analysis or revise the taxonomy claim so that the abstract, the body, and the conclusion agree.
- [§5.7, Table 25] The audio-domain comparison provides a direct counterexample to any cross-modal reading of the LLM-superiority claim: on SingMOS, the learning-based RAMP model achieves LCC 0.505 and SRCC 0.480, whereas the ALM-based SALMONN(vic1.5)-Lora achieves only LCC 0.372 and SRCC 0.347. The text acknowledges this drop, but the surrounding discussion still emphasizes ALM-based generalization benefits. If the superiority claim is intended to apply across modalities, this table contradicts it; if it is restricted to the seven NLG benchmarks, the scope should be stated explicitly in the abstract and conclusion.
minor comments (6)
- [Front matter] The CCS Concepts and Additional Key Words fields still contain template placeholder text ('Do Not Use This Code', 'Do, Not, Us, This, Code, Put, the, Correct, Terms, for,Your, Paper') and must be replaced with the correct ACM terms.
- [Throughout] There are numerous typographical errors that should be corrected, including 'attributations' (Section 1), 'CRTLEval'/'CTRLEval' (Table 1 and §3.1.3), 'Scibendi'/'Scribendi' (Table 1), 'Perpleity'/'Perplexity' (Table 1), 'InforLM'/'InfoLM' (Table 2), and 'SaftyBench'/'SafetyBench' (Table 5).
- [Table 5] Reference numbering is inconsistent: 'HumanEval' is cited as [27] in Table 5 but as [29] in the text and reference list, and 'GAOKAO-Bench' is cited as [488] in Table 5 but as [480] in the text. These should be harmonized.
- [§3.3 and §4.2] Several table entries are not discussed in the prose, including BEER, LEIC, RUSE, and SentBLEU in Table 3 and DreamSim in Table 13. Readers would benefit from a sentence for each or an explicit statement that these entries are included only for completeness.
- [§4.7, Table 17] The 'three clear trends' in the text-to-image comparison are based on a single pair of benchmarks with no confidence intervals, sample sizes, or significance tests. Please add this information or soften the wording from 'clear trends' to 'observations'.
- [Tables 24 and 25] Table headers contain spacing/typing issues ('NISQAA VG', 'VoiceMOS-BVCCTest', 'SOMOSTest'), and the Qwen2-Audio-Lora row in Table 25 is entirely empty; the authors should either fill the row or state why it is omitted.
Circularity Check
No circularity: the survey's comparative claims rest on external benchmark numbers, not on the survey's own framework; the only self-citations are illustrative.
full rationale
No load-bearing circular step is present. The paper's central empirical claim—that LLM-based automated evaluation methods currently outperform other approaches significantly (§3.5.2)—is supported by meta-evaluation results over external benchmarks (SummEval, Topical-Chat, FED, WMT-22, OpenMEVA, BAGEL, WebNLG) with human judgments as the reference standard. These numbers are not derived from the authors' five-paradigm taxonomy; the taxonomy is a classification scheme, not an input to the evaluations. The quantitative comparisons in Tables 9, 10, 17, 24, and 25 report correlations or accuracies of independently published methods, and none of the reported quantities is a fitted parameter renamed as a prediction. The inclusion of CriticEval, which may be author-affiliated, does not make the cross-paradigm claim circular: CriticEval appears only in Table 10, where only LLM-based methods are tested, so it is not used to establish the claim that LLM-based methods beat heuristic, embedding-based, or learning-based methods; that claim rests on Table 9. Other self-citations (e.g., HD-Eval, MultiCritique, UniCBE) are used as illustrative examples in qualitative discussion or as related work, not as the evidential basis of a derivation. Concerns about unspecified evaluation protocols, selective missing entries, and heterogeneity of original papers are reproducibility and correctness issues, not circularity. The first-principles content of the survey—the unified five-paradigm organization—is a categorization effort, and its applicability across text, vision, and audio is asserted by classification rather than by a derivation that assumes what it proves. No step reduces to its own inputs by construction.
Assumptions & free parameters
assumptions (2)
- domain assumption The five-paradigm taxonomy (heuristic, embedding, learning, LLM-based, benchmark-based) is an exhaustive and non-overlapping categorization of evaluation methods across text, vision, and speech.
- domain assumption The 12 selected meta-evaluation benchmarks and the specific representative methods in Tables 9 and 10 are sufficient to compare paradigm performance.
Cite this review
Pith. "Pith review of A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations." pith.science (2026). https://pith.science/paper/ERC6QWE6
@misc{pith2026250610019,
author = {Pith},
title = {Pith review of: A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations},
year = {2026},
howpublished = {\url{https://pith.science/paper/ERC6QWE6}},
note = {Machine review of arXiv:2506.10019}
}
read the original abstract
Recent advances in deep learning have significantly enhanced generative AI capabilities across text, images, and audio. However, automatically evaluating the quality of these generated outputs presents ongoing challenges. Although numerous automatic evaluation methods exist, current research lacks a systematic framework that comprehensively organizes these methods across text, visual, and audio modalities. To address this issue, we present a comprehensive review and a unified taxonomy of automatic evaluation methods for generated content across all three modalities; We identify five fundamental paradigms that characterize existing evaluation approaches across these domains. Our analysis begins by examining evaluation methods for text generation, where techniques are most mature. We then extend this framework to image and audio generation, demonstrating its broad applicability. Finally, we discuss promising directions for future research in cross-modal evaluation methodologies.
Figures
Reference graph
Works this paper leans on
-
[1]
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al. 2024. Seed-tts: A family of high-quality versatile speech generation models.arXiv preprint arXiv:2406.02430(2024)
arXiv 2024
-
[2]
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. Spice: Semantic propositional image caption evaluation. InComputer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part V 14. Springer, 382–398
2016
-
[3]
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016. SPICE: Semantic Propositional Image Caption Evaluation. InComputer Vision – ECCV 2016, Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.). Springer International Publishing, Cham, 382–398
2016
-
[4]
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu. 2024. Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words.arXiv preprint arXiv:2406.13340(2024)
arXiv 2024
-
[5]
Siddhant Arora, Zhiyun Lu, Chung-Cheng Chiu, Ruoming Pang, and Shinji Watanabe. 2025. Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics.arXiv preprint arXiv:2503.01174(2025). Manuscript submitted to ACM 52 Lan et al
arXiv 2025
-
[6]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. Program Synthesis with Large Language Models. arXiv:2108.07732 [cs.PL]
arXiv 2021
-
[7]
Kaito Baba, Wataru Nakata, Yuki Saito, and Hiroshi Saruwatari. 2024. The T05 system for the VoiceMOS challenge 2024: Transfer learning from deep image classifier to naturalness MOS prediction of high-quality synthetic speech. In2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 818–824
2024
-
[8]
Sher Badshah and Hassan Sajjad. 2024. Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text. arXiv:2408.09235 [cs.CL] https://arxiv.org/abs/2408.09235
arXiv 2024
Show all 295 references
-
[9]
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems33 (2020), 12449–12460
2020
-
[10]
Ge Bai, Jie Liu, Xingyuan Bu, Yancheng He, Jiaheng Liu, Zhanhui Zhou, Zhuoran Lin, Wenbo Su, Tiezheng Ge, Bo Zheng, et al. 2024. MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.arXiv preprint arXiv:2402.14762(2024)
2024 arXiv
-
[11]
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. 2023. LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.arXiv preprint arXiv:2308.14508(2023)
2023 arXiv
-
[12]
Eslam Mohamed Bakr, Pengzhan Sun, Xiaoqian Shen, Faizan Farooq Khan, Li Erran Li, and Mohamed Elhoseiny. 2023. Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models. InProceedings of the IEEE/CVF International Conference on Computer Vision. 20041–20053
2023
-
[13]
Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Jade Goldstein...
2005
-
[14]
Manik Bhandari, Pranav Narayan Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig. 2020. Re-evaluating Evaluation in Text Summarization. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
-
[15]
Mikołaj Bińkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. 2018. Demystifying mmd gans.arXiv preprint arXiv:1801.01401(2018)
2018 arXiv
-
[16]
Fan Bu, Yuhao Zhang, Xidong Wang, Benyou Wang, Qun Liu, and Haizhou Li. 2024. Roadmap towards superhuman speech understanding using large language models.arXiv preprint arXiv:2410.13268(2024)
2024 arXiv
-
[17]
Jannis Bulian, Christian Buck, Wojciech Gajewski, Benjamin Börschinger, and Tal Schuster. 2022. Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, ...
2022 doi
-
[18]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiaoyi Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang,...
2024 arXiv
-
[19]
Maosong Cao, Alexander Lam, Haodong Duan, Hongwei Liu, Songyang Zhang, and Kai Chen. 2024. CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution. arXiv:2410.16256 [cs.CL] https://arxiv.org/abs/2410.16256
2024 arXiv
-
[20]
Yupeng Cao, Haohang Li, Yangyang Yu, Shashidhar Reddy Javaji, Yueru He, Jimin Huang, Zining Zhu, Qianqian Xie, Xiao-yang Liu, Koduvayur Subbalakshmi, et al. 2025. FinAudio: A Benchmark for Audio Large Language Models in Financial Applications.arXiv preprint arXiv:2503.20990 (2025)
2025
-
[22]
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. arXiv:2308.07201 [cs.CL] https://arxiv.org/abs/2308.07201
2023 arXiv
-
[24]
Boxing Chen and Hongyu Guo. 2015. Representation Based Translation Evaluation Metrics. InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), ...
2015 doi
-
[25]
Chen Chen, Yuchen Hu, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang, Chao-Han Huck Yang, and Eng Siong Chng. 2025. Audio Large Language Models Can Be Descriptive Speech Quality Evaluators.arXiv preprint arXiv:2501.17202(2025). Manuscript submitted to ACM A Survey of Automat...
2025 arXiv
-
[26]
Hong Chen, Duc Vo, Hiroya Takamura, Yusuke Miyao, and Hideki Nakayama. 2022. StoryER: Automatic Story Evaluation via Ranking, Rating and Reasoning. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Y...
2022 doi
-
[27]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...
2021 arXiv
-
[28]
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al
-
[29]
Wang Chen, Piji Li, and Irwin King. 2021. A Training-free and Reference-free Summarization Evaluation Metric via Centrality-weighted Relevance and Self-referenced Redundancy. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th...
2021
-
[30]
Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia. 2023. Theoremqa: A theorem-driven question answering dataset. InThe 2023 Conference on Empirical Methods in Natural Language Processing
2023
-
[31]
Xiaoqiao Chen, Qingyi Zhang, Manhui Lin, Guangyi Yang, and Chu He. 2019. No-reference color image quality assessment: From entropy to perceptual quality.EURASIP Journal on Image and Video Processing2019 (2019), 1–14
2019
-
[32]
Yixiong Chen, Li Liu, and Chris Ding. 2023. X-iqe: explainable image quality evaluation for text-to-image generation with visual large language models.arXiv preprint arXiv:2305.10843(2023)
2023 arXiv
-
[33]
Yiming Chen, Xianghu Yue, Xiaoxue Gao, Chen Zhang, Luis Fernando D’Haro, Robby Tan, and Haizhou Li. 2024. Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models. InFindings of the Association for Computational Linguistics: EMNLP 2024. 10917–10930
2024
-
[34]
Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao, Robby T Tan, and Haizhou Li. 2024. Voicebench: Benchmarking llm-based voice assistants. arXiv preprint arXiv:2410.17196(2024)
2024 arXiv
-
[35]
Zehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, et al
-
[36]
Zhaorun Chen, Yichao Du, Zichen Wen, Yiyang Zhou, Chenhang Cui, Zhenzhen Weng, Haoqin Tu, Chaoqi Wang, Zhengwei Tong, Qinglan Huang, et al. 2024. MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?arXiv preprint arXiv:2407.04842 (2024)
2024 arXiv
-
[37]
Xize Cheng, Ruofan Hu, Xiaoda Yang, Jingyu Lu, Dongjie Fu, Zehan Wang, Shengpeng Ji, Rongjie Huang, Boyang Zhang, Tao Jin, et al . 2025. VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?. InThe Thirteenth International Conference on Learning Representations
2025
-
[38]
Jaemin Cho, Yushi Hu, Roopal Garg, Peter Anderson, Ranjay Krishna, Jason Baldridge, Mohit Bansal, Jordi Pont-Tuset, and Su Wang. 2023. Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-image generation.arXiv preprint arXiv:2310.18235(2023)
2023 arXiv
-
[39]
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. 2019. Sentence Mover’s Similarity: Automatic Evaluation for Multi-Sentence Texts. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Ed...
2019 doi
-
[40]
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge. arXiv:1803.05457 [cs.AI]
2018 arXiv
-
[41]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems.arXiv preprint arXiv:2110.14168 (2021)
2021 arXiv
-
[42]
Roi Cohen, May Hamri, Mor Geva, and Amir Globerson. 2023. LM vs LM: Detecting Factual Errors via Cross Examination. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 12621–12640
2023
-
[43]
Pierre Colombo, Chloe Clave, and Pablo Piantanida. 2021. InfoLM: A New Metric to Evaluate Summarization & Data2Text Generation. InAAAI Conference on Artificial Intelligence. https://api.semanticscholar.org/CorpusID:244896426
2021
-
[44]
Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida. 2021. Automatic Text Evaluation through the Lens of Wasserstein Barycenters. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, ...
2021 doi
-
[45]
OpenCompass Contributors. 2023. OpenCompass: A Universal Evaluation Platform for Foundation Models. https://github.com/open-compass/ opencompass
2023
-
[46]
Erica Cooper, Wen-Chin Huang, Tomoki Toda, and Junichi Yamagishi. 2022. Generalization ability of MOS prediction networks. InICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 8442–8446
2022
-
[47]
Erica Cooper, Wen-Chin Huang, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. 2023. The VoiceMOS Challenge 2023: Zero-shot subjective speech quality prediction for multiple domains. In2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 1–7
2023
-
[48]
Liam Cripwell, Joël Legrand, and Claire Gardent. 2023. Simplicity Level Estimate (SLE): A Learned Reference-Less Metric for Sentence Simplification. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda Bouamor, Juan Pino, and Kalika B...
2023 doi
-
[49]
Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. UltraFeedback: Boosting Language Models with High-quality Feedback. arXiv:2310.01377 [cs.CL]
2023 arXiv
-
[50]
Wenqian Cui, Xiaoqi Jiao, Ziqiao Meng, and Irwin King. 2025. VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models.arXiv preprint arXiv:2501.04962(2025)
2025 arXiv
-
[51]
Yin Cui, Guandao Yang, Andreas Veit, Xun Huang, and Serge Belongie. 2018. Learning to Evaluate Image Captioning. arXiv:1806.06422 [cs.CV] https://arxiv.org/abs/1806.06422
2018 arXiv
-
[52]
Fredrik Cumlin, Xinyu Liang, Victor Ungureanu, Chandan KA Reddy, Christian Schüldt, and Saikat Chatterjee. 2025. Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality Models. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and S...
2025
-
[53]
DeepSeek-AI, Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, Bing Xue, Bingxuan Wang, Bochao Wu, Bei F...
2025 arXiv
-
[54]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2024 arXiv
-
[55]
Soham Deshmukh, Dareen Alharthi, Benjamin Elizalde, Hannes Gamper, Mahmoud Al Ismail, Rita Singh, Bhiksha Raj, and Huaming Wang. 2024. Pam: Prompting audio-language models for audio quality assessment.arXiv preprint arXiv:2402.00282(2024)
2024 arXiv
-
[56]
Soham Deshmukh, Shuo Han, Hazim Bukhari, Benjamin Elizalde, Hannes Gamper, Rita Singh, and Bhiksha Raj. 2025. Audio Entailment: Assessing deductive reasoning for audio understanding. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 23769–23777
2025
-
[57]
Soham Deshmukh, Shuo Han, Rita Singh, and Bhiksha Raj. 2025. ADIFF: Explaining audio difference using natural language.arXiv preprint arXiv:2502.04476(2025)
2025 arXiv
-
[58]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019 doi
-
[59]
Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William Cohen. 2019. Handling Divergent Reference Texts when Evaluating Table-to-Text Generation. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Kor...
2019 doi
-
[60]
George Doddington. 2002. Automatic evaluation of machine translation quality using n-gram co-occurrence statistics. InProceedings of the Second International Conference on Human Language Technology Research(San Diego, California)(HLT ’02). Morgan Kaufmann Publishers Inc., San ...
2002
-
[61]
Xuan Dong and Donald S Williamson. 2020. A Pyramid Recurrent Network for Predicting Crowdsourced Speech-Quality Ratings of Real-World Signals. (2020)
2020
-
[62]
Esin Durmus, He He, and Mona Diab. 2020. FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schlu...
2020 doi
-
[63]
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Nikol...
2024
-
[64]
Alexander Richard Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. SummEval: Re- evaluating Summarization Evaluation.Transactions of the Association for Computational Linguistics9 (2021), 391–409
2021
-
[65]
Alexander Richard Fabbri, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022. QAFactEval: Improved QA-Based Factual Consistency Evaluation for Summarization. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: H...
2022
-
[66]
Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Songyang Zhang, Kai Chen, Zongwen Shen, and Jidong Ge. 2023. LawBench: Benchmarking Legal Knowledge of Large Language Models. arXiv:2309.16289 [cs.CL] https://arxiv.org/abs/2309.16289
2023 arXiv
-
[67]
Tiantian Feng, Jihwan Lee, Anfeng Xu, Yoonjeong Lee, Thanathai Lertpetchpun, Xuan Shi, Helin Wang, Thomas Thebaud, Laureano Moro-Velazquez, Dani Byrd, et al. 2025. Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits.arXiv prepr...
2025 arXiv
-
[68]
Weixi Feng, Jiachen Li, Michael Saxon, Tsu-jui Fu, Wenhu Chen, and William Yang Wang. 2024. TC-Bench: Benchmarking Temporal Compositionality in Text-to-Video and Image-to-Video Generation.arXiv preprint arXiv:2406.08656(2024)
2024 arXiv
-
[69]
Cédric Févotte, Rémi Gribonval, and Emmanuel Vincent. 2005. BSS_EVAL Toolbox User Guide–Revision 2.0. (2005)
2005
-
[70]
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022. Results of WMT22 Metrics Shared Task: Stop Using BLEU – Neural Metrics Are Better and More Robust. InProceedings o...
2022
-
[71]
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. GPTScore: Evaluate as You Desire. arXiv:2302.04166 [cs.CL] https://arxiv.org/abs/ 2302.04166
2023 arXiv
-
[72]
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. 2023. DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data.Advances in Neural Information Processing Systems36 (2023), 50742–50768
2023
-
[73]
Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, and Hsin-Min Wang. 2018. Quality-Net: An end-to-end non-intrusive speech quality assessment model based on BLSTM.arXiv preprint arXiv:1808.05344(2018)
2018 arXiv
-
[74]
Bofei Gao, Zefan Cai, Runxin Xu, Peiyi Wang, Ce Zheng, Runji Lin, Keming Lu, Junyang Lin, Chang Zhou, Wen Xiao, et al. 2024. LLM critics help catch bugs in mathematics: Towards a better mathematical verifier with natural language feedback.CoRR(2024)
2024
-
[75]
Kuofeng Gao, Shu-Tao Xia, Ke Xu, Philip Torr, and Jindong Gu. 2024. Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models.arXiv preprint arXiv:2412.05167(2024). Manuscript submitted to ACM 56 Lan et al
2024 arXiv
-
[76]
Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020. Dialogue Response Ranking Training with Large-Scale Human Feedback Data. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Y...
2020 doi
-
[77]
Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and Bill Dolan. 2020. Dialogue Response RankingTraining with Large-Scale Human Feedback Data. InEMNLP
2020
-
[78]
Yiming Gao, Bin Wang, Chengwei Wei, Shuo Sun, and AiTi Aw. 2025. IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models.arXiv preprint arXiv:2505.16774(2025)
2025
-
[79]
Yang Gao, Wei Zhao, and Steffen Eger. 2020. SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summa- rization. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schlute...
2020 doi
-
[80]
Zorik Gekhman, Jonathan Herzig, Roee Aharoni, Chen Elkind, and Idan Szpektor. 2023. TrueTeacher: Learning Factual Consistency Evaluation with Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2053–2070
2023
-
[81]
Sarik Ghazarian, Johnny Wei, Aram Galstyan, and Nanyun Peng. 2019. Better Automatic Evaluation of Open-Domain Dialogue Systems with Contextualized Embeddings. InProceedings of the Workshop on Methods for Optimizing and Evaluating Neural Language Generation, Antoine Bosselut, A...
2019 doi
-
[82]
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 15180–15190
2023
-
[83]
Tejas Gokhale, Hamid Palangi, Besmira Nushi, Vibhav Vineet, Eric Horvitz, Ece Kamar, Chitta Baral, and Yezhou Yang. 2022. Benchmarking spatial relationships in text-to-image generation.arXiv preprint arXiv:2212.10015(2022)
2022 arXiv
-
[84]
Paul Grimal, Hervé Le Borgne, Olivier Ferret, and Julien Tourille. 2024. TIAM-A metric for evaluating alignment in Text-to-Image generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2890–2899
2024
-
[85]
Max Grusky, Mor Naaman, and Yoav Artzi. 2018. Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies. InProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volu...
2018 doi
-
[86]
Yuzhe Gu, Ziwei Ji, Wenwei Zhang, Chengqi Lyu, Dahua Lin, and Kai Chen. 2024. ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models.arXiv preprint arXiv:2407.04693(2024)
2024 arXiv
-
[87]
Jian Guan and Minlie Huang. 2020. UNION: An Unreferenced Metric for Evaluating Open-ended Story Generation. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association ...
2020 doi
-
[88]
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021. OpenMEVA: A Benchmark for Evaluating Open-ended Story Generation Metrics. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics ...
2021
-
[89]
Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou. 2024. Adapting frechet audio distance for generative music evaluation. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1331–1335
2024
-
[90]
Yinuo Guo, Chong Ruan, and Junfeng Hu. 2018. Meteor++: Incorporating Copy Knowledge into Machine Translation Evaluation. InProceedings of the Third Conference on Machine Translation: Shared Task Papers, Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Mark Fishel, Yvette G...
2018
-
[91]
Rohit Gupta, Constantin Orăsan, and Josef van Genabith. 2015. ReVal: A Simple and Effective Machine Translation Evaluation Metric Based on Recurrent Neural Networks. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lluís Màrquez, Chris ...
2015 doi
-
[92]
Francisco Guzmán, Shafiq Joty, Lluís Màrquez, Alessandro Moschitti, Preslav Nakov, and Massimo Nicosia. 2014. Learning to Differentiate Better from Worse Translations. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Alessandro ...
2014 doi
-
[93]
Francisco Guzmán, Shafiq Joty, Lluís Màrquez, and Preslav Nakov. 2015. Pairwise Neural Machine Translation Evaluation. InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Proce...
2015 doi
-
[94]
Seungju Han, Beomsu Kim, and Buru Chang. 2022. Measuring and Improving Semantic Diversity of Dialogue Generation. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Ling...
2022 doi
-
[95]
Hardy Hardy, Shashi Narayan, and Andreas Vlachos. 2019. HighRES: Highlight-based Reference-less Evaluation of Summarization. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Associa...
2019 doi
-
[96]
Hosein Hasanbeig, Hiteshi Sharma, Leo Betthauser, Felipe Vieira Frujeri, and Ida Momennejad. 2023. ALLURE: Auditing and Improving LLM-based Evaluation of Text using Iterative In-Context-Learning. arXiv:2309.13701 [cs.CL] https://arxiv.org/abs/2309.13701
2023 arXiv
-
[97]
Tomoki Hayashi, Ryuichi Yamamoto, Takenori Yoshimura, Peter Wu, Jiatong Shi, Takaaki Saeki, Yooncheol Ju, Yusuke Yasuda, Shinnosuke Takamichi, and Shinji Watanabe. 2021. Espnet2-tts: Extending the edge of tts research.arXiv preprint arXiv:2110.07840(2021)
2021 arXiv
-
[98]
Yuze He, Yushi Bai, Matthieu Lin, Wang Zhao, Yubin Hu, Jenny Sheng, Ran Yi, Juanzi Li, and Yong-Jin Liu. 2023. T3 Bench: Benchmarking Current Progress in Text-to-3D Generation.arXiv preprint arXiv:2310.02977(2023)
2023 arXiv
-
[99]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding.Proceedings of the International Conference on Learning Representations (ICLR)(2021)
2021
-
[100]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring Mathematical Problem Solving With the MATH Dataset.NeurIPS(2021)
2021
-
[101]
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. 2017. CNN architectures for large-scale audio classification. In2017 ieee international conference on acoustics, sp...
2017
-
[102]
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718(2021)
2021 arXiv
-
[103]
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems30 (2017)
2017
-
[104]
Tobias Hinz, Stefan Heinrich, and Stefan Wermter. 2020. Semantic object accuracy for generative text-to-image synthesis.IEEE transactions on pattern analysis and machine intelligence44, 3 (2020), 1552–1565
2020
-
[105]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. arXiv:2006.11239 [cs.LG] https://arxiv.org/abs/2006.11239
2020 arXiv
-
[106]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. arXiv:1904.09751 [cs.CL] https://arxiv.org/abs/1904.09751
2020 arXiv
-
[107]
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021. 𝑄 2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering. InProceedings of the 2021 Conference on Empirical Methods in Natural...
2021
-
[108]
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing...
2021
-
[109]
Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024. Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Ban...
2024 doi
-
[110]
Xinyu Hu, Li Lin, Mingqi Gao, Xunjian Yin, and Xiaojun Wan. 2024. Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability. arXiv:2406.18365 [cs.CL] https://arxiv.org/abs/2406.18365
2024 arXiv
-
[111]
Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A Smith. 2023. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2...
2023
-
[112]
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang, Andy T Liu, Chen-An Li, Yu-Xiang Lin, Wei-Cheng Tseng, Anuj Diwan, Yi-Jen Shih, Jiatong Shi, et al. 2024. Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 18...
2024 arXiv
-
[113]
Chien-yu Huang, Min-Han Shih, Ke-Han Lu, Chi-Yuan Hsiao, and Hung-yi Lee. 2025. SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Process...
2025
-
[114]
Kexin Huang, Xiangyang Liu, Qianyu Guo, Tianxiang Sun, Jiawei Sun, Yaru Wang, Zeyang Zhou, Yixu Wang, Yan Teng, Xipeng Qiu, Yingchun Wang, and Dahua Lin. 2023. Flames: Benchmarking Value Alignment of Chinese Large Language Models. arXiv:2311.06899 [cs.CL]
2023 arXiv
-
[115]
Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. 2023. T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.Advances in Neural Information Processing Systems36 (2023), 78723–78747
2023
-
[116]
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020. GRADE: Automatic Graph-Enhanced Coherence Metric for Evaluating Open-Domain Dialogue Systems. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webbe...
2020 doi
-
[117]
Wen-Chin Huang, Erica Cooper, and Tomoki Toda. 2024. Mos-bench: Benchmarking generalization abilities of subjective speech quality assessment models.arXiv preprint arXiv:2411.03715(2024). Manuscript submitted to ACM 58 Lan et al
2024 arXiv
-
[118]
Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. 2022. The voicemos challenge 2022.arXiv preprint arXiv:2203.11389(2022)
2022 arXiv
-
[119]
Wen-Chin Huang, Erica Cooper, Junichi Yamagishi, and Tomoki Toda. 2022. Ldnet: Unified listener dependent modeling in mos prediction for synthetic speech. InICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 896–900
2022
-
[120]
Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao. 2024. The VoiceMOS challenge 2024: Beyond speech quality prediction. In2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 803–810
2024
-
[121]
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. 2023. C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.arXiv preprint ar...
2023 arXiv
-
[122]
Md Asadul Islam and Enrico Magnani. 2021. Is this the end of the gold standard? A straightforward reference-less grammatical error correction metric. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang,...
2021 doi
-
[123]
Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023. Multi-Dimensional Evaluation of Text Summarization with In-Context Learning. InFindings of the Association for Computational Linguistics: ACL 2...
2023 doi
-
[124]
Wissam A Jassim, Jan Skoglund, Michael Chinen, and Andrew Hines. 2021. WARP-Q: Quality prediction for generative neural speech codecs. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 401–405
2021
-
[125]
Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, and Sanjiv Kumar. 2024. Rethinking fid: Towards a better evaluation metric for image generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9307–9315
2024
-
[126]
Mercer, Lalit R
Frederick Jelinek, Robert L. Mercer, Lalit R. Bahl, and Janet M. Baker. 1977. Perplexity—a measure of the difficulty of speech recognition tasks. Journal of the Acoustical Society of America62 (1977). https://api.semanticscholar.org/CorpusID:121680873
1977
-
[127]
Jesper Jensen and Cees H Taal. 2016. An algorithm for predicting the intelligibility of speech masked by modulated noise maskers.IEEE/ACM Transactions on Audio, Speech, and Language Processing24, 11 (2016), 2009–2022
2016
-
[128]
Ziwei Ji, Yuzhe Gu, Wenwei Zhang, Chengqi Lyu, Dahua Lin, and Kai Chen. 2024. ANAH: Analytical Annotation of Hallucinations in Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, ...
2024 doi
-
[129]
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2024. TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks. arXiv:2310.00752 [cs.CL] https://arxiv.org/abs/2310.00752
2024 arXiv
-
[130]
Feng Jiang, Zhiyu Lin, Fan Bu, Yuhao Du, Benyou Wang, and Haizhou Li. 2025. S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following with Paralinguistic Information.arXiv preprint arXiv:2503.05085(2025)
2025 arXiv
-
[131]
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao. 2019. Tiger: Text-to-image grounding for image caption evaluation.arXiv preprint arXiv:1909.02050(2019)
2019 arXiv
-
[132]
Christian Johnson. 2022. Binary Encoded Word Mover‘s Distance. InProceedings of the 7th Workshop on Representation Learning for NLP, Spandana Gella, He He, Bodhisattwa Prasad Majumder, Burcu Can, Eleonora Giunchiglia, Samuel Cahyawijaya, Sewon Min, Maximilian Mozes, Xiang Lorr...
2022 doi
-
[133]
Weld, and Luke Zettlemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Association for Computati...
2017
-
[134]
Moussa Kamal Eddine, Guokan Shang, Antoine Tixier, and Michalis Vazirgiannis. 2022. FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metrics for Automatic Text Generation. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vo...
2022 doi
-
[135]
Mintong Kang, Chejian Xu, and Bo Li. 2024. AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models.arXiv preprint arXiv:2412.08608(2024)
2024 arXiv
-
[136]
James M Kates and Kathryn H Arehart. 2010. The hearing-aid speech quality index (HASQI).Journal of the Audio Engineering Society58, 5 (2010), 363–381
2010
-
[137]
James M Kates and Kathryn H Arehart. 2014. The hearing-aid speech perception index (HASPI).Speech Communication65 (2014), 75–93
2014
-
[138]
Pei Ke, Bosi Wen, Andrew Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al. 2024. Critiquellm: Towards an informative critique generation model for evaluation of large language model generation. InProceedings of the 62nd Ann...
2024
-
[139]
Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2024. CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation. arXiv:23...
2024 arXiv
-
[140]
Pei Ke, Hao Zhou, Yankai Lin, Peng Li, Jie Zhou, Xiaoyan Zhu, and Minlie Huang. 2022. CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text Generation. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1:...
2022 doi
-
[141]
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. 2018. Fr\’echet audio distance: A metric for evaluating music enhancement algorithms.arXiv preprint arXiv:1812.08466(2018)
2018 arXiv
-
[142]
Heeseung Kim, Che Hyun Lee, Sangkwon Park, Jiheum Yeom, Nohil Park, Sangwon Yu, and Sungroh Yoon. 2025. Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models.arXiv preprint arXiv:2502.19759(2025)
2025 arXiv
-
[143]
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. InThe Twelfth International Conference on Learning Repre...
2024
-
[144]
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2024. Prometheus: Inducing Fine-grained Evaluation Capability in Language Models. arXiv:2310.08491 [cs.CL] https: //arxiv.org/ab...
2024 arXiv
-
[145]
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. InProceedings of the 2024 Confere...
2024 doi
-
[146]
Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models. arXiv:2405.01535 [cs.CL] https://...
2024 arXiv
-
[147]
Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. 2023. Pick-a-pic: An open dataset of user preferences for text-to-image generation.Advances in Neural Information Processing Systems36 (2023), 36652–36663
2023
-
[148]
Neema Kotonya, Saran Krishnasamy, Joel Tetreault, and Alejandro Jaimes. 2023. Little Giants: Exploring the Potential of Small LLMs as Evaluation Metrics in Summarization in the Eval4NLP 2023 Shared Task. InProceedings of the 4th Workshop on Evaluation and Comparison of NLP Sys...
2023 doi
-
[149]
Max Ku, Dongfu Jiang, Cong Wei, Xiang Yue, and Wenhu Chen. 2023. Viescore: Towards explainable metrics for conditional image synthesis evaluation.arXiv preprint arXiv:2312.14867(2023)
2023 arXiv
-
[150]
Max Ku, Tianle Li, Kai Zhang, Yujie Lu, Xingyu Fu, Wenwen Zhuang, and Wenhu Chen. 2023. Imagenhub: Standardizing the evaluation of conditional image generation models.arXiv preprint arXiv:2310.01596(2023)
2023 arXiv
-
[151]
Chun-Yi Kuan, Wei-Ping Huang, and Hung-yi Lee. 2024. Understanding sounds, missing the questions: The challenge of object hallucination in large audio-language models.arXiv preprint arXiv:2406.08402(2024)
2024 arXiv
-
[152]
Chun-Yi Kuan and Hung-yi Lee. 2025. Can large audio-language models truly hear? Tackling hallucinations with multi-task assessment and stepwise audio reasoning. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2025
-
[153]
Robert Kubichek. 1993. Mel-cepstral distance measure for objective speech quality assessment. InProceedings of IEEE pacific rim conference on communications computers and signal processing, Vol. 1. IEEE, 125–128
1993
-
[154]
Anurag Kumar, Ke Tan, Zhaoheng Ni, Pranay Manocha, Xiaohui Zhang, Ethan Henderson, and Buye Xu. 2023. Torchaudio-squim: Reference-less speech quality and intelligibility measures in torchaudio. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Pr...
2023
-
[155]
Marie Kunešová, Jindřich Matoušek, Jan Lehečka, Jan Švec, Josef Michálek, Daniel Tihelka, Martin Bulín, Zdeněk Hanzlíček, and Markéta Řezáčková
-
[156]
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev. 2024. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack. arXiv:2406.10149
2024 arXiv
-
[157]
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015. From Word Embeddings To Document Distances. InProceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37), Francis Bach and David Blei (Eds.). PMLR, ...
2015
-
[158]
InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Ensemble of deep neural network models for MOS prediction. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2023
-
[159]
Tuomas Kynkäänniemi, Tero Karras, Samuli Laine, Jaakko Lehtinen, and Timo Aila. 2019. Improved precision and recall metric for assessing generative models.Advances in neural information processing systems32 (2019)
2019
-
[160]
Bennett, and Marti A
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization.Transactions of the Association for Computational Linguistics10 (2022), 163–177. Manuscript submitted to ACM 60 Lan et al
2022
-
[161]
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019. Natural questions: a benchmark for question answering research.Transactions of the Association for Co...
2019
-
[162]
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017. Race: Large-scale reading comprehension dataset from examinations. arXiv preprint arXiv:1704.04683(2017)
2017 arXiv
-
[163]
Sanyam Lakhanpal, Shivang Chopra, Vinija Jain, Aman Chadha, and Man Luo. 2024. Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation.arXiv preprint arXiv:2403.16422(2024)
2024 arXiv
-
[164]
Bennett, and Marti A
Philippe Laban, Tobias Schnabel, Paul N. Bennett, and Marti A. Hearst. 2022. SummaC: Re-Visiting NLI-based Models for Inconsistency Detection in Summarization.Transactions of the Association for Computational Linguistics10 (2022), 163–177. https://doi.org/10.1162/tacl_a_00453
2022 doi
-
[165]
Guillaume Lample and Alexis Conneau. 2019. Cross-lingual Language Model Pretraining. arXiv:1901.07291 [cs.CL]
2019 arXiv
-
[166]
Tian Lan, Deng Cai, Yan Wang, Yixuan Su, Heyan Huang, and Xian-Ling Mao. 2024. Exploring Dense Retrieval for Dialogue Response Selection. ACM Trans. Inf. Syst.42, 3, Article 84 (Jan. 2024), 29 pages. https://doi.org/10.1145/3632750
2024 doi
-
[167]
Smith, and Hannaneh Hajishirzi
Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi. 2024. RewardBench: Evaluating Reward Models for Language Modeling. arXiv:2403.13787 [cs.LG] ...
2024 arXiv
-
[169]
Tian Lan, Wenwei Zhang, Chengqi Lyu, Shuaibin Li, Chen Xu, Heyan Huang, Dahua Lin, Xian-Ling Mao, and Kai Chen. 2024. Training Language Models to Critique With Multi-agent Feedback. arXiv:2410.15287 [cs.CL] https://arxiv.org/abs/2410.15287
2024 arXiv
-
[170]
Tian Lan, Xian-Ling Mao, Wei Wei, Xiaoyan Gao, and Heyan Huang. 2020. PONE: A Novel Automatic Evaluation Metric for Open-Domain Generative Dialogue Systems. arXiv:2004.02399 [cs.CL] https://arxiv.org/abs/2004.02399
2020 arXiv
-
[171]
Teven Le Scao and Claire Gardent. 2023. Joint Representations of Text and Knowledge Graphs for Retrieval and Evaluation. InFindings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings), Jong C. Park, Yuki Arase, Baotian Hu, Wei Lu, Derry Wijaya, Ayu Pu...
2023 doi
-
[172]
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Ras- togi, and Sushant Prakash. 2024. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. arXiv:2309.00267...
2024 arXiv
-
[173]
Tian Lan, Wenwei Zhang, Chen Xu, Heyan Huang, Dahua Lin, Kai Chen, and Xian ling Mao. 2024. CriticEval: Evaluating Large Language Model as Critic. arXiv:2402.13764 [cs.CL] https://arxiv.org/abs/2402.13764
2024 arXiv
-
[174]
Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, and Youngjae Yu. 2024. Aligning Large Language Models by On-Policy Self-Judgment. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei...
2024 doi
-
[175]
Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Minjoon Seo, Kang Min Yoo, and Youngjae Yu. 2024. Aligning Large Language Models by On-Policy Self-Judgment. arXiv:2402.11253 [cs.LG] https://arxiv.org/abs/2402.11253
2024 arXiv
-
[176]
Jinhyuk Lee, Feiyang Chen, Sahil Dua, Daniel Cer, Madhuri Shanbhogue, Iftekhar Naim, Gustavo Hernández Ábrego, Zhe Li, Kaifeng Chen, Henrique Schechter Vera, et al. 2025. Gemini embedding: Generalizable embeddings from gemini.arXiv preprint arXiv:2503.07891(2025)
2025 arXiv
-
[177]
Yichong Leng, Xu Tan, Sheng Zhao, Frank Soong, Xiang-Yang Li, and Tao Qin. 2021. MBNet: MOS prediction for synthesized speech with mean-bias network. InICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 391–395
2021
-
[178]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer
-
[179]
Yang Lei, Jiangtong Li, Dawei Cheng, Zhijun Ding, and Changjun Jiang. 2024. CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model. arXiv:2311.05812 [cs.CL] https://arxiv.org/abs/2311.05812
2024 arXiv
-
[180]
Dawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi, Chengshuai Zhao, Zhen Tan, Amrita Bhattacharjee, Yuxuan Jiang, Canyu Chen, Tianhao Wu, Kai Shu, Lu Cheng, and Huan Liu. 2024. From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. arXiv:2411.16594...
2024
-
[181]
Huayang Li, Tian Lan, Zihao Fu, Deng Cai, Lemao Liu, Nigel Collier, Taro Watanabe, and Yixuan Su. 2023. Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective. InThirty-seventh Conference on Neural Information Processing Systems....
2023
-
[182]
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023. CMMLU: Measuring massive multitask language understanding in Chinese. arXiv:2306.09212 [cs.CL]
2023 arXiv
-
[183]
Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Emily Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. 2024. GenAI-Bench: A Holistic Benchmark for Compositional Text-to-Visual Generation. InSynthetic Data for Computer Vision Workshop@ CVPR 2024
2024
-
[184]
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A Diversity-Promoting Objective Function for Neural Conversation Models. InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Lang...
2016 doi
-
[185]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. arXiv:2301.12597 [cs.CV] https://arxiv.org/abs/2301.12597
2023 arXiv
-
[186]
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Pengfei Liu, et al. 2024. Generative Judge for Evaluating Alignment. InThe Twelfth International Conference on Learning Representations
2024
-
[187]
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. Halueval: A large-scale hallucination evaluation benchmark for large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 6449–6464. Manuscript s...
2023
-
[188]
Junlong Li, Fan Zhou, Shichao Sun, Yikai Zhang, Hai Zhao, and Pengfei Liu. 2024. Dissecting Human and LLM Preferences. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikum...
2024 doi
-
[189]
Kai Li, Can Shen, Yile Liu, Jirui Han, Kelong Zheng, Xuechao Zou, Zhe Wang, Xingjian Du, Shun Zhang, Hanjun Luo, et al. 2025. AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models.arXiv preprint arXiv:2505.16211(2025)
2025
-
[190]
Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024. SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models. InFindings of the Association for Computational Linguistics ACL 2024, Lun-Wei Ku, Andre...
2024
-
[191]
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. 2023. Generative Judge for Evaluating Alignment. arXiv:2310.05470 [cs.CL] https://arxiv.org/abs/2310.05470
2023 arXiv
-
[192]
Ruosen Li, Teerth Patel, and Xinya Du. 2024. PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations. arXiv:2307.02762 [cs.CL] https://arxiv.org/abs/2307.02762
2024 arXiv
-
[193]
Zitong Li and Wei Li. 2023. MOSLight: A Lightweight Data-Efficient System for Non-Intrusive Speech Quality Assessment. InProc. INTERSPEECH, Vol. 2023. 5386–5390
2023
-
[194]
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, Yuxuan Lai, Chongyang Tao, and Shuai Ma. 2024. Leveraging Large Language Models for NLG Evaluation: Advances and Challenges. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-O...
2024 doi
-
[195]
Rongao Li, Jie Fu, Bo-Wen Zhang, Tao Huang, Zhihong Sun, Chen Lyu, Guang Liu, Zhi Jin, and Ge Li. 2023. TACO: Topics in Algorithmic COde generation dataset.arXiv preprint arXiv:2312.14852(2023)
2023 arXiv
-
[196]
Qiao Liang, Ying Shen, Tiantian Chen, Lin Zhang, and Shengjie Zhao. 2025. ADTMOS–Synthesized Speech Quality Assessment Based on Audio Distortion Tokens.IEEE Transactions on Audio, Speech and Language Processing(2025)
2025
-
[197]
Xinyu Liang, Fredrik Cumlin, Christian Schüldt, and Saikat Chatterjee. 2023. DeePMOS: deep posterior mean-opinion-score of speech. InProceedings of INTERSPEECH. 526–530
2023
-
[198]
Xun Liang, Shichao Song, Zifan Zheng, Hanyu Wang, Qingchen Yu, Xunkai Li, Rong-Hua Li, Yi Wang, Zhonghao Wang, Feiyu Xiong, and Zhiyu Li
-
[199]
Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou. 2021. Conversations Are Not Flat: Modeling the Dynamic Information Flow across Dialogue Utterances. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat...
2021
-
[200]
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe
-
[201]
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning. InFindings of the Association for Computational Linguistics: EMNLP 2020, Trevo...
2020 doi
-
[202]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. InText Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013
2004
-
[203]
Guan-Ting Lin, Cheng-Han Chiang, and Hung-Yi Lee. 2024. Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pap...
2024
-
[206]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arXiv:2109.07958 [cs.CL] https://arxiv.org/abs/2109.07958
2022 arXiv
-
[207]
arXiv:2305.20050 [cs.LG] https://arxiv.org/abs/2305.20050
Let’s Verify Step by Step. arXiv:2305.20050 [cs.LG] https://arxiv.org/abs/2305.20050
-
[208]
Yu-Xiang Lin, Chih-Kai Yang, Wei-Chih Chen, Chen-An Li, Chien-yu Huang, Xuanjun Chen, and Hung-yi Lee. 2025. A Preliminary Exploration with GPT-4o Voice Mode.arXiv preprint arXiv:2502.09940(2025)
2025 arXiv
-
[209]
Zicheng Lin, Zhibin Gou, Tian Liang, Ruilin Luo, Haowei Liu, and Yujiu Yang. 2024. CriticBench: Benchmarking LLMs for Critique-Correct Reasoning. InFindings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Associ...
2024 doi
-
[210]
Zhiqiu Lin, Deepak Pathak, Baiqi Li, Jiayao Li, Xide Xia, Graham Neubig, Pengchuan Zhang, and Deva Ramanan. 2025. Evaluating text-to-visual generation with image-to-text generation. InEuropean Conference on Computer Vision. Springer, 366–384
2025
-
[211]
Guan-Ting Lin, Jiachen Lian, Tingle Li, Qirui Wang, Gopala Anumanchipalli, Alexander H Liu, and Hung-yi Lee. 2025. Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.arXiv preprint arXiv:2503.04721(2025)
2025 arXiv
-
[212]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring How Models Mimic Human Falsehoods. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 3214–3252. Manuscript submitted to ACM 62 Lan et al
2022
-
[213]
Chris Yuhao Liu, Liang Zeng, Jiacai Liu, Rui Yan, Jujie He, Chaojie Wang, Shuicheng Yan, Yang Liu, and Yahui Zhou. 2024. Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.arXiv preprint arXiv:2410.18451(2024)
2024 arXiv
-
[214]
Yen-Ting Lin and Yun-Nung Chen. 2023. LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. InProceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023), Yun-Nung Chen and Abhinav Rastogi (Eds.). ...
2023 doi
-
[215]
Mianxin Liu, Jinru Ding, Jie Xu, Weiguo Hu, Xiaoyang Li, Lifeng Zhu, Zhian Bai, Xiaoming Shi, Benyou Wang, Haitao Song, Pengfei Liu, Xiaofan Zhang, Shanshan Wang, Kang Li, Haofen Wang, Tong Ruan, Xuanjing Huang, Xin Sun, and Shaoting Zhang. 2024. MedBench: A Comprehensive, Sta...
2024 arXiv
-
[216]
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, and Lifu Huang. 2024. X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects. InProceedings of the 2024 Conference of the Nort...
2024
-
[217]
Minqian Liu, Ying Shen, Zhiyang Xu, Yixin Cao, Eunah Cho, Vaibhav Kumar, Reza Ghanadan, and Lifu Huang. 2024. X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects. arXiv:2311.08788 [cs.CL] https://arxiv.org/ abs/...
2024 arXiv
-
[218]
Cheng Liu, Hui Wang, Jinghua Zhao, Shiwan Zhao, Hui Bu, Xin Xu, Jiaming Zhou, Haoqin Sun, and Yong Qin. 2025. MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech...
2025
-
[219]
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. InProceedings of the 2016 Conference on Empirica...
2016 doi
-
[220]
Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. 2024. Evalcrafter: Benchmarking and evaluating large video generation models. InProceedings of the IEEE/CVF Conference on Computer Vision and Patter...
2024
-
[221]
Ganjun Liu, Xiaohui Hou, Meng Ge, Tao Zhang, and Haizhou Li. 2024. A non-intrusive approach to assessing dysarthria severity: Advancing clinical diagnosis. InCompanion Proceedings of the ACM Web Conference 2024. 1134–1137
2024
-
[222]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. arXiv:2303.16634 [cs.CL] https://arxiv.org/abs/2303.16634
2023 arXiv
-
[223]
Yuanxin Liu, Lei Li, Shuhuai Ren, Rundong Gao, Shicheng Li, Sishuo Chen, Xu Sun, and Lu Hou. 2024. Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation.Advances in Neural Information Processing Systems36 (2024)
2024
-
[224]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Ro{BERT}a: A Robustly Optimized {BERT} Pretraining Approach. https://openreview.net/forum?id=SyxS0T4tvS
2020
-
[225]
Siyang Liu, Sahand Sabour, Yinhe Zheng, Pei Ke, Xiaoyan Zhu, and Minlie Huang. 2022. Rethinking and Refining the Distinct Metric. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Smaranda Muresan, Preslav Nakov...
2022 doi
-
[226]
Xiao Liu, Xuanyu Lei, Shengyuan Wang, Yue Huang, Zhuoer Feng, Bosi Wen, Jiale Cheng, Pei Ke, Yifan Xu, Weng Lam Tam, Xiaohan Zhang, Lichao Sun, Xiaotao Gu, Hongning Wang, Jing Zhang, Minlie Huang, Yuxiao Dong, and Jie Tang. 2024. AlignBench: Benchmarking Chinese Alignment of L...
2024 arXiv
-
[227]
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. 2024. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models. arXiv:2402.17177 [cs.CV] h...
2024 arXiv
-
[228]
Yixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, and Dragomir Radev. 2023. Towards Interpretable and Efficient Automatic Reference-Based Summarization Evaluation. InProceedings of the 2023 Conference on Empirical Methods in Natu...
2023 doi
-
[229]
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang. 2019. Mosnet: Deep learning based objective assessment for voice conversion.arXiv preprint arXiv:1904.08352(2019)
2019 arXiv
-
[230]
Chi-kiu Lo. 2017. MEANT 2.0: Accurate semantic MT evaluation for any output language. InProceedings of the Second Conference on Machine Translation, Ondřej Bojar, Christian Buck, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno ...
2017
-
[231]
Chi-kiu Lo. 2019. YiSi - a Unified Semantic MT Quality Evaluation and Estimation Metric for Languages with Different Levels of Available Resources. InProceedings of the Fourth Conference on Machine Translation (Volume 2: Shared Task Papers, Day 1), Ondřej Bojar, Rajen Chatterj...
2019
-
[232]
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024. HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition. arXiv:2402.15754 [cs.CL] https://arxiv.org/abs/2402.15754
2024 arXiv
-
[233]
Yantao Liu, Zijun Yao, Rui Min, Yixin Cao, Lei Hou, and Juanzi Li. 2024. RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style. arXiv:2410.16184 [cs.CL] https://arxiv.org/abs/2410.16184 Manuscript submitted to ACM A Survey of Automatic Evaluation Meth...
2024 arXiv
-
[234]
Ke-Han Lu, Chun-Yi Kuan, and Hung-yi Lee. 2025. Speech-ifeval: Evaluating instruction-following and quantifying catastrophic forgetting in speech-aware language models.arXiv preprint arXiv:2505.19037(2025)
2025 arXiv
-
[235]
Zijun Liu, Peiyi Wang, Runxin Xu, Shirong Ma, Chong Ruan, Peng Li, Yang Liu, and Yu Wu. 2025. Inference-Time Scaling for Generalist Reward Modeling. arXiv:2504.02495 [cs.CL] https://arxiv.org/abs/2504.02495
2025
-
[236]
Yujie Lu, Xianjun Yang, Xiujun Li, Xin Eric Wang, and William Yang Wang. 2024. Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation.Advances in Neural Information Processing Systems36 (2024)
2024
-
[237]
Yi-Fan Lu, Xian-Ling Mao, Tian Lan, Chen Xu, and Heyan Huang. 2024. Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models. arXiv:2410.09418 [cs.CL] https://arxiv.org/abs/2410.09418
2024 arXiv
-
[238]
Yi-Fan Lu, Xian-Ling Mao, Tian Lan, Tong Zhang, Yu-Shi Zhu, and Heyan Huang. 2025. SEOE: A Scalable and Reliable Semantic Evaluation Framework for Open Domain Event Detection. arXiv:2503.03303 [cs.CL] https://arxiv.org/abs/2503.03303
2025 arXiv
-
[239]
David Lopez-Paz and Maxime Oquab. 2016. Revisiting classifier two-sample tests.arXiv preprint arXiv:1610.06545(2016)
2016 arXiv
-
[240]
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017. Towards an Automatic Turing Test: Learning to Evaluate Dialogue Responses. InProceedings of the 55th Annual Meeting of the Association for Computational Linguis...
2017 doi
-
[241]
Chang Ma, Junlei Zhang, Zhihao Zhu, Cheng Yang, Yujiu Yang, Yaohui Jin, Zhenzhong Lan, Lingpeng Kong, and Junxian He. 2024. AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents. arXiv:2401.13178 [cs.CL]
2024 arXiv
-
[242]
Wong, and Dacheng Tao
Qingyu Lu, Liang Ding, Liping Xie, Kanjian Zhang, Derek F. Wong, and Dacheng Tao. 2022. Toward Human-Like Evaluation for Natural Language Generation with Error Analysis. arXiv:2212.10179 [cs.CL] https://arxiv.org/abs/2212.10179
2022 arXiv
-
[243]
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Yong Hu, Yu-Shi Zhu, Tong Zhang, Heyan Huang, and Xian-Ling Mao. 2025. Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines. arXiv:2411.16365 [cs.CL] https://arxiv.org/abs/2411.16365
2025 arXiv
-
[244]
Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Shu-Hang Liu, Heyan Huang, Zhijing Wu, Chen Xu, and Xian-Ling Mao. 2025. T2I-Eval-R1: Reinforcement Learning-Driven Reasoning for Interpretable Text-to-Image Evaluation. arXiv:2505.17897 [cs.AI] https://arxiv.org/abs/2505.17897
2025 arXiv
-
[245]
Mounica Maddela, Yao Dou, David Heineman, and Wei Xu. 2023. LENS: A Learnable Evaluation Metric for Text Simplification. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoa...
2023 doi
-
[246]
Liangchen Luo, Zi Lin, Yinxiao Liu, Lei Shu, Yun Zhu, Jingbo Shang, and Lei Meng. 2023. Critique Ability of Large Language Models. arXiv:2310.04815 [cs.LG] https://arxiv.org/abs/2310.04815
2023 arXiv
-
[247]
Yi Luo and Nima Mesgarani. 2018. Tasnet: time-domain audio separation network for real-time, single-channel speech separation. In2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 696–700
2018
-
[248]
François Mairesse, Milica Gašić, Filip Jurčíček, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010. Phrase-Based Statistical Language Generation Using Graphical Models and Active Learning. InProceedings of the 48th Annual Meeting of the Association for Computational ...
2010
-
[249]
Ziyang Ma, Yinghao Ma, Yanqiao Zhu, Chen Yang, Yi-Wen Chao, Ruiyang Xu, Wenxi Chen, Yuanzhe Chen, Zhuo Chen, Jian Cong, et al. 2025. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix.arXiv preprint arXiv:2505.13032(2025)
2025 arXiv
-
[250]
Potsawee Manakul, Adian Liusie, and Mark Gales. 2023. SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 9004–9017
2023
-
[251]
Georgia Maniati, Alexandra Vioni, Nikolaos Ellinas, Karolos Nikitaras, Konstantinos Klapsas, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, and Pirros Tsiakoulis. 2022. SOMOS: The samsung open mos dataset for the evaluation of neural text-to-speech synthesis.arXiv preprint a...
2022 arXiv
-
[252]
Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li, and Alex Smola. 2025. EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge.arXiv preprint arXiv:2505.23009(2025)
2025 arXiv
-
[253]
Dakota Mahan, Duy Van Phung, Rafael Rafailov, Chase Blagden, Nathan Lile, Louis Castricato, Jan-Philipp Fränken, Chelsea Finn, and Alon Albalak. 2024. Generative Reward Models. arXiv:2410.12832 [cs.LG] https://arxiv.org/abs/2410.12832
2024 arXiv
-
[254]
Gallil Maimon, Amit Roth, and Yossi Adi. 2025. Salmon: A Suite for Acoustic Language Model Evaluation. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2025
-
[255]
Shikib Mehri and Maxine Eskenazi. 2020. Unsupervised Evaluation of Interactive Dialog with DialoGPT. InProceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, Olivier Pietquin, Smaranda Muresan, Vivian Chen, Casey Kennington, David Vandy...
2020 doi
-
[256]
Soumi Maiti, Yifan Peng, Takaaki Saeki, and Shinji Watanabe. 2023. Speechlmscore: Evaluating speech generation using speech language model. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5. Manuscript submitted to...
2023
-
[257]
Fanqing Meng, Wenqi Shao, Lixin Luo, Yahong Wang, Yiran Chen, Quanfeng Lu, Yue Yang, Tianshuo Yang, Kaipeng Zhang, Yu Qiao, et al. 2024. PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.arXiv preprint arXiv:2406.11802(2024)
2024 arXiv
-
[258]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781
2013 arXiv
-
[259]
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation. InProceedings of the 2023 Conference on Empirica...
2023
-
[260]
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019. Putting Evaluation in Context: Contextual Embeddings Improve Machine Translation Evaluation. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís ...
2019 doi
-
[261]
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike. 2024. LLM Critics Help Catch LLM Bugs. arXiv:2407.00215 [cs.SE] https://arxiv.org/abs/2407.00215
2024 arXiv
-
[262]
Gabriel Mittag and Sebastian Möller. 2021. Deep learning based assessment of synthetic speech naturalness.arXiv preprint arXiv:2104.11673(2021)
2021 arXiv
-
[263]
Shikib Mehri and Maxine Eskenazi. 2020. USR: An Unsupervised and Reference Free Evaluation Metric for Dialog Generation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (E...
2020 doi
-
[264]
Hyeongdon Moon, Yoonseok Yang, Hangyeol Yu, Seunghyun Lee, Myeongho Jeong, Juneyoung Park, Jamin Shin, Minsam Kim, and Seungtaek Choi. 2022. Evaluating the Knowledge Dependency of Questions. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processi...
2022 doi
-
[265]
Francesco Moramarco, Alex Papadopoulos Korfiatis, Mark Perera, Damir Juric, Jack Flann, Ehud Reiter, Anya Belz, and Aleksandar Savkov. 2022. Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation. InProceedings of the 60th Annual Meeting of the...
2022
-
[266]
Gonçalo Mordido and Christoph Meinel. 2020. Mark-Evaluate: Assessing Language Generation using Population Estimation Methods. InProceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). International Commi...
2020 doi
-
[267]
Christoph Minixhofer, Ondřej Klejch, and Peter Bell. 2024. TTSDS-Text-to-Speech Distribution Score. In2024 IEEE Spoken Language Technology Workshop (SLT). IEEE, 766–773
2024
-
[268]
Gabriel Mittag, Ross Cutler, Yasaman Hosseinkashi, Michael Revow, Sriram Srinivasan, Naglakshmi Chande, and Robert Aichner. 2020. DNN No-Reference PSTN Speech Quality Prediction. InProc. Interspeech 2020. 2867–2871
2020
-
[269]
Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, Vu Le, Nick McKenna, Carina Suzana Negreanu, Chris Parnin, and Advait Sarkar. 2024. Evaluating the Evaluator: Measuring LLMs’ Adherence to Task Evaluation Instructions. arXiv:2408.08781 [cs.AI] https://arxiv.org/abs/2408.08781
2024
-
[270]
Gabriel Mittag, Babak Naderi, Assmaa Chehadi, and Sebastian Möller. 2021. NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets.arXiv preprint arXiv:2104.09494(2021)
2021 arXiv
-
[271]
Cohen, and Mirella Lapata
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. arXiv:1808.08745 [cs.CL]
2018 arXiv
-
[272]
Matteo Negri, Marco Turchi, José G. C. de Souza, and Daniele Falavigna. 2014. Quality Estimation for Automatic Speech Recognition. InProceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, Junichi Tsujii and Jan Hajic (Eds....
2014
-
[273]
Preksha Nema and Mitesh M. Khapra. 2018. Towards a Better Metric for Evaluating Question Generation Systems. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ichi Tsujii (Eds.). Ass...
2018 doi
-
[274]
John Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi. 2020. TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Sys...
2020 doi
-
[275]
mrfakename, Vaibhav Srivastav, Clémentine Fourrier, Lucain Pouget, Yoach Lacombe, main, and Sanchit Gandhi. 2024. Text to Speech Arena. https://huggingface.co/spaces/TTS-AGI/TTS-Arena
2024
-
[276]
Cheng Niu, Yuanhao Wu, Juno Zhu, Siliang Xu, KaShun Shum, Randy Zhong, Juntong Song, and Tong Zhang. 2024. RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computati...
2024 doi
-
[277]
Keerthiram Murugesan, Sarathkrishna Swaminathan, Soham Dan, Subhajit Chaudhury, Chulaka Gunasekara, Maxwell Crouse, Diwakar Mahajan, Ibrahim Abdelaziz, Achille Fokoue, Pavan Kapanipathi, Salim Roukos, and Alexander Gray. 2023. MISMATCH: Fine-grained Evaluation of Machine- gene...
2023
-
[278]
OpenBMB. 2025. UltraEval-Audio: An Easy-to-Use, Fast, and Easily Integrable Tool for Evaluating Audio LLM. https://github.com/OpenBMB/ UltraEval-Audio. Accessed: 2025-04-18
2025
-
[279]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022 arXiv
-
[280]
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 5206–5210. Manuscript submitted to ACM 66 Lan et al
2015
-
[281]
Jun-Ping Ng and Viktoria Abrecht. 2015. Better Summarization Evaluation with Word Embeddings for ROUGE. InProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lluís Màrquez, Chris Callison-Burch, and Jian Su (Eds.). Association for Computatio...
2015 doi
-
[282]
Tuan Nguyen, Corinne Fredouille, Alain Ghio, Mathieu Balaguer, and Virginie Woisard. 2024. Exploring pathological speech quality assessment with ASR-powered Wav2Vec2 in data-scarce context.arXiv preprint arXiv:2403.20184(2024)
2024 arXiv
-
[283]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds....
2002
-
[284]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[285]
Dong Huk Park, Samaneh Azadi, Xihui Liu, Trevor Darrell, and Anna Rohrbach. 2021. Benchmark for compositional text-to-image synthesis. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)
2021
-
[286]
Junsoo Park, Seungyeon Jwa, Ren Meiying, Daeyoung Kim, and Sanghyuk Choi. 2024. OffsetBias: Leveraging Debiased Data for Tuning Evaluators. InFindings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Associ...
2024 doi
-
[287]
Junsoo Park, Seungyeon Jwa, Meiying Ren, Daeyoung Kim, and Sanghyuk Choi. 2024. OffsetBias: Leveraging Debiased Data for Tuning Evaluators. arXiv:2407.06551 [cs.CL] https://arxiv.org/abs/2407.06551
2024 arXiv
-
[288]
Prabhat Pandey, Rupak Vignesh Swaminathan, KV Girish, Arunasish Sen, Jian Xie, Grant P Strimel, and Andreas Schwarz. 2025. SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning.arXiv preprint arXiv:2504.09081(2025)
2025 arXiv
-
[289]
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The LAMBADA dataset: Word prediction requiring a broad discourse context.arXiv preprint arXiv:1606.06031(2016)
2016 arXiv
-
[290]
Zifan Peng, Yule Liu, Zhen Sun, Mingchen Li, Zeren Luo, Jingyi Zheng, Wenhan Dong, Xinlei He, Xuechao Wang, Yingjie Xue, Shengmin Xu, and Xinyi Huang. 2025. JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models. arXiv:2505.17568 [cs.CR]
2025
-
[291]
ChaeHun Park, Seungil Lee, Daniel Rim, and Jaegul Choo. 2023. DEnsity: Open-domain Dialogue Evaluation Metric using Density Estimation. In Findings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Associati...
2023 doi
-
[292]
Vitou Phy, Yang Zhao, and Akiko Aizawa. 2020. Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems. InProceedings of the 28th International Conference on Computational Linguistics, Donia Scott, Nuria Bel, and Chengqing Zong (Eds.). Inter...
2020 doi
-
[293]
Jaden Pieper and Stephen Voran. 2024. AlignNet: Learning dataset score alignment functions to enable better training of speech quality estimators. InProc. Interspeech 2024. 82–86
2024
-
[295]
Brian Patton, Yannis Agiomyrgiannakis, Michael Terry, Kevin Wilson, Rif A Saurous, and D Sculley. 2016. AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.arXiv preprint arXiv:1611.09207(2016)
2016 arXiv
-
[296]
Abhirama Subramanyam Penamakuri, Kiran Chhatre, and Akshat Jain. 2025. Audiopedia: Audio QA with Knowledge. InICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5
2025
-
[298]
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. InProceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Alessandro Moschitti, Bo Pang, and Walter Daelemans (Eds.). Assoc...
2014 doi
-
[2020]
InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.)
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). ...
-
[2022]
Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing16, 6 (2022), 1505–1518
2022
-
[2023]
T-Eval: Evaluating the Tool Utilization Capability Step by Step.arXiv preprint arXiv:2312.14033(2023)
2023 arXiv
-
[2024]
arXiv:2407.14507 [cs.CL] https://arxiv.org/abs/2407.14507
Internal Consistency and Self-Feedback in Large Language Models: A Survey. arXiv:2407.14507 [cs.CL] https://arxiv.org/abs/2407.14507
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.