REVIEW 4 major objections 7 minor 54 references
An Empirical Study of Evaluating Long-form Question Answering
T0 review · 4 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Fine-grained GPT-4o judging aligns with human scores of long-form answers far better than ROUGE or BERTScore, a 2,079-answer study finds, while LLM judges show their own biases.
desk verdict Useful meta-evaluation with a clear practical message, but the headline averages are miscalculated and the human ground truth is unreported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The meta-evaluation protocol: 5,236 generated answers from seven LLMs across ASQA, ANTIQUE, and WikiEval, with 2,079 human-rated on correctness and informativeness, are compared against seven automatic metrics using Spearman and Kendall correlations and pairwise win-rate agreement. The improvement mechanism is fine-grained prompting, which decomposes the evaluation prompt into four components (task, data, output requirements, criteria) and supplies explicit per-criterion rubrics; this raises GPT-4o's Spearman correlation from 42.0 to 55.0 and is the central lever the paper identifies for making LLM judges align with humans.
What would settle it
Re-run the meta-evaluation on a fresh sample of long-form answers labeled by multiple independent annotators with measured inter-annotator agreement; if fine-grained GPT-4o's Spearman correlation with the multi-annotator gold scores does not significantly exceed ROUGE-L and BERTScore, the paper's central conclusion fails. A cheaper check on the released data is to estimate label noise from a small re-annotation subset and test whether the reported 55.0 versus 42.0 fine-grained gain survives noise correction.
Extended reading notes
Core claim
Fine-grained evaluation with GPT-4o improves its Spearman correlation with human ratings on ASQA from 42.0 to 55.0, far exceeding deterministic metrics like Rouge-L (11.4), Exact Match (41.4), and Disambig-F1 (23.0). The authors argue that providing more detailed instructions, decomposing the judge prompt into task, data, output format, and explicit 1 to 5 rubrics for accuracy and informativeness, is the decisive factor in making LLM-based LFQA evaluation usable. At the same time, they show that LLM judges systematically favor longer answers, give higher scores to their own outputs, shift with prompt wording and temperature, and that no single prompt configuration is optimal for all evaluated models.
Load-bearing premise
The 2,079 human ratings of correctness and informativeness are treated as reliable ground truth, but the paper reports no inter-annotator agreement, no annotator count, and no validation of the single-point 1 to 5 rating scale, so any noise or bias in those labels propagates into every correlation and bias claim.
Editorial extensions
If this is right
- Deterministic metrics such as ROUGE-L and BERTScore are unreliable as primary LFQA measures, especially for non-factoid and open-ended answers, so results built on them should be interpreted with caution.
- LLM judges with fine-grained, multi-criteria instructions can approximate human consistency on correctness and informativeness, making large-scale LFQA evaluation feasible without per-item human annotation.
- Because LLM judges favor longer answers and their own outputs, practitioners should rank models rather than trust raw scores, and should avoid letting a model grade its own responses.
- Judge prompt structure matters: a four-component prompt (task, data, output, criteria) generally improves consistency, but no single prompt works best for all LLMs, so prompt selection affects leaderboard conclusions.
- Judge temperature can change model rankings on tightly scored datasets, so LFQA evaluations should fix a low temperature and report the setting.
Reading between the lines
- The paper's self-preference and length-bias results imply that single-judge LFQA leaderboards may systematically rank verbose answers from the judge's own model family higher; a direct test would be to measure rank changes when the judge is held out of the generation pool.
- Since fine-grained rubrics raise consistency, a natural next step is to combine the rubric decomposition with claim-level atomic evaluation, splitting answers into verifiable units before scoring, a hybrid the paper does not test but its criteria-decomposition finding directly suggests.
- The absence of inter-annotator agreement data means the true ceiling for any automatic metric is unknown; re-annotating a subset of the released 2,079 answers with multiple judges per item would recalibrate all reported correlations and quantify label noise.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an empirical meta-evaluation of automatic metrics for long-form question answering (LFQA). Using answers generated by seven LLMs on ASQA, ANTIQUE, and WikiEval, the authors collect human ratings of correctness and informativeness for 2,079 QA pairs and compare these with deterministic metrics (Rouge-L, Exact Match, Disambig-F1, BERTScore, RAGAS answer relevance) and LLM-based judges (GPT-4o, Claude-3.5, Gemini-2.0, in coarse-grained and fine-grained variants). The main findings are that LLM-based judges correlate more strongly with human ratings than deterministic metrics, that deterministic metrics are sensitive to answer length and question type, that LLM judges exhibit self-preference and other biases, and that fine-grained prompt design improves LLM-judge agreement with humans, e.g., GPT-4o improves from 42.0 to 55.0 Spearman correlation on ASQA. The paper also proposes decomposing prompts into task, data, output, and criteria components. Code and data are released.
Significance. If the results are reliable, this is a practically useful contribution: it provides a moderate-sized human-annotated LFQA evaluation resource (2,079 QA pairs, 4,158 ratings), compares a broad range of metrics on three diverse datasets, and offers concrete evidence on the value of fine-grained LLM-based evaluation. The analysis of length, question-type, and self-preference biases is relevant to anyone using LLM judges. Credit is due for the explicit release of code and data and for acknowledging limitations such as the single-point rating scale. However, the quantitative conclusions currently rest on a human ground truth whose reliability is not reported, and several aggregate numbers in the main tables are internally inconsistent. These issues must be resolved before the headline claims can be accepted.
major comments (4)
- [3.1.3] The human annotations are the reference for every correlation and bias analysis in Sections 3.2 through 3.4, but the manuscript reports no annotator count, no redundancy in labeling, no adjudication procedure, and no inter-annotator agreement statistic such as Cohen's kappa or Krippendorff's alpha. Without this information, the reported Spearman and Kendall coefficients are correlations against an unvalidated target, and any annotator leniency, drift, or answer-length-related bias is confounded with metric behavior. The Conclusion acknowledges the single-point scale as a limitation but not annotator reliability. To make the central claims load-bearing, the authors should report the annotation protocol in detail or, if the data cannot be recovered, substantially temper the quantitative claims and provide an uncertainty analysis.
- [Tables 1 and 2] The Average rows in Tables 1 and 2 are not the means of the model-level rows printed in the same tables. For example, in Table 2 the ANTIQUE CG average is printed as 83.0, but the mean of the seven listed values (36.7, 53.2, 74.9, 69.3, 65.6, 73.0, 54.3) is 61.0; the printed ASQA GPT-4o average of 42.0 differs from the mean of 40.5; and the ASQA EM average of 41.4 differs from the mean of 42.9. If the averages are computed through a nonlinear transformation such as Fisher z-transformation and back-transformation, this must be stated and justified; otherwise the reader cannot interpret the aggregates. This issue directly affects the headline claim in Section 3.2 that fine-grained GPT-4o improves from 42.0 to 55.0, since both numbers come from these average rows.
- [3.2] The manuscript repeatedly states that LLM-based metrics show "significantly higher consistency" with human judgments, but no significance tests, confidence intervals, or effect-size uncertainties are reported. The per-model sample sizes differ across datasets (about 50 answers per model for ASQA and WikiEval versus about 200 per model for ANTIQUE), and many of the reported differences, for instance the 42.0 versus 33.0 difference between GPT-4o and Claude-3.5 on ASQA, could be within sampling noise. The authors should provide confidence intervals or at least a significance test for the main correlation comparisons, or soften the language accordingly.
- [3.4.3] The self-reinforcement analysis concludes that LLM-based evaluators "assign significantly higher scores to their own outputs," but the evidence is win-rate comparisons from Figure 6 without a statistical test and without a control for genuine answer-quality differences. If GPT-4o, Claude-3.5, and Gemini-2.0 simply produce better answers to most ASQA questions, the same win-rate pattern would be observed even with a perfectly unbiased judge. The claim of a "clear evaluation bias" therefore needs a stronger design, such as comparing judges on outputs of matched quality, or at least an inter-judge consistency analysis that accounts for model strength. As written, the bias conclusion is not fully supported.
minor comments (7)
- [2.1] There is a typo in "abstract lonng-form summarization" (should be "long-form").
- [3.2] The phrase "fatoid QA" should be "factoid QA."
- [3.1.3] The sentence "A high-quality response should should be informative enough" contains a duplicated "should."
- [Table 5] Table 5 is difficult to parse: the header "Kendall (%) LLMs Prompts" mixes the three default LLM-judge columns with the nine prompt-condition columns, and the caption says "Variation of correlation coefficients" but the table appears to show absolute correlations. Please restructure the table into clear blocks and define exactly what each column reports.
- [Figure 6] Figure 6 is visually cluttered because the three win-rate matrices are overlaid and the row/column labels are partially repeated. Separate subplots or a table with model labels would make the self-reinforcement comparison readable.
- [3.3.2] Table 3 reports temperature sensitivity for ASQA and WikiEval but not ANTIQUE; the text does not explain this omission. Please either include ANTIQUE or state why it is excluded.
- [3.4.4] The IDF analysis in Section 3.4.4 is qualitative. Please specify the exact statistic used (e.g., Spearman correlation between average IDF and metric score) and report the corresponding values, since the text uses words like "strongly correlate" without numerical support.
Circularity Check
No circularity: the paper is an empirical meta-evaluation against external human judgments, with no fitted-parameter, self-citation, or definitional reductions.
full rationale
The paper's derivation chain is an empirical comparison, not a formal derivation. It collects 5,236 model-generated answers, obtains 2,079 human ratings of correctness and informativeness, and then measures Spearman/Kendall correlations and win-rate agreements between those external human labels and various automatic metrics (ROUGE-L, EM, BERTScore, RAGAS AR, GPT-4o, Claude-3.5, Gemini-2.0). No quantity that is claimed as a result is defined in terms of the conclusion: e.g., 'the fine-grained evaluation with GPT-4o improves its score from 42.0 to 55.0' is a measured correlation coefficient against human ratings, not an identity or a fitted parameter renamed as a prediction. The same-model judge/generator setup in Section 3.4.3 is explicitly treated as the object of study (self-preference bias) rather than as evidence for the main accuracy claim. There are no load-bearing self-citations: the references to LLM-EVAL, G-EVAL, RAGAS, and prior meta-evaluation work are external prior work, and none of the paper's authors' own prior results is invoked to forbid alternatives or to justify a central premise. The paper's limitations, such as the unvalidated single-point human rating scale and lack of reported inter-annotator agreement, are threats to the reliability of the human ground truth, but they do not make any step circular because the human reference is external to the automatic metrics being studied. The empirical findings are therefore self-contained against an external benchmark, and the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Human judgments of correctness and informativeness are an appropriate ground truth for LFQA evaluation.
- domain assumption The sampled testbeds (ASQA, ANTIQUE, WikiEval) are representative of ambiguous factoid, non-factoid, and factoid LFQA.
- domain assumption Correlation coefficients and win-rate agreement are sufficient meta-evaluation measures.
- domain assumption The selected prompt perturbations and temperature values are sufficient to assess robustness.
Cite this review
Pith. "Pith review of An Empirical Study of Evaluating Long-form Question Answering." pith.science (2026). https://pith.science/paper/GW7J7RCD
@misc{pith2026250418413,
author = {Pith},
title = {Pith review of: An Empirical Study of Evaluating Long-form Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/GW7J7RCD}},
note = {Machine review of arXiv:2504.18413}
}
read the original abstract
\Ac{LFQA} aims to generate lengthy answers to complex questions. This scenario presents great flexibility as well as significant challenges for evaluation. Most evaluations rely on deterministic metrics that depend on string or n-gram matching, while the reliability of large language model-based evaluations for long-form answers remains relatively unexplored. We address this gap by conducting an in-depth study of long-form answer evaluation with the following research questions: (i) To what extent do existing automatic evaluation metrics serve as a substitute for human evaluations? (ii) What are the limitations of existing evaluation metrics compared to human evaluations? (iii) How can the effectiveness and robustness of existing evaluation methods be improved? We collect 5,236 factoid and non-factoid long-form answers generated by different large language models and conduct a human evaluation on 2,079 of them, focusing on correctness and informativeness. Subsequently, we investigated the performance of automatic evaluation metrics by evaluating these answers, analyzing the consistency between these metrics and human evaluations. We find that the style, length of the answers, and the category of questions can bias the automatic evaluation metrics. However, fine-grained evaluation helps mitigate this issue on some metrics. Our findings have important implications for the use of large language models for evaluating long-form question answering. All code and datasets are available at https://github.com/bugtig6351/lfqa_evaluation.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Rachith Aiyappa, Jisun An, Haewoon Kwak, and Yong-Yeol Ahn. 2023. Can we trust the evaluation on ChatGPT? arXiv:2303.12767 (March 2023). https: //doi.org/10.48550/arXiv.2303.12767 arXiv:2303.12767 [cs]
-
[2]
Meghana Moorthy Bhat, Rui Meng, Ye Liu, Yingbo Zhou, and Semih Yavuz
-
[3]
Bruce Croft, and Mark Sanderson
Valeriia Bolotova, Vladislav Blinov, Falk Scholer, W. Bruce Croft, and Mark Sanderson. 2022. A Non-Factoid Question-Answering Taxonomy. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’22) . Association for Computing Machinery, New York, NY, USA, 1196–1207. https://doi.org/10.1145/3...
arXiv 2022
-
[4]
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al
-
[5]
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: A Preliminary Empirical Study. arXiv:2304.00723 (April 2023). https://doi.org/ 10.48550/arXiv.2304.00723 arXiv:2304.00723 [cs]
-
[6]
Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv. arXiv preprint arXiv:2303.12712 (2023)
arXiv 2023
-
[7]
Cheng-Han Chiang and Hung-yi Lee. 2023. A closer look into automatic evalua- tion using large language models. arXiv preprint arXiv:2310.05657 (2023)
arXiv 2023
-
[8]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv:2305.01937 (May 2023). https: //doi.org/10.48550/arXiv.2305.01937 arXiv:2305.01937 [cs]
Show all 54 references
-
[9]
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023. Ra- gas: Automated evaluation of retrieval augmented generation. arXiv preprint arXiv:2309.15217 (2023)
2023 arXiv
- [10]
-
[11]
Yuchen Fan, Xin Zhong, Yazhe Wan, Chengsi Wang, Haonan Cheng, Gaoche Wu, Ning Ding, and Bowen Zhou. 2024. EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation. arXiv preprint arXiv:2407.04969 (2024)
2024
-
[12]
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019. ELI5: Long Form Question Answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Florence, ...
2019
-
[13]
Team GLM, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie H...
-
[14]
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023. Gptscore: Evaluate as you desire. arXiv preprint arXiv:2302.04166 (2023)
2023 arXiv
- [15]
-
[16]
Helia Hashemi, Mohammad Aliannejadi, Hamed Zamani, and W Bruce Croft
- [17]
-
[18]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, De- vendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al . 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[19]
Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Yungi Kim, Hyeonju Lee, Jihoo Kim, Changbae Ahn, Seonghoon Yang, Sukyung Lee, Hyunbyung Park, Gyoungjin Gim, Mikyoung Cha, Hwalsuk Lee, and Sunghun Kim. 2023. SOLAR 10.7B: Scaling Large L...
2023 arXiv
-
[20]
Kalpesh Krishna, Aurko Roy, and Mohit Iyyer. 2021. Hurdles to Progress in Long-form Question Answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, Kristina Toutanova, Anna...
2021
-
[21]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Wenliang Dai, Andrea Madotto, and Pascale Fung. 2023. Survey of Hallucination in Natural Language Generation. Comput. Surveys 55, 12 (Dec. 2023), 1–38. https://doi.org/10.1145/3571730 ar...
2023 arXiv
-
[22]
Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michi- hiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Nav...
-
[23]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74–81
2004
-
[24]
Yen-Ting Lin and Yun-Nung Chen. 2023. Llm-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models. arXiv preprint arXiv:2305.13711 (2023)
2023 arXiv
-
[25]
Saiful Bari, Mizanur Rahman, Md Am- ran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang
Md Tahmid Rahman Laskar, M. Saiful Bari, Mizanur Rahman, Md Am- ran Hossen Bhuiyan, Shafiq Joty, and Jimmy Xiangji Huang. 2023. A Sys- tematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets. arXiv:2305.18486 (July 2023). https://doi.org/10.48550/arXiv.2305...
- [26]
-
[27]
Ben Mann, N Ryder, M Subbiah, J Kaplan, P Dhariwal, A Neelakantan, P Shyam, G Sastry, A Askell, S Agarwal, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 1 (2020)
2020 arXiv
- [28]
-
[29]
OpenAI. 2023. GPT-3.5-Turbo-Instruct. https://www.openai.com Accessed: 2024-07-18
2023
-
[30]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. G-eval: Nlg evaluation using gpt-4 with better human alignment. arXiv preprint arXiv:2303.16634 (2023)
2023 arXiv
-
[31]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
2002
-
[32]
Ehud Reiter and Anja Belz. 2009. An Investigation into the Validity of Some Metrics for Automatically Evaluating Natural Language Generation Systems. Computational Linguistics 35, 4 (Dec. 2009), 529–558. https://doi.org/10.1162/ coli.2009.35.4.35405
2009
-
[33]
Ivan Stelmakh, Yi Luan, Bhuwan Dhingra, and Ming-Wei Chang. 2022. ASQA: Factoid Questions Meet Long-Form Answers. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for ...
2022 doi
-
[34]
Dan Su, Xiaoguang Li, Jindi Zhang, Lifeng Shang, Xin Jiang, Qun Liu, and Pascale Fung. 2022. Read before Generate! Faithful Long Form Question Answering with Machine Reading. arXiv:2203.00343 [cs.CL] https://arxiv.org/abs/2203.00343
2022 arXiv
-
[35]
OpenAI. 2025. Explore developer resources, tutorials, API docs, and dynamic examples to get the most out of OpenAI’s platform. https://platform.openai.com Accessed: 2024-01-24
2025
-
[36]
Meta LLaMA Team. 2024. Introducing Meta Llama 3: The most capable openly available LLM to date. https://ai.meta.com/blog/meta-llama-3/
2024
-
[37]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yas- mine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhos- ale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[38]
Tu Vu, Kalpesh Krishna, Salaheddin Alzubi, Chris Tar, Manaal Faruqui, and Yun- Hsuan Sung. 2024. Foundational autoraters: Taming large language models for An Empirical Study of Evaluating Long-form Question Answering SIGIR ’25, July 13–18, 2025, Padua, Italy better automatic e...
2024 arXiv
-
[39]
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023. Is chatgpt a good nlg evaluator? a preliminary study. arXiv preprint arXiv:2303.04048 (2023)
2023 arXiv
-
[40]
Mingxu Tao, Dongyan Zhao, and Yansong Feng. 2024. Chain-of-Discussion: A Multi-Model Framework for Complex Evidence-Based Question Answering. arXiv:2402.16313 [cs.CL] https://arxiv.org/abs/2402.16313
2024 arXiv
- [41]
-
[42]
Junchao Wu, Shu Yang, Runzhe Zhan, Yulin Yuan, Lidia Sam Chao, and Derek Fai Wong. 2025. A survey on LLM-generated text detection: Necessity, methods, and future directions. Computational Linguistics (2025), 1–65
2025
-
[43]
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi. 2023. A critical evaluation of evaluations for long-form question answering. arXiv preprint arXiv:2305.18201 (2023)
2023 arXiv
-
[44]
Qinyuan Ye, Maxamed Axmed, Reid Pryzant, and Fereshte Khani. 2023. Prompt engineering a prompt engineer. arXiv preprint arXiv:2311.05661 (2023)
2023 arXiv
-
[45]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, and Zhifang Sui. 2023. Large Language Models are not Fair Evaluators. arXiv:2305.17926 [cs.CL] https://arxiv.org/abs/2305.17926
2023 arXiv
-
[46]
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019. Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019)
2019 arXiv
- [47]
-
[48]
Meyer, and Stef- fen Eger
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Stef- fen Eger. 2019. MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing ...
2019
-
[49]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2024)
2024
- [50]
-
[2020]
In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42
ANTIQUE: A non-factoid question answering benchmark. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part II 42 . Springer, 166–173
2020
-
[2022]
arXiv:2211.09110 (Nov
Holistic Evaluation of Language Models. arXiv:2211.09110 (Nov. 2022). http://arxiv.org/abs/2211.09110 arXiv:2211.09110 [cs]
2022 arXiv
- [2023]
-
[2024]
arXiv:2406.12793
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools. arXiv:2406.12793
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.