Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

The Science of Evaluating Foundation Models

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read LLM evaluation should start from the use case, not from generic benchmarks, and the paper turns this into an ABCD checklist with pruning and documentation stages.

desk verdict A readable checklist and survey with an overstated 'no actionable guideline exists' claim that its own references contradict; fix the framing and it's worth a serious referee. read the letter →

arxiv 2502.09670 v1 pith:3KACSY7B submitted 2025-02-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords foundationmodelslargelanguagemodelevaluationframeworkABCDchecklistbenchmarkselectiondomainexpertisereproducible
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the field lacks an actionable guideline that tells practitioners how to evaluate a large language model end-to-end, and it sets out to provide one. The proposed ABCD framework organizes evaluation around Algorithm, Big Data, Computation Resources, and Domain Expertise, then translates those four lenses into a step-by-step checklist, an applicability-analysis stage where evaluators prune and weight evaluation dimensions for their use case, and documentation standards for reproducibility. The paper also reviews recent evaluation dimensions, metrics, and tools, positioning them inside this process. A sympathetic reader should take the central claim as: evaluation becomes a repeatable, context-aware procedure rather than a collection of loosely connected benchmarks.

What carries the argument

The central mechanism is the ABCD framework (Algorithm, Big Data, Computation Resources, Domain Expertise) used as an organizing alphabet for evaluation, together with the three-stage workflow it feeds: a preparation checklist, applicability analysis, and documentation. The checklist maps each preparation step to the relevant ABCD letter, so choices about models, datasets, metrics, baselines, ethics and safety, and resources are made explicit before experiments begin. The applicability-analysis stage supplies the paper's main operational idea: not every evaluation dimension is needed for every task, so evaluators should assign relative weights to dimensions and prune to what is feasible, then disclose those weights in documentation. This weighting-and-disclosure mechanism is what converts the framework from a taxonomy into a decision procedure.

What would settle it

Search the literature and practitioner tooling for an existing, widely available evaluation process that already includes step-by-step instructions for defining objectives, selecting datasets and metrics, setting baselines, addressing ethics and safety, allocating resources, and documenting results; finding one that practitioners can follow end-to-end in a new domain would falsify the paper's claim that no such actionable guideline exists.

Watch

Extended reading notes

Core claim

The claim is that there is no actionable evaluation guideline incorporating a cohesive process for large language models, and that evaluation should be driven by use-case context rather than generic leaderboards. The paper formalizes the process with the ABCD framework: Algorithm covers model choices and baselines, Big Data covers selection and diversity of evaluation datasets, Computation Resources covers memory, GPU, storage, and inference constraints, and Domain Expertise covers contextually meaningful metrics and human evaluation. These four letters anchor a checklist of eight preparation steps, an applicability-analysis stage in which evaluators weight dimensions and prune unnecessary evaluations, and documentation standards that include model cards and data sheets. If the paper is right, evaluating an LLM becomes a disciplined method that can be repeated, audited, and adapted to domains such as healthcare or law.

Load-bearing premise

The load-bearing premise is the gap claim introduced in Section 1: that no prior work offers an actionable, cohesive evaluation guideline, even though Section 7 itself lists existing frameworks and tools that could be read as exactly such guidelines.

Editorial extensions

If this is right

  • Following the ABCD checklist before running experiments makes model-selection decisions traceable to the stated use case.
  • Resource-constrained teams can prune low-priority evaluation dimensions and still produce a defensible evaluation report.
  • Disclosing dimension weights makes benchmark results comparable between teams with different priorities.
  • Domain experts gain a defined role in evaluation through choosing datasets, metrics, and qualitative checks that automated benchmarks miss.
  • New evaluation metrics and tools can be placed into a single process instead of being treated as competing leaderboards.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If ABCD becomes common practice, questions like 'which model is best' would shift from aggregate leaderboard rankings to context-specific fitness-for-use statements, reducing the misleading simplicity of a single number.
  • The weighting step implies a research program of eliciting and validating stakeholder weights for evaluation dimensions, since the paper acknowledges those weights are subjective.
  • The framework's advice implies a testable hypothesis: teams using the checklist produce more reproducible and decision-relevant evaluations than teams relying on generic benchmarks, which a controlled comparison could check.
  • The suggested multi-agent evaluation direction could operationalize ABCD by assigning each letter to an agent with distinct responsibilities, an extension the paper mentions but does not implement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes that existing LLM evaluation literature lacks an actionable, cohesive process that integrates use-case context with ethical and operational considerations. To fill this gap, it introduces the ABCD framework (Algorithm, Big Data, Computation Resources, Domain Expertise), a nine-step evaluation checklist in Table 2, a three-stage workflow of checklist, applicability analysis, and documentation, and a targeted survey of evaluation dimensions (performance, robustness, fairness, explainability, safety) and metrics. The paper explicitly disclaims offering a new evaluation method or an exhaustive survey, and instead emphasizes formalizing the process and providing practical tools.

Significance. The paper's strength is its compact organization of a large body of evaluation literature into usable categories, and the ABCD checklist is a coherent, readable starting point for practitioners. Several existing frameworks (HELM, LalaEval, fmeval, Chang et al., Peng et al.) are cited, which makes the paper a useful pointer into the space. However, the central claim of a missing actionable guideline is contradicted by those same citations, and the paper provides no worked example or empirical demonstration that the ABCD checklist improves evaluation decisions. With a reframed contribution and a concrete application, the material could serve as a useful tutorial or position piece, but as written it does not substantiate the claimed novelty.

major comments (3)
  1. [Section 1 / Section 7] The paper's load-bearing claim in Section 1 that 'there exists no actionable evaluation guideline incorporating a cohesive process' is internally inconsistent with Section 7. Section 7.2 describes LalaEval as 'a holistic human evaluation framework for domain-specific LLMs, encompassing domain specification, criteria establishment, benchmark dataset creation, evaluation rubric construction, and thorough analysis of evaluation outcomes,' which is precisely an actionable, domain-aware process. Section 7.1 describes Peng et al.'s two-stage framework from core abilities to agent applications and Chang et al.'s categorization of evaluation methods, and Section 7.2 describes fmeval as an open-source library covering both performance and responsible-AI dimensions. Since these existing frameworks and tools provide structured, context-aware evaluation processes, the gap claim as stated is contradicted by the paper's own survey. The contribution should be reframed as a synthesis or operational checklist that consolidates existing guidelines, with an explicit paragraph stating what ABCD adds beyond terminology.
  2. [Section 2.3 / Table 1] The memory-requirement guidance is internally inconsistent. Section 2.3 first states that a 7B-parameter model 'requires approximately 28 GB of memory, assuming 4 bytes per parameter,' then immediately gives a rule of thumb of approximately 2 × X GB for X billion parameters in bfloat16/float16. Table 1 lists the 7B row as 14 GB and all rows follow the 2 bytes-per-parameter scaling. If the 4 bytes-per-parameter figure refers to FP32 weights and the table refers to BF16 weights, this distinction must be stated; if the table is meant to include runtime activation memory, the relationship is mislabeled. Because Table 1 is presented as planning guidance and the checklist includes 'Allocate Resources (C),' an inconsistent resource model weakens the paper's practical utility.
  3. [Section 5 / Table 2] The paper claims to formalize the evaluation process, but Table 2's checklist consists of generic project-management steps (define objectives, prioritize dimensions, select datasets, identify metrics, establish baselines, address ethics, allocate resources, document, iterate) with no illustration of how the ABCD decomposition changes a concrete evaluation decision. Section 5.2 offers informal examples of selectively weighting dimensions, but there is no end-to-end use case, case study, or comparison with an existing framework such as LalaEval or HELM. Adding at least one worked example (e.g., evaluating a model for healthcare question answering or code generation) and, if feasible, a comparison with a baseline evaluation practice would substantiate the claim that the framework is actionable and useful.
minor comments (5)
  1. [Section 1] The phrase 'how to systemically approach LLM evaluation' should be 'systematically,' and the sentence structure in 'current research [10, 58] lacks a comprehensive...' should be rephrased to make clear that the references do not themselves lack comprehensiveness.
  2. [Section 2.3] The claim that models 'exceeding 100 billion parameters demand exponentially more memory' is inaccurate relative to Table 1, which shows linear growth at 2 bytes per parameter; replace 'exponentially' with 'proportionally' or specify which overheads become nonlinear.
  3. [Figure 1] The workflow diagram shows five unlabeled boxes and no arrow labels or stage names; annotate it to match Sections 5.1–5.3 so that the relationship between the checklist, applicability analysis, and documentation is clear.
  4. [Section 5.3] The bullet beginning 'Employing standardized documentation tools...' is a sentence continuation rather than a parallel bullet item; merge it into the previous line or rewrite it as a proper bullet.
  5. [Section 3.1] The heading 'Entity/Word Extraction are tasks' should read 'Entity/Word Extraction is a task category,' and the subsequent sentence beginning 'This category encompasses...' should be adjusted for number agreement.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a survey and framework proposal with no fitted quantities, no predictions derived from inputs, and no load-bearing self-citation chain.

full rationale

This is a review/framework paper rather than a derivation: it proposes the ABCD evaluation framework, surveys existing evaluation dimensions, metrics, and tools, and provides a checklist. There are no equations, fitted parameters, or empirical predictions whose output is equivalent to an input by construction. The paper's self-citations ([87], [92]) are background references to the authors' prior work on LLM-as-evaluator benchmarks and healthcare data augmentation; they are not used to justify the central framework or to forbid alternatives. The main contestable claim is the Section 1 assertion that 'there exists no actionable evaluation guideline incorporating a cohesive process,' which is a novelty/gap claim and, if anything, a correctness concern given the frameworks cited in Section 7 (HELM, Chang et al., Peng et al., fmeval, LalaEval). But an overstated gap claim is not circularity: the paper does not define its contribution in terms of that gap, nor does it fit a parameter and then rename the fit as a prediction. The ABCD checklist is a generic organizational device, not a renamed empirical result presented as unification. Accordingly, no circular step can be quoted and exhibited under the specified patterns, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central contribution is a framework, not a derivation, so there are no fitted parameters or invented physical entities. The burden sits in three unproved premises: the sufficiency of ABCD, the subjectivity and prunability of dimensions, and the claimed gap in prior work. The gap claim is the strongest assumption and is partially contradicted by the paper's own related-work section.

assumptions (3)
  • ad hoc to paper The ABCD decomposition (Algorithm, Big Data, Computation Resources, Domain Expertise) is a sufficient organizing basis for LLM evaluation.
    Proposed in Section 2 without empirical support, then used to structure the rest of the paper; if it omitted a salient factor such as societal impact, the framework would be incomplete.
  • domain assumption A subset of evaluation dimensions can be pruned for a task, and the relative importance and weights are inherently subjective.
    Section 5.2 argues selective pruning and transparent weighting are needed, but no method is given for determining which dimensions to omit or how to weight them.
  • domain assumption Existing literature provides no actionable evaluation guideline incorporating a cohesive process.
    Stated in the Introduction to motivate the work; contradicted in part by the paper's own related-work section listing HELM, Chang et al., Peng et al., fmeval, and LalaEval.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Science of Evaluating Foundation Models." pith.science (2026). https://pith.science/paper/3KACSY7B

@misc{pith2026250209670,
  author       = {Pith},
  title        = {Pith review of: The Science of Evaluating Foundation Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3KACSY7B}},
  note         = {Machine review of arXiv:2502.09670}
}
read the original abstract

The emergent phenomena of large foundation models have revolutionized natural language processing. However, evaluating these models presents significant challenges due to their size, capabilities, and deployment across diverse applications. Existing literature often focuses on individual aspects, such as benchmark performance or specific tasks, but fails to provide a cohesive process that integrates the nuances of diverse use cases with broader ethical and operational considerations. This work focuses on three key aspects: (1) Formalizing the Evaluation Process by providing a structured framework tailored to specific use-case contexts, (2) Offering Actionable Tools and Frameworks such as checklists and templates to ensure thorough, reproducible, and practical evaluations, and (3) Surveying Recent Work with a targeted review of advancements in LLM evaluation, emphasizing real-world applications.

Figures

Figures reproduced from arXiv: 2502.09670 by the authors.

Figure 1
Figure 1. Workflow of evaluation a comprehensive lens to view the multifaceted nature of LLM evaluation, offering a foundation for the more detailed discussions that follow. 2.1 Algorithm – Models LLMs can be classified into closed-source and open-source models, each with distinct traits influencing deployment, accessibility, and adaptability. Closed-source models (e.g., GPT1 , Claude2 , and Gemini3 ) are proprietary systems … view at source ↗
Figure 2
Figure 2. Dimensions of Evaluation across various domains. These benchmarks assess the rea￾soning capabilities of models, requiring them to comprehend context, infer unstated connections, and resolve ambiguities. Retrieval and ranking tasks test a model’s ability to identify and rank the most relevant documents or passages from a corpus given a query. Datasets like MS MARCO [5], TREC [84], and Natural Questions (NQ) [37] are … view at source ↗
Figure 3
Figure 3. Evaluation Methodologies These evaluations vary depending on the specific goals of each task, with distinct metrics tailored to capture perfor￾mance effectively. In the following, we outline common NLU tasks and the metrics used to evaluate them. In tasks such as sentiment analysis, topic classification, and named entity recognition (NER), models are assessed using metrics like accuracy, precision, recall, and F1-sc… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

    cs.CV 2026-04 unverdicted novelty 8.0 of 10

    DF3DV-1K supplies 1,048 scenes with clean and cluttered image pairs plus a challenging 41-scene subset to benchmark and improve distractor-free radiance field methods.

  2. Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch

    cs.LG 2025-11 conditional novelty 6.0 of 10

    Tree-Based Invariant Kernels fix the floating-point reduction order across GPUs, making LLM logits and sampled tokens bitwise identical for tensor-parallel sizes 1/2/4/8 and exactly matching vLLM (TP=4) with FSDP (TP=1).

Reference graph

Works this paper leans on

109 extracted references · 9 canonical work pages · cited by 2 Pith papers

  1. [1]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 tech- nical report. arXiv preprint arXiv:2303.08774 (2023)

  2. [2]

    Hizkiel Mitiku Alemayehu, Hamada M Zahera, and Axel- Cyrille Ngonga Ngomo. 2024. Error Analysis of Multilingual Lan- guage Models in Machine Translation: A Case Study of English- Amharic Translation. In Proceedings of the 2024 Conference on Empir- ical Methods in Natural Language Processing . 19758–19768

  3. [3]

    Zico Kolter, Matt Fredrikson, et al

    Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian, Derek Duenas, Maxwell Lin, Justin Wang, Dan Hendrycks, Andy Zou, Yuan et al. Zico Kolter, Matt Fredrikson, et al. 2024. Agentharm: A benchmark for measuring harmfulness of llm agents. arXiv preprint arXiv:2410.09024 (2024)

  4. [4]

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609 (2023)

  5. [5]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosenberg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2018. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. arXiv:1611.09268 [cs.CL] https: //arxiv.org/abs/1611.09268

  6. [6]

    Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrin- sic Evaluation Measures for Machine Translation and/or Summariza- tion, Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare Voss (Eds.). Association for Computati...

  7. [7]

    Barry Becker and Ronny Kohavi. 1996. Adult. UCI Machine Learning Repository. DOI: https://doi.org/10.24432/C5XW20

  8. [8]

    Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christo- pher D Manning. 2015. A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326 (2015)

Show all 109 references
  1. [9]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Ka- plan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al . 2020. Language Models are Few- Shot Learners. In Advances in Neural Information Processing Systems , Vol. 33. 1877–1901

  2. [10]

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2024. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology 15, 3 (2024), 1–45

  3. [11]

    Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, and Kathleen McKeown. 2023. Do models explain themselves? counterfactual simulatability of natural language explanations. arXiv preprint arXiv:2307.08678 (2023)

  4. [12]

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas An- gelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al . 2024. Chatbot arena: An open platform for evaluating llms by human preference. arXiv preprint arXiv:2403.04132 (2024)

  5. [13]

    George Chrysostomou and Nikolaos Aletras. 2021. Improving the faithfulness of attention-based explanations with task-specific infor- mation for text classification. arXiv preprint arXiv:2105.02657 (2021)

  6. [14]

    Nick Craswell. 2009. Mean Reciprocal Rank . Springer US, Boston, MA, 1703–1703. https://doi.org/10.1007/978-0-387-39940-9_488

  7. [15]

    Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: a standardized adversarial ro- bustness benchmark. arXiv:2010.09670 [cs.LG] https://arxiv.org/abs/ 2010.09670

  8. [16]

    Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al. 2024. Risk taxonomy, mitigation, and assessment benchmarks of large language model systems. arXiv preprint arXiv:2401.05778 (2024)

  9. [17]

    Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai. 2019. Bias in bios: A case study of semantic representation bias in a high-stakes setting. In proceedings of th...

  10. [18]

    Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace. 2019. ERASER: A benchmark to evaluate rationalized NLP models. arXiv preprint arXiv:1911.03429 (2019)

  11. [19]

    William Dieterich, Christina Mendoza, and Tim Brennan. 2016. COM- PAS risk scales: Demonstrating accuracy equity and predictive parity. Northpointe Inc 7, 4 (2016), 1–36

  12. [20]

    Esin Durmus, He He, and Mona Diab. 2020. FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization. In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics , Dan Jurafsky, Joyce Chai, Natalie S...

  13. [21]

    Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018. Hot- Flip: White-Box Adversarial Examples for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Com- putational Linguistics (Volume 2: Short Papers) , Iryna Gurevych and Yusuke Miy...

  14. [22]

    James Foulds, Rashidul Islam, Kamrun Naher Keya, and Shimei Pan. 2019. An Intersectional Definition of Fairness. arXiv:1807.08362 [cs.LG] https://arxiv.org/abs/1807.08362

  15. [23]

    Yonatan Geifman and Ran El-Yaniv. 2017. Selective Classification for Deep Neural Networks. arXiv:1705.08500 [cs.LG] https://arxiv.org/ abs/1705.08500

  16. [24]

    Tanya Goyal and Greg Durrett. 2020. Evaluating Factuality in Gen- eration with Dependency-level Entailment. In Findings of the As- sociation for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguis- tics, Onl...

  17. [25]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. arXiv:1706.04599 [cs.LG] https://arxiv.org/abs/1706.04599

  18. [26]

    Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Large lan- guage model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680 (2024)

  19. [27]

    Zishan Guo, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Linhao Yu, Yan Liu, Jiaxuan Li, Bojian Xiong, Deyi Xiong, et al. 2023. Evaluating large language models: A comprehensive survey. arXiv preprint arXiv:2310.19736 (2023)

  20. [28]

    Giwon Hong, Aryo Pradipta Gema, Rohit Saxena, Xiaotang Du, Ping Nie, Yu Zhao, Laura Perez-Beltrachini, Max Ryabinin, Xu- anli He, Clémentine Fourrier, and Pasquale Minervini. 2024. The Hallucinations Leaderboard - An Open Effort to Measure Halluci- nations in Large Language Mo...

  21. [29]

    Hsin-Yi Hsieh, Shih-Cheng Huang, and Richard Tsai. 2024. TWBias: A Benchmark for Assessing Social Bias in Traditional Chinese Large Language Models through a Taiwan Cultural Lens. In Findings of the Association for Computational Linguistics: EMNLP 2024 . 8688–8704

  22. [30]

    Simon Hughes, Minseok Bae, and Miaoran Li. 2023. Vectara Hal- lucination Leaderboard. https://github.com/vectara/hallucination- leaderboard

  23. [31]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bam- ford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mis- tral 7B. arXiv preprint arXiv:2310.06825 (2023)

  24. [32]

    Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is BERT Really Robust? A Strong Baseline for Natural Language Attack The Science of Evaluating Foundation Models on Text Classification and Entailment. arXiv:1907.11932 [cs.CL] https://arxiv.org/abs/1907.11932

  25. [33]

    Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer

  26. [34]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, An- ton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...

  27. [35]

    Earnshaw, Imran S

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, Tony Lee, Etienne David, Ian Stavness, Wei Guo, Berton A. Earnshaw, Imran S. Haque, Sara Beery, Jure Leskovec, A...

  28. [36]

    Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020. Evaluating the Factual Consistency of Abstractive Text Summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , Bonnie Webber, Trevor Cohn, Yul...

  29. [37]

    Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polo- sukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkor- eit, Quoc Le, and ...

  30. [38]

    Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. 2023. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702 (2023)

  31. [39]

    Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Miza- nur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, et al. 2024. A systematic survey and critical review on evaluating large language models: Challeng...

  32. [40]

    Mina Lee, Megha Srivastava, Amelia Hardy, John Thickstun, Esin Durmus, Ashwin Paranjape, Ines Gerard-Ursin, Xiang Lisa Li, Faisal Ladhak, Frieda Rong, et al. 2022. Evaluating human-language model interaction. arXiv preprint arXiv:2212.09746 (2022)

  33. [41]

    Noah Lee, Na Min An, and James Thorne. 2023. Can Large Language Models Capture Dissenting Human Voices? arXiv:2305.13788 [cs.CL] https://arxiv.org/abs/2305.13788

  34. [42]

    Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji- Rong Wen. 2023. Halueval: A large-scale hallucination evaluation benchmark for large language models.arXiv preprint arXiv:2305.11747 (2023)

  35. [43]

    Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan

  36. [45]

    Manning, Christo- pher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christo- pher Ré, Diana Acosta-Nav...

  37. [46]

    Yuanzhi Liang, Linchao Zhu, and Yi Yang. 2024. AntEval: Quanti- tatively Evaluating Informativeness and Expressiveness of Agent Social Interactions. arXiv preprint arXiv:2401.06509 (2024)

  38. [47]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evalua- tion of Summaries. In Text Summarization Branches Out . Associ- ation for Computational Linguistics, Barcelona, Spain, 74–81. https: //aclanthology.org/W04-1013/

  39. [48]

    Qingyu Lu, Baopu Qiu, Liang Ding, Kanjian Zhang, Tom Kocmi, and Dacheng Tao. 2023. Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models. arXiv preprint arXiv:2303.13809 (2023)

  40. [49]

    Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies . 142–150

  41. [50]

    Marta Marchiori Manerba, Karolina Stańczak, Riccardo Guidotti, and Isabelle Augenstein. 2023. Social bias probing: Fairness benchmarking for language models. arXiv preprint arXiv:2311.09090 (2023)

  42. [51]

    Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher

  43. [52]

    Moin Nadeem, Anna Bethke, and Siva Reddy. 2020. StereoSet: Measur- ing stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456 (2020)

  44. [53]

    Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al

  45. [54]

    Pointer sentinel mixture models.arXiv preprint arXiv:1609.07843 (2016)

  46. [55]

    Cohen, and Mirella Lapata

    Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In Proceedings of the Yuan et al. 2018 Conference on Empirical Methods in Natural Language Processing , El...

  47. [56]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Compu- tational Linguistics (ACL). 311–318

  48. [57]

    arXiv preprint arXiv:1602.06023 (2016)

    Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023 (2016)

  49. [58]

    Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman

  50. [59]

    Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013. Towards robust linguistic analysis using ontonotes. In Proceedings of the Seventeenth Conference on Computational Natural Language Learning. 143–152

  51. [60]

    Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michi- hiro Yasunaga, and Diyi Yang. 2023. Is ChatGPT a General-Purpose Natural Language Processing Task Solver? arXiv:2302.06476 [cs.CL] https://arxiv.org/abs/2302.06476

  52. [61]

    Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, and Dong Yu

  53. [62]

    Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman

  54. [63]

    Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Car- son Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jack- son Kernion, Kamil ˙e Lukoši ¯ut˙e, et al . 2023. Question decomposi- tion improves the faithfulness of model-generated reasoning. arXiv preprint arXi...

  55. [64]

    Ji-Lun Peng, Sijia Cheng, Egil Diau, Yung-Yu Shih, Po-Heng Chen, Yen-Ting Lin, and Yun-Nung Chen. 2024. A Survey of Useful LLM Evaluation. arXiv preprint arXiv:2406.00936 (2024)

  56. [65]

    Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang

  57. [66]

    Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, and Alan W Black. 2021. NoiseQA: Challenge Set Evaluation for User-Centric Question Answering. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational...

  58. [67]

    Rush, Sumit Chopra, and Jason Weston

    Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015. A Neural Attention Model for Abstractive Sentence Summarization. Proceed- ings of the 2015 Conference on Empirical Methods in Natural Language Processing (2015). https://doi.org/10.18653/v1/d15-1044

  59. [68]

    Erik F Sang and Fien De Meulder. 2003. Introduction to the CoNLL- 2003 shared task: Language-independent named entity recognition. arXiv preprint cs/0306050 (2003)

  60. [69]

    Han Qiu, Jiaxing Huang, Peng Gao, Qin Qi, Xiaoqin Zhang, Ling Shao, and Shijian Lu. 2024. LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models. arXiv preprint arXiv:2410.09962 (2024)

  61. [70]

    Yaozong Shen, Lijie Wang, Ying Chen, Xinyan Xiao, Jing Liu, and Hua Wu. 2022. An Interpretability Evaluation Benchmark for Pre-trained Language Models. arXiv preprint arXiv:2207.13948 (2022)

  62. [71]

    P Rajpurkar. 2016. Squad: 100,000+ questions for machine compre- hension of text. arXiv preprint arXiv:1606.05250 (2016)

  63. [72]

    Till Speicher, Hoda Heidari, Nina Grgic-Hlaca, Krishna P Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar. 2018. A unified approach to quantifying algorithmic unfairness: Measuring individual &group unfairness via inequality indices. In Proceedings of the 24th AC...

  64. [73]

    In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)

    SQuAD: 100,000+ Questions for Machine Comprehension of Text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.). Association for Computational Linguistics, Austin, Texas, 2383–2392. https...

  65. [74]

    Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li. 2024. Comparing Criteria Development Across Domain Experts, Lay Users, and Models in Large Language Model Evaluation. arXiv preprint arXiv:2410.02054 (2024)

  66. [75]

    Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V Stolyar, Katelyn Polanska, Karleigh R McCarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, et al. 2024. A framework for human evaluation of large language models in healthcare derived from literatu...

  67. [76]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al . 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)

  68. [77]

    Pola Schwöbel, Luca Franceschi, Muhammad Bilal Zafar, Keerthan Vasist, Aman Malhotra, Tomer Shenhar, Pinal Tailor, Pinar Yilmaz, Michael Diamond, and Michele Donini. 2024. Evaluating Large Lan- guage Models with fmeval. arXiv preprint arXiv:2407.12872 (2024)

  69. [78]

    Alex Wang. 2018. Glue: A multi-task benchmark and analy- sis platform for natural language understanding. arXiv preprint arXiv:1804.07461 (2018)

  70. [79]

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christo- pher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Re- cursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural lang...

  71. [80]

    Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2024. DecodingTrust: A Com...

  72. [81]

    Chongyan Sun, Ken Lin, Shiwei Wang, Hulong Wu, Chengfei Fu, and Zhen Wang. 2024. LalaEval: A Holistic Human Evaluation Frame- work for Domain-Specific Large Language Models. arXiv preprint arXiv:2408.13338 (2024)

  73. [82]

    Junyang Wang, Yuhang Wang, Guohai Xu, Jing Zhang, Yukai Gu, Haitao Jia, Ming Yan, Ji Zhang, and Jitao Sang. 2023. An llm-free The Science of Evaluating Foundation Models multi-dimensional benchmark for mllms hallucination evaluation. arXiv preprint arXiv:2311.07397 (2023)

  74. [83]

    Longyue Wang, Chenyang Lyu, Tianbo Ji, Zhirui Zhang, Dian Yu, Shuming Shi, and Zhaopeng Tu. 2023. Document-Level Machine Translation with Large Language Models. arXiv:2304.02210 [cs.CL] https://arxiv.org/abs/2304.02210

  75. [84]

    Smith, and Teruko Mitamura

    Mengqiu Wang, Noah A. Smith, and Teruko Mitamura. 2007. What is the Jeopardy Model? A Quasi-Synchronous Grammar for QA. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Lan- guage Learning (EMNLP-CoNLL), ...

  76. [85]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman

  77. [86]

    Advances in Neural Information Processing Systems 36 (2024)

    Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems 36 (2024)

  78. [87]

    Yicheng Wang, Jiayi Yuan, Yu-Neng Chuang, Zhuoer Wang, Yingchi Liu, Mark Cusick, Param Kulkarni, Zhengping Ji, Yasser Ibrahim, and Xia Hu. 2024. DHP Benchmark: Are LLMs Good NLG Evaluators? arXiv preprint arXiv:2408.13704 (2024)

  79. [88]

    Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020. Asking and Answering Questions to Evaluate the Factual Consistency of Sum- maries. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel T...

  80. [89]

    Adina Williams, Nikita Nangia, and Samuel R Bowman. 2017. A broad-coverage challenge corpus for sentence understanding through inference. arXiv preprint arXiv:1704.05426 (2017)

  81. [90]

    Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jian- feng Gao, Ahmed Hassan Awadallah, and Bo Li. 2022. Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Lan- guage Models. arXiv:2111.02840 [cs.CL] https://arxiv.org/abs/2111. 02840

  82. [91]

    Wangsong Yin, Mengwei Xu, Yuanchun Li, and Xuanzhe Liu

  83. [92]

    Jiayi Yuan, Ruixiang Tang, Xiaoqian Jiang, and Xia Hu. 2024. Large language models for healthcare data augmentation: An example on patient-trial matching. In AMIA Annual Symposium Proceedings, Vol. 2023. 1324

  84. [93]

    Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, et al. 2024. R-judge: Benchmarking safety risk awareness for llm agents. arXiv preprint arXiv:2401.10019 (2024)

  85. [94]

    Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, Qinzhuo Wu, Zhengyan Li, Chong Zhang, Ruotian Ma, Zichu Fei, Ruijian Cai, Jun Zhao, Xingwu Hu, Zhiheng Yan, Yiding Tan, Yuan Hu, Qiyuan Bian, Zhihua Liu, Shan Qin...

  86. [95]

    Yining Wang, Liwei Wang, Yuanzhi Li, Di He, Tie-Yan Liu, and Wei Chen. 2013. A Theoretical Analysis of NDCG Type Ranking Measures. arXiv:1304.6480 [cs.LG] https://arxiv.org/abs/1304.6480

  87. [96]

    Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems 28 (2015)

  88. [97]

    Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu. 2023. Align- ing large language models with human: A survey. arXiv preprint arXiv:2307.12966 (2023)

  89. [98]

    Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Zhenqiang Gong, and Xing Xie. 2024. PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts. arXiv:2306.04528 [cs.CL] https:/...

  90. [99]

    Lin Xu, Zhiyuan Hu, Daquan Zhou, Hongyu Ren, Zhen Dong, Kurt Keutzer, See Kiong Ng, and Jiashi Feng. 2024. Magic: Investigation of large language model powered multi-agent in cognition, adaptability, rationality and collaboration. In Proceedings of the 2024 Conference on Empir...

  91. [100]

    Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015. Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books. arXiv:1506.06724 [cs.CV] https://arxiv. org/abs/1506.06724

  92. [101]

    arXiv preprint arXiv:2403.11805 (2024)

    Llm as a system service on mobile devices. arXiv preprint arXiv:2403.11805 (2024)

  93. [104]

    Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen, Xiaofeng Mao, Longtao Huang, Hui Xue, Wenhai Wang, Kui Ren, and Jingyi Wang. 2024. S-Eval: Automatic and Adaptive Test Generation for Benchmarking Safety Evaluation of Large Language Models. arXiv preprint arXiv:2405.14191 (2024)

  94. [105]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. arXiv:1904.09675 [cs.CL] https://arxiv.org/abs/1904.09675

  95. [107]

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. 2024. Explainability for large language models: A survey.ACM Transactions on Intelligent Systems and Technology 15, 2 (2024), 1–38

  96. [109]

    Kaijie Zhu, Qinlin Zhao, Hao Chen, Jindong Wang, and Xing Xie. 2024. PromptBench: A Unified Library for Evaluation of Large Language Models. arXiv:2312.07910 [cs.AI] https://arxiv.org/abs/2312.07910

  97. [2016]

    A Diversity-Promoting Objective Function for Neural Conver- sation Models. InProceedings of the 2016 Conference of the North Amer- ican Chapter of the Association for Computational Linguistics: Human Language Technologies, Kevin Knight, Ani Nenkova, and Owen Ram- bow (Eds.). A...

  98. [2017]

    In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.)

    TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension. In Proceedings of the 55th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Regina Barzilay and Min-Yen Kan (Eds.). Associa- tion for Computatio...

  99. [2020]

    arXiv preprint arXiv:2010.00133 (2020)

    CrowS-pairs: A challenge dataset for measuring social biases in masked language models. arXiv preprint arXiv:2010.00133 (2020)

  100. [2021]

    arXiv preprint arXiv:2110.08193 (2021)

    BBQ: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193 (2021)

  101. [2024]

    arXiv preprint arXiv:2401.03601 (2024)

    Infobench: Evaluating instruction following ability in large language models. arXiv preprint arXiv:2401.03601 (2024)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.