Pith. sign in

REVIEW 3 major objections 6 minor 104 references

A Conceptual Framework for AI Capability Evaluations

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper proposes a descriptive seven-element framework for analyzing AI capability evaluations, arguing that it makes diverse evaluations transparent, comparable, and interpretable without imposing a new taxonomy.

desk verdict A genuinely useful and clearly written descriptive framework for AI capability evaluations whose comprehensiveness claim outruns its evidence; it deserves serious review, and would come out stronger with claims matched to evidence in revision. read the letter →

arxiv 2506.18213 v1 pith:7EUAVMOO submitted 2025-06-23 cs.AI

classification cs.AI
keywords AIcapabilityevaluationLLMframeworktransparencycomparabilitygovernancebenchmarkdesignmeta-evaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI capability evaluations—tests of what a system can and cannot do—are multiplying faster than the field can compare them, and governance decisions increasingly lean on them. This paper argues that any such evaluation can be analyzed through a shared descriptive framework made of seven elements: Evaluation Target, Task, Evaluated Subject, System Inputs, Evaluation Instance, Measurement, and Result Analysis, plus transversal elements such as practitioners, ethics, pilot tests, and ablation studies. Because the framework systematizes existing terminology rather than inventing a new taxonomy, the authors claim it makes diverse evaluations transparent, comparable, and interpretable for researchers, practitioners, and policymakers, and helps expose methodological gaps like missing statistical analysis or unreported prompts. A reader should care because regulators and deployers need a common way to scrutinize evidence about model capabilities and risks.

What carries the argument

The load-bearing object is the framework itself: a seven-element decomposition with named sub-elements (Evaluation Target, Task, Evaluated Subject, System Inputs, Evaluation Instance, Measurement, Result Analysis) and transversal elements (Evaluation Practitioners, Ethical Considerations, Pilot Tests, Ablation Studies). Each sub-element names a concrete design choice an evaluation makes, such as closed-ended versus open-ended task mode, identity conditions, prompt templates, evaluator selection, reference-based versus reference-free metrics, baselines, and significance testing. The framework does descriptive work: it places any evaluation onto a common grid so evaluations can be compared cell by cell, design gaps become visible, and results can be interpreted by non-specialists, without prescribing an order of steps or a rigid format.

What would settle it

Have independent coders apply the framework to a large, diverse sample of published capability evaluations and tally any substantive design choice that cannot be assigned to an element without stretching definitions, such as a dynamic test set that adapts to model responses or a metric defined over interaction trajectories. The claim of comprehensiveness would fail if such cases are frequent, or if a single published evaluation resists description in principle.

Watch

Extended reading notes

Core claim

The paper's central claim is that the unstructured landscape of AI capability evaluations can be mapped by a single conceptual framework that abstracts the essential features of an evaluation without prescribing how one should be run. The framework decomposes an evaluation into the capability and objective being targeted; the task specification, including task mode, steps, and interactions; the evaluated subject, including its identity conditions and operational context; the system inputs, including input source, test dataset, and prompt techniques; the evaluation instance, including evaluation criteria, evaluator subject, and evaluator task; measurement, including metric and baselines; and result analysis, including qualitative, statistical, and future-work components. Transversal elements cover evaluation practitioners, ethical considerations, pilot tests, and ablation studies. The authors adopt a broad definition of capability evaluations, exclude real-world impact evaluations, and demonstrate the framework's use by describing three published evaluations in an appendix, including an LLM-as-examiner benchmark, a red-teaming meta-evaluation, and a multi-agent code-generation study.

Load-bearing premise

Everything rests on the premise that the seven elements, identified through qualitative analysis of a non-exhaustive set of examples, genuinely cover every substantive aspect an AI capability evaluation can have, so no evaluation must be forced or distorted to fit.

Editorial extensions

If this is right

  • Different evaluations—say, a reasoning benchmark and a safety red-teaming study—become comparable element by element, so hidden differences in prompts, evaluators, metrics, and analysis are made explicit.
  • Methodological weaknesses surface as missing or underspecified elements, such as an evaluation that reports no statistical analysis, no prompt details, or no account of how evaluators were chosen.
  • Policymakers and auditors gain a plain-language checklist for scrutinizing evaluation reports, which matters as regulations begin to embed benchmarks into legal requirements.
  • The framework provides a natural organizing scheme for a living repository of evaluations, letting researchers track how open challenges such as data contamination are being addressed.
  • Practitioners designing new evaluations can use the seven elements as a design checklist, ensuring each load-bearing decision is made and reported.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors deliberately stop at description; a natural step they do not take is to turn the seven elements into a reporting standard—an 'evaluation card' that developers, third-party auditors, or regulators fill out, which would make the framework normative rather than descriptive.
  • Because the framework is distilled from existing evaluation practice, it may carry the blind spots of that practice; a genuinely novel evaluation paradigm that does not fit the seven elements would expose a limit of the framework rather than a flaw in the paradigm.
  • The framework's treatment of meta-evaluations—where the subject is the evaluation method itself—suggests it could serve as an audit tool: one could use it to check whether an evaluation's stated objective is actually supported by its task, metric, and analysis choices.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes a descriptive conceptual framework for analyzing AI capability evaluations, organized around seven core elements (Evaluation Target, Task, Evaluated Subject, System Inputs, Evaluation Instance, Measurement, Result Analysis) plus transversal elements (practitioners, ethics, pilot tests, ablation studies). The framework is built from existing terminology and illustrated on three published evaluations in Appendix 2. The authors claim that the framework supports transparency, comparability, and interpretability across diverse evaluations without imposing new taxonomies or rigid formats, while explicitly excluding impact evaluations.

Significance. If the claimed coverage holds, this framework is a useful synthesis: it integrates widely used evaluation concepts, connects them to documented challenges such as data contamination, prompt sensitivity, and statistical analysis, and makes evaluation structure accessible to non-specialists including policymakers. The paper's strengths are its careful reuse of existing terminology, its explicit integration of evaluation challenges into the element structure, and the worked examples in Appendix 2. However, the central promise—that the seven elements capture the full range of aspects an evaluation method may encompass—is asserted rather than demonstrated. Because transparency and comparability across evaluations depend on this coverage, the framework's significance is conditional on a validation step that the paper itself defers to future work.

major comments (3)
  1. [Section 1] The claim that the framework is 'intended to be comprehensive—capturing the full range of aspects that an evaluation method may encompass' is load-bearing for the paper's stated benefits, but the only evidence offered is a qualitative reading of a non-exhaustive set of evaluations. Section 4 explicitly postpones systematic validation by proposing to 'analyze a range of evaluations using this framework as a set of use cases and to iteratively refine it.' As written, the comprehensiveness claim is a plausible hypothesis, not an established result; the paper should either weaken the claim to reflect the current evidence or provide a systematic methodology, such as a structured corpus analysis with coverage statistics and failure cases.
  2. [Appendix 2 and Section 4] The three use cases in Appendix 2 are self-admittedly 'not exhaustive' and 'simplified,' and they are all described by the authors in terms that fit the framework without tension. There is no discussion of borderline or adversarial cases, such as simulator-based environments, dynamic multi-agent protocols, or elicitation procedures that cut across System Inputs and Operational Context. Without such cases, the paper does not establish that the framework can represent diverse evaluations without forcing aspects into ill-fitting categories or silently dropping them. A concrete test would be to apply the framework to a balanced corpus of evaluations—including agent evaluations, red-teaming, meta-evaluations, and benchmark-based evaluations—and report inter-annotator agreement, coverage failures, and revisions made in response.
  3. [Section 3 (element boundaries)] The boundaries between several elements are not operationally defined. For example, prompt techniques are placed under System Inputs (Section 3.4), while access methods, configuration settings, and auxiliary tools are placed under Operational Context (Section 3.3), and multi-step task structure is placed under Task (Section 3.2). The paper does not provide decision rules for assigning a given evaluation feature to one element rather than another, and it does not discuss how to handle features that plausibly span multiple elements. Because comparability across evaluations requires that different analysts map the same evaluation onto the same structure, the absence of an annotation protocol or inter-rater reliability evidence is a gap in the central claim.
minor comments (6)
  1. [Appendix 2 table] The row for 'Evaluated Subject' in Paper 1 contains the typo 'descentralized' (should be 'decentralized'), and the term appears several times in the table.
  2. [Appendix 2 table, Measurement] The phrase 'Liket Scale' should be 'Likert Scale,' and 'the evaluator was also asked' should be 'the evaluators were also asked.'
  3. [Appendix A.7] The text contains 'ANOV A' with a spurious space; this should read 'ANOVA.'
  4. [Appendix B] The table cites 'Vidgen et al., 2021' for the RoBERTa hate speech classifier, but this reference is missing from the reference list.
  5. [Section 3.6] The phrase 'diverity' appears in the Appendix 2 table under the 'Diversity' metric; this should be corrected to 'diversity.'
  6. [Global] The paper uses footnotes for broad definitions of capability evaluations, but some footnotes extend to multiple lines and could be moved to the main text or made more concise for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the framework is a descriptive synthesis of external evaluation literature, with no fitted inputs, derived predictions, or load-bearing self-citations.

full rationale

The paper does not derive a quantitative result from fitted parameters, and it contains no equations whose outputs are their own inputs. Its central product is a descriptive conceptual framework whose seven elements and transversal components are assembled from externally published work and from the authors' qualitative reading of existing evaluations; Appendix 2 then applies that framework to three published papers as illustrative use cases. The claimed benefit of the framework—supporting transparency, comparability, and interpretability—is presented as a design goal and as an assertion about the usefulness of the framework, not as a prediction inferred from the same framework. The paper explicitly acknowledges that its survey of evaluations is non-exhaustive and that future validation through additional use cases is needed, which is a limitation in evidence for comprehensiveness rather than a circular derivation. No step in the paper reduces a stated result to its own definition, no parameter is fitted and then renamed as a prediction, and no load-bearing premise is justified solely by a self-citation. Accordingly, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities: the framework introduces no new constructs with independent empirical handles, only a reorganization of existing concepts. The main assumptions are about comprehensiveness, usefulness of the descriptive stance, and integrability of existing terminology. No new repository, metric, or evaluation method is built.

assumptions (4)
  • domain assumption The seven-element decomposition is comprehensive enough to capture the full range of evaluation aspects.
    Section 1 states the framework is 'intended to be comprehensive,' but no systematic selection or validation procedure is given.
  • domain assumption A descriptive, non-sequential framework is more useful for analysis than a prescriptive sequential one.
    Section 2 criticizes Laskar et al. (2024) and Paskov et al. (2025) for prescribing a sequence, implicitly assuming non-sequential mapping is preferable.
  • domain assumption Existing terminology from cited works can be coherently integrated without loss of meaning.
    The framework merges terms from Harding & Sharadin (2024), Burden et al. (2025), Reuel et al. (2024a), and others; the paper does not test inter-source consistency.
  • domain assumption The three Appendix 2 use cases are representative enough to illustrate applicability.
    Appendix 2 states the selection 'is not exhaustive and some details may be simplified or open to interpretation.'

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Conceptual Framework for AI Capability Evaluations." pith.science (2026). https://pith.science/paper/7EUAVMOO

@misc{pith2026250618213,
  author       = {Pith},
  title        = {Pith review of: A Conceptual Framework for AI Capability Evaluations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7EUAVMOO}},
  note         = {Machine review of arXiv:2506.18213}
}
read the original abstract

As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence about system capabilities and risks. Yet there remains a lack of clarity on how to perform these assessments both comprehensively and reliably. To address this gap, we propose a conceptual framework for analyzing AI capability evaluations, offering a structured, descriptive approach that systematizes the analysis of widely used methods and terminology without imposing new taxonomies or rigid formats. This framework supports transparency, comparability, and interpretability across diverse evaluations. It also enables researchers to identify methodological weaknesses, assists practitioners in designing evaluations, and provides policymakers with an accessible tool to scrutinize, compare, and navigate complex evaluation landscapes.

Figures

Figures reproduced from arXiv: 2506.18213 by the authors.

Figure 1
Figure 1. Overview of the proposed Conceptual Framework. 3.3. Evaluated Subject The evaluated subject is the subject of the capability claim—in most cases, an AI system, and in the case of meta-evaluations, an evaluation method itself (e.g. Hong et al. (2024)). Identity Conditions. The evaluated system should be in￾dividualized by defining its identity conditions (Harding & Sharadin, 2024; Biderman et al., 2024). This include… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

104 extracted references · 51 canonical work pages

  1. [1]

    Early insights from developing question-answer evaluations for frontier AI , 2024

    AISI. Early insights from developing question-answer evaluations for frontier AI , 2024. Accessed May 10, 2025

  2. [2]

    Benchmarking foundation models with language-model-as-an-examiner

    Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al. Benchmarking foundation models with language-model-as-an-examiner. Advances in Neural Information Processing Systems, 36: 0 78142--78167, 2023

  3. [3]

    Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation

    Barnett, P. and Thiergart, L. Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation. arXiv preprint arXiv:2411.12820, 2024

  4. [4]

    A quantitative study of nlp approaches to question difficulty estimation

    Benedetto, L. A quantitative study of nlp approaches to question difficulty estimation. In International Conference on Artificial Intelligence in Education, pp.\ 428--434. Springer, 2023

  5. [5]

    Evaluating ai for law: Bridging the gap with open-source solutions

    Bhambhoria, R., Dahan, S., Li, J., and Zhu, X. Evaluating ai for law: Bridging the gap with open-source solutions. arXiv preprint arXiv:2404.12349, 2024

  6. [6]

    F., Ammanamanchi, P

    Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., et al. Lessons from the trenches on reproducible evaluation of language models. CoRR, 2024

  7. [7]

    R., Steunebrink, B

    Bieger, J., Th \'o risson, K. R., Steunebrink, B. R., and Thorarensen, T. Evaluation of general-purpose artificial intelligence: Why, what & how. In Proceedings of the IJCAI Workshop on Evaluating General-Purpose Artificial Intelligence (EGPAI 2016), 2016. URL https://alumni.media.mit.edu/ kris/ftp/EGPAI_2016_paper_9.pdf. Accessed May 10, 2025

  8. [8]

    T., Li, Y., Lundberg, S., et al

    Bubeck, S., Chadrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023

Show all 104 references
  1. [9]

    Evaluating ai evaluation: Perils and prospects

    Burden, J. Evaluating ai evaluation: Perils and prospects. arXiv preprint arXiv:2407.09221, 2024

  2. [10]

    Paradigms of ai evaluation: Mapping goals, methodologies and culture

    Burden, J., Te s i \'c , M., Pacchiardi, L., and Hern \'a ndez-Orallo, J. Paradigms of ai evaluation: Mapping goals, methodologies and culture. arXiv preprint arXiv:2502.15620, 2025

  3. [11]

    R., and Cheung, S.-C

    Cao, J., Chan, Y.-K., Ling, Z., Wang, W., Li, S., Liu, M., Qiao, R., Han, Y., Wang, C., Yu, B., He, P., Wang, S., Zheng, Z., Lyu, M. R., and Cheung, S.-C. How should we build a benchmark? revisiting 274 code-related benchmarks for llms, 2025. URL https://arxiv.org/abs/2501.10711

  4. [12]

    L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al

    Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al. Black-box access is insufficient for rigorous ai audits. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, ...

  5. [13]

    Cheng, Y., Georgopoulos, M., Cevher, V., and Chrysos, G. G. Leveraging the context through multi-round interactions for jailbreaking attacks. arXiv preprint arXiv:2402.09177, 2024

  6. [14]

    N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J

    Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., et al. Chatbot arena: An open platform for evaluating llms by human preference. In Forty-first International Conference on Machine Learning, 2024

  7. [15]

    On the limitations of reference-free evaluations of generated text

    Deutsch, D., Dror, R., and Roth, D. On the limitations of reference-free evaluations of generated text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp.\ 10960--10977, 2022

  8. [16]

    R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M

    Didolkar, A., Goyal, A., Ke, N. R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M. C., and Arora, S. Metacognitive capabilities of llms: An exploration in mathematical problem solving. Advances in Neural Information Processing Systems, 37: 0 1978...

  9. [17]

    Generalization or memorization: Data contamination and trustworthy evaluation for large language models

    Dong, Y., Jiang, X., Liu, H., Jin, Z., Gu, B., Yang, M., and Li, G. Generalization or memorization: Data contamination and trustworthy evaluation for large language models. In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 12039--12050, 2024

  10. [18]

    W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H

    Dow, A., Vaughan, J. W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H. Dimensions of generative ai evaluation design. In Proceedings of the NeurIPS 2024 Workshop on Evaluating Evaluations: Examining Best Practices for Measuring Broader Impacts of Generative AI, 20...

  11. [19]

    Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation

    Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation. arXiv preprint arXiv:2502.06559, 2025

  12. [20]

    Second draft of the general purpose AI code of practice, April 2024

    European Commission . Second draft of the general purpose AI code of practice, April 2024. Written by independent experts. Accessed May 10, 2025

  13. [21]

    Issue brief: Early best practices for frontier AI safety evaluations, 2024

    Frontier Model Forum . Issue brief: Early best practices for frontier AI safety evaluations, 2024. Accessed May 10, 2025

  14. [22]

    Llm-based nlg evaluation: Current status and challenges

    Gao, M., Hu, X., Yin, X., Ruan, J., Pu, X., and Wan, X. Llm-based nlg evaluation: Current status and challenges. Computational Linguistics, pp.\ 1--28, 2025

  15. [23]

    A case for better evaluation standards in nlg

    Gehrmann, S., Clark, E., and Sellam, T. A case for better evaluation standards in nlg. In Workshop on Setting up ML Evaluation Standards to Accelerate Progress at ICLR 2022, 2022. URL https://iclr.cc/virtual/2022/7328. Poster presentation

  16. [24]

    Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text

    Gehrmann, S., Clark, E., and Sellam, T. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. Journal of Artificial Intelligence Research, 77: 0 103--166, 2023

  17. [25]

    Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models

    Guha, N., Nyarko, J., Ho, D., R \'e , C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing S...

  18. [26]

    R., Hullman, J., and Subramonyam, H

    Gupta, N. R., Hullman, J., and Subramonyam, H. A conceptual framework for ethical evaluation of machine learning systems. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp.\ 534--546, 2024

  19. [27]

    Deception abilities emerged in large language models

    Hagendorff, T. Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences, 121 0 (24): 0 e2317967121, 2024

  20. [28]

    Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., and Schulz, E. Machine psychology, 2024. URL https://arxiv.org/abs/2303.13988

  21. [29]

    a m \"a l \

    H \"a m \"a l \"a inen, M. and Alnajjar, K. Human evaluation of creative NLG systems: An interdisciplinary survey on recent papers. In Bosselut, A., Durmus, E., Gangal, V. P., Gehrmann, S., Jernite, Y., Perez-Beltrachini, L., Shaikh, S., and Xu, W. (eds.), Proceedings of the 1...

  22. [30]

    and Sharadin, N

    Harding, J. and Sharadin, N. What is it for a machine learning model to have a capability? The British Journal for the Philosophy of Science, 2024. Advance online publication. Available at https://doi.org/10.1086/732153

  23. [31]

    Hofst \"a tter, F., Teoh, J., van der Weij, T., and Ward, F. R. The elicitation game: Stress-testing capability elicitation techniques. In Workshop on Socially Responsible Language Modelling Research, 2024. URL https://openreview.net/forum?id=zy6LB5t62f

  24. [32]

    R., Srivastava, A., and Agrawal, P

    Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J. R., Srivastava, A., and Agrawal, P. Curiosity-driven red-teaming for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?...

  25. [33]

    and Zhou, X.-H

    Hu, T. and Zhou, X.-H. Unveiling llm evaluation focused on metrics: Challenges and solutions. arXiv preprint arXiv:2404.09135, 2024

  26. [34]

    On the limitations of fine-tuned judge models for llm evaluation

    Huang, H., Qu, Y., Zhou, H., Liu, J., Yang, M., Xu, B., and Zhao, T. On the limitations of fine-tuned judge models for llm evaluation. arXiv preprint arXiv:2403.02839, 2024

  27. [35]

    M ath P rompter: Mathematical reasoning using large language models

    Imani, S., Du, L., and Shrivastava, H. M ath P rompter: Mathematical reasoning using large language models. In Sitaram, S., Beigman Klebanov, B., and Williams, J. D. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Indu...

  28. [36]

    Reference-free evaluation metrics for text generation: A survey

    Ito, T., van Deemter, K., and Suzuki, J. Reference-free evaluation metrics for text generation: A survey. arXiv preprint arXiv:2501.12011, 2025

  29. [37]

    Ivanova, A. A. Running cognitive evaluations on large language models: The do's and the don'ts. arXiv preprint arXiv:2312.01276, 2023

  30. [38]

    Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks

    Jacovi, A., Caciularu, A., Goldman, O., and Goldberg, Y. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp.\ 507...

  31. [39]

    Cladder: assessing causal reasoning in language models

    Jin, Z., Chen, Y., Leeb, F., Gresele, L., Kamal, O., Lyu, Z., Blin, K., Gonzalez, F., Kleiman-Weiner, M., Sachan, M., and Sch\" o lkopf, B. Cladder: assessing causal reasoning in language models. In Proceedings of the 37th International Conference on Neural Information Process...

  32. [40]

    T., and Sch \"o lkopf, B

    Jin, Z., Liu, J., LYU, Z., Poff, S., Sachan, M., Mihalcea, R., Diab, M. T., and Sch \"o lkopf, B. Can large language models infer causation from correlation? In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=vqIH0ObdqL

  33. [41]

    R., Rockt \"a schel, T., and Perez, E

    Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., Grefenstette, E., Bowman, S. R., Rockt \"a schel, T., and Perez, E. Debating with more persuasive llms leads to more truthful answers. In Proceedings of the 41st International Conference on Machine L...

  34. [42]

    Causal reasoning and large language models: Opening a new frontier for causality

    Kiciman, E., Ness, R., Sharma, A., and Tan, C. Causal reasoning and large language models: Opening a new frontier for causality. Transactions on Machine Learning Research, 2023

  35. [43]

    Ai agent governance: A field guide

    Kraprayoon, J., Williams, Z., and Fayyaz, R. Ai agent governance: A field guide. arXiv preprint arXiv:2505.21808, 2025

  36. [44]

    Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation

    Kuhn, L., Gal, Y., and Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=VD-AYtP0dve

  37. [45]

    P., Wu, H., and Yu, H

    Lalor, J. P., Wu, H., and Yu, H. Building an evaluation scale using item response theory. In Proceedings of the conference on empirical methods in natural language processing. Conference on empirical methods in natural language processing, volume 2016, pp.\ 648, 2016

  38. [46]

    Laskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H., Jahan, I., Bhuiyan, A., Tan, C. W., Parvez, M. R., et al. A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations. In Proceedin...

  39. [47]

    J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M

    Lee, S., Kim, M., Cherif, L., Dobre, D., Lee, J., Hwang, S. J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M. Learning diverse attacks on large language models for robust red-teaming and safety tuning. In Red Teaming GenAI: What Can We Learn from Adversaries?,...

  40. [48]

    Leveraging large language models for nlg evaluation: Advances and challenges

    Li, Z., Xu, X., Shen, T., Xu, C., Gu, J.-C., Lai, Y., Tao, C., and Ma, S. Leveraging large language models for nlg evaluation: Advances and challenges. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 16028--16045, 2024

  41. [49]

    D., Re, C., Acosta-Navas, D., Hudson, D

    Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F....

  42. [50]

    Liao, Q. V. and Xiao, Z. Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100, 2023

  43. [51]

    D., and Schmidt, L

    Liao, T., Taori, R., Raji, I. D., and Schmidt, L. Are we learning yet? a meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/fo...

  44. [52]

    Against the achilles' heel: A survey on red teaming for generative models

    Lin, L., Mu, H., Zhai, Z., Wang, M., Wang, Y., Wang, R., Gao, J., Zhang, Y., Che, W., Baldwin, T., et al. Against the achilles' heel: A survey on red teaming for generative models. Journal of Artificial Intelligence Research, 82: 0 687--775, 2025

  45. [53]

    Datasets for large language models: A comprehensive survey

    Liu, Y., Cao, J., Liu, C., Ding, K., and Jin, L. Datasets for large language models: A comprehensive survey. arXiv preprint arXiv:2402.18041, 2024

  46. [54]

    R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M

    McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M. N. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. arXiv preprint arXiv:2402.09880, 2024

  47. [55]

    W., and Meisen, T

    Meyes, R., Lu, M., de Puiseau, C. W., and Meisen, T. Ablation studies in artificial neural networks. arXiv preprint arXiv:1901.08644, 2019

  48. [56]

    Adding error bars to evals: A statistical approach to language model evaluations

    Miller, E. Adding error bars to evals: A statistical approach to language model evaluations. arXiv preprint arXiv:2411.00640, 2024

  49. [57]

    Auditing large language models: a three-layered approach

    M \"o kander, J., Schuett, J., Kirk, H., and Floridi, L. Auditing large language models: a three-layered approach. AI and Ethics, 4 0 (4), 2023

  50. [58]

    Evaluating the performance of large language models via debates

    Moniri, B., Hassani, H., and Dobriban, E. Evaluating the performance of large language models via debates. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, pp.\ 2040--2075, Albuquerque, New Mexico, April 2...

  51. [59]

    and Kapoor, S

    Narayanan, A. and Kapoor, S. Gpt-4 and professional benchmarks: The wrong answer to the right question. AI Snake Oil (Substack), March 2023. URL https://www.aisnakeoil.com/p/gpt-4-and-professional-benchmarks. Accessed May 10, 2025

  52. [60]

    Oecd framework for the classification of ai systems

    OECD. Oecd framework for the classification of ai systems. OECD Digital Economy Papers, No. 323, 2022

  53. [61]

    and Kang, E

    Orr, W. and Kang, E. B. Ai as a sport: On the competitive epistemologies of benchmarking. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 1875--1884, 2024

  54. [62]

    Llm evaluators recognize and favor their own generations

    Panickssery, A., Bowman, S., and Feng, S. Llm evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems, 37: 0 68772--68802, 2024

  55. [63]

    T., and Soder, L

    Paskov, P., Berglund, L., Smith, E. T., and Soder, L. Gpai evaluations standards taskforce: towards effective ai governance. In Workshop on Socially Responsible Language Modelling Research, 2024

  56. [64]

    Preliminary suggestions for rigorous gpai model evaluations

    Paskov, P., Byun, M., Wei, K., and Webster, T. Preliminary suggestions for rigorous gpai model evaluations. Technical Report PEA3971-1, RAND Corporation, April 2025. URL https://www.rand.org/pubs/perspectives/PEA3971-1.html. Accessed May 10, 2025

  57. [65]

    Discovering language model behaviors with model-written evaluations

    Perez, E., Ringer, S., Lukosiute, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pp.\ 133...

  58. [66]

    and Jud, H

    Pfister, R. and Jud, H. Understanding and benchmarking artificial intelligence: Openai's o3 is not agi. arXiv preprint arXiv:2501.07458, 2025

  59. [67]

    The roots search tool: Data transparency for llms

    Piktus, A., Akiki, C., Villegas, P., Lauren c on, H., Dupont, G., Luccioni, S., Jernite, Y., and Rogers, A. The roots search tool: Data transparency for llms. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstra...

  60. [68]

    D., Denton, E., Bender, E

    Raji, I. D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. AI and the everything in the whole wide world benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=j...

  61. [69]

    Large language model evaluation via multi ai agents: Preliminary results

    Rasheed, Z., Waseem, M., Syst \"a , K., and Abrahamsson, P. Large language model evaluation via multi ai agents: Preliminary results. In International Conference on Learning Representations, pp.\ 1--12, 2024

  62. [70]

    A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al

    Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al. Gaps in the safety evaluation of generative ai. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp.\...

  63. [71]

    Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices

    Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024 ...

  64. [72]

    Reuel, A., Soder, L., Bucknall, B., and Undheim, T. A. Position: technical research and talent is needed for effective ai governance. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024 b

  65. [73]

    S., Rajkumar, N., Moës, N., Ladish, J., Bau, D., Bricman, P., Guha, N., Newman, J., Bengio, Y., South, T., Pentland, A., Koyejo, S., Kochenderfer, M

    Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., Luccioni, A. S., Rajkumar, N., Moës, N., Ladi...

  66. [74]

    Better than random: reliable nlg human evaluation with constrained active sampling

    Ruan, J., Pu, X., Gao, M., Wan, X., and Zhu, Y. Better than random: reliable nlg human evaluation with constrained active sampling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp.\ 18915--18923, 2024

  67. [75]

    K., Saha, S., Jain, V., Mondal, S., and Chadha, A

    Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. A systematic survey of prompt engineering in large language models: Techniques and applications. CoRR, abs/2402.07927, 2024. URL https://doi.org/10.48550/arXiv.2402.07927

  68. [76]

    L., and Agirre, E

    Sainz, O., Campos, J., Garc \' a-Ferrero, I., Etxaniz, J., de Lacalle, O. L., and Agirre, E. Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp.\ 10776--10787, 2023

  69. [77]

    Targeting the benchmark: On methodology in current natural language processing research

    Schlangen, D. Targeting the benchmark: On methodology in current natural language processing research. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume ...

  70. [78]

    S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G

    Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., Li, Y., Gupta, A., Han, H., Schulhoff, S., Dulepet, P. S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G. C., Li, F., Tao, H., Srivastava, A., Costa, H. D., Gupta, S., Rogers, M. L., Goncearenco, I.,...

  71. [79]

    Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

    Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations, 2024. URL https://op...

  72. [80]

    Model evaluation for extreme risks

    Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023

  73. [81]

    CHOPS : CH at with customer profile systems for customer service with LLM s

    Shi, J., Li, J., Ma, Q., Yang, Z., Ma, H., and Li, L. CHOPS : CH at with customer profile systems for customer service with LLM s. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=9Wmdk94oKF

  74. [82]

    Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms

    Sirdeshmukh, V., Deshpande, K., Mols, J., Jin, L., Cardona, E.-Y., Lee, D., Kritz, J., Primack, W., Yue, S., and Xing, C. Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms. arXiv preprint arXiv:2501.17399, 2025

  75. [83]

    K., Grundy, E

    Slattery, P., Saeri, A. K., Grundy, E. A., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., and Thompson, N. The ai risk repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. CoRR, 2024

  76. [84]

    A study of translation edit rate with targeted human annotation

    Snover, M., Dorr, B., Schwartz, R., Micciulla, L., and Makhoul, J. A study of translation edit rate with targeted human annotation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers, pp.\ 223--231, 2006

  77. [85]

    Audit cards: Contextualizing ai evaluations

    Staufer, L., Yang, M., Reuel, A., and Casper, S. Audit cards: Contextualizing ai evaluations. arXiv preprint arXiv:2504.13839, 2025

  78. [86]

    Comprehensive reassessment of large-scale evaluation outcomes in llms: A multifaceted statistical approach

    Sun, K., Wang, R., Liu, H., and Søgaard, A. Comprehensive reassessment of large-scale evaluation outcomes in llms: A multifaceted statistical approach. CoRR, abs/2403.15250, 2024. URL https://doi.org/10.48550/arXiv.2403.15250

  79. [87]

    Measuring data science automation: A survey of evaluation tools for ai assistants and agents

    Testini, I., Hern \'a ndez-Orallo, J., and Pacchiardi, L. Measuring data science automation: A survey of evaluation tools for ai assistants and agents. arXiv preprint arXiv:2506.08800, 2025

  80. [88]

    Thurnherr, B. C. Who should develop which ai evaluations?, April 2024. Accessed May 10, 2025

  81. [89]

    Best practices for the human evaluation of automatically generated text

    Van Der Lee, C., Gatt, A., Van Miltenburg, E., Wubben, S., and Krahmer, E. Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation, pp.\ 355--368, 2019

  82. [90]

    M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S

    Vania, C., Htut, P. M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S. R. Comparing test sets with item response theory. In Annual Meeting of the Association for Computational Linguistics, 2021

  83. [91]

    Mint: Evaluating llms in multi-turn interaction with tools and language feedback

    Wang, X., Wang, Z., Liu, J., Chen, Y., Yuan, L., Peng, H., and Ji, H. Mint: Evaluating llms in multi-turn interaction with tools and language feedback. In 12th International Conference on Learning Representations, ICLR 2024, 2024

  84. [92]

    A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al

    Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al. Sociotechnical safety evaluation of generative ai systems. arXiv preprint arXiv:2310.11986, 2023

  85. [93]

    D., Wallach, H., Mitchell, M., Wang, A., Salaudeen, O., Bommasani, R., Ganguli, D., Koyejo, S., and Isaac, W

    Weidinger, L., Raji, I. D., Wallach, H., Mitchell, M., Wang, A., Salaudeen, O., Bommasani, R., Ganguli, D., Koyejo, S., and Isaac, W. Toward an evaluation science for generative ai systems. arXiv preprint arXiv:2503.05336, 2025

  86. [94]

    An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping

    Xia, B., Lu, Q., Zhu, L., and Xing, Z. An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping. In Proceedings of the 1st ACM International Conference on AI-Powered Software, pp.\ 74--78, 2024

  87. [95]

    A critical review of causal inference benchmarks for large language models

    Yang, L., Clivio, O., Shirvaikar, V., and Falck, F. A critical review of causal inference benchmarks for large language models. In AAAI 2024 Workshop on ''Are Large Language Models Simply Causal Parrots?'', 2023. URL https://openreview.net/forum?id=mRwgczYZFJ

  88. [96]

    Evaluatology: The science and engineering of evaluation

    Zhan, J., Wang, L., Gao, W., Li, H., Wang, C., Huang, Y., Li, Y., Yang, Z., Kang, G., Luo, C., Ye, H., Dai, S., and Zhang, Z. Evaluatology: The science and engineering of evaluation. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 4 0 (1): 0 100162, 2024. I...

  89. [97]

    K., Klyman, K., Mai, Y., Levine, Y., Zhang, Y., Bommasani, R., and Liang, P

    Zhang, A. K., Klyman, K., Mai, Y., Levine, Y., Zhang, Y., Bommasani, R., and Liang, P. Language model developers should report train-test overlap. arXiv preprint arXiv:2410.08385, 2024 a

  90. [98]

    Q., Shaw, R., Anthis, J

    Zhang, A. Q., Shaw, R., Anthis, J. R., Milton, A., Tseng, E., Suh, J., Ahmad, L., Kumar, R. S. S., Posada, J., Shestakofsky, B., et al. The human factor in ai red teaming: Perspectives from social and collaborative computing. In Companion Publication of the 2024 Conference on ...

  91. [99]

    Pacost: Paired confidence significance testing for benchmark contamination detection in large language models

    Zhang, H., Lin, Y., and Wan, X. Pacost: Paired confidence significance testing for benchmark contamination detection in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp.\ 1794--1809, 2024 c

  92. [100]

    and Kanayet, F

    Zhang, Y. and Kanayet, F. Genai evaluation maturity framework (gemf) to assess and improve genai evaluations. In Workshop on Evaluating Evaluations: Examining Best Practices for Measuring Broader Impacts of Generative AI at NeurIPS 2024 , 2024. URL https://neurips.cc/virtual/2...

  93. [101]

    Llmeval: A preliminary study on how to evaluate large language models

    Zhang, Y., Zhang, M., Yuan, H., Liu, S., Shi, Y., Gui, T., Zhang, Q., and Huang, X. Llmeval: A preliminary study on how to evaluate large language models. Proceedings of the AAAI Conference on Artificial Intelligence, 38 0 (17): 0 19615--19622, Mar. 2024 d . doi:10.1609/aaai.v...

  94. [102]

    L., Trischler, A., Daum \'e III, H., Suleman, K., and Olteanu, A

    Zhou, K., Blodgett, S. L., Trischler, A., Daum \'e III, H., Suleman, K., and Olteanu, A. Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computa...

  95. [103]

    X., Chen, X., Lin, Y., Wen, J.-R., and Han, J

    Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J. Don't make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023

  96. [104]

    Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks

    Zhou, X., Weyssow, M., Widyasari, R., Zhang, T., He, J., Lyu, Y., Chang, J., Zhang, B., Huang, D., and Lo, D. Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks. arXiv preprint arXiv:2502.06215, 2025

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.