REVIEW 3 major objections 6 minor 104 references
A Conceptual Framework for AI Capability Evaluations
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper proposes a descriptive seven-element framework for analyzing AI capability evaluations, arguing that it makes diverse evaluations transparent, comparable, and interpretable without imposing a new taxonomy.
desk verdict A genuinely useful and clearly written descriptive framework for AI capability evaluations whose comprehensiveness claim outruns its evidence; it deserves serious review, and would come out stronger with claims matched to evidence in revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the framework itself: a seven-element decomposition with named sub-elements (Evaluation Target, Task, Evaluated Subject, System Inputs, Evaluation Instance, Measurement, Result Analysis) and transversal elements (Evaluation Practitioners, Ethical Considerations, Pilot Tests, Ablation Studies). Each sub-element names a concrete design choice an evaluation makes, such as closed-ended versus open-ended task mode, identity conditions, prompt templates, evaluator selection, reference-based versus reference-free metrics, baselines, and significance testing. The framework does descriptive work: it places any evaluation onto a common grid so evaluations can be compared cell by cell, design gaps become visible, and results can be interpreted by non-specialists, without prescribing an order of steps or a rigid format.
What would settle it
Have independent coders apply the framework to a large, diverse sample of published capability evaluations and tally any substantive design choice that cannot be assigned to an element without stretching definitions, such as a dynamic test set that adapts to model responses or a metric defined over interaction trajectories. The claim of comprehensiveness would fail if such cases are frequent, or if a single published evaluation resists description in principle.
Extended reading notes
Core claim
The paper's central claim is that the unstructured landscape of AI capability evaluations can be mapped by a single conceptual framework that abstracts the essential features of an evaluation without prescribing how one should be run. The framework decomposes an evaluation into the capability and objective being targeted; the task specification, including task mode, steps, and interactions; the evaluated subject, including its identity conditions and operational context; the system inputs, including input source, test dataset, and prompt techniques; the evaluation instance, including evaluation criteria, evaluator subject, and evaluator task; measurement, including metric and baselines; and result analysis, including qualitative, statistical, and future-work components. Transversal elements cover evaluation practitioners, ethical considerations, pilot tests, and ablation studies. The authors adopt a broad definition of capability evaluations, exclude real-world impact evaluations, and demonstrate the framework's use by describing three published evaluations in an appendix, including an LLM-as-examiner benchmark, a red-teaming meta-evaluation, and a multi-agent code-generation study.
Load-bearing premise
Everything rests on the premise that the seven elements, identified through qualitative analysis of a non-exhaustive set of examples, genuinely cover every substantive aspect an AI capability evaluation can have, so no evaluation must be forced or distorted to fit.
Editorial extensions
If this is right
- Different evaluations—say, a reasoning benchmark and a safety red-teaming study—become comparable element by element, so hidden differences in prompts, evaluators, metrics, and analysis are made explicit.
- Methodological weaknesses surface as missing or underspecified elements, such as an evaluation that reports no statistical analysis, no prompt details, or no account of how evaluators were chosen.
- Policymakers and auditors gain a plain-language checklist for scrutinizing evaluation reports, which matters as regulations begin to embed benchmarks into legal requirements.
- The framework provides a natural organizing scheme for a living repository of evaluations, letting researchers track how open challenges such as data contamination are being addressed.
- Practitioners designing new evaluations can use the seven elements as a design checklist, ensuring each load-bearing decision is made and reported.
Reading between the lines
- The authors deliberately stop at description; a natural step they do not take is to turn the seven elements into a reporting standard—an 'evaluation card' that developers, third-party auditors, or regulators fill out, which would make the framework normative rather than descriptive.
- Because the framework is distilled from existing evaluation practice, it may carry the blind spots of that practice; a genuinely novel evaluation paradigm that does not fit the seven elements would expose a limit of the framework rather than a flaw in the paradigm.
- The framework's treatment of meta-evaluations—where the subject is the evaluation method itself—suggests it could serve as an audit tool: one could use it to check whether an evaluation's stated objective is actually supported by its task, metric, and analysis choices.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a descriptive conceptual framework for analyzing AI capability evaluations, organized around seven core elements (Evaluation Target, Task, Evaluated Subject, System Inputs, Evaluation Instance, Measurement, Result Analysis) plus transversal elements (practitioners, ethics, pilot tests, ablation studies). The framework is built from existing terminology and illustrated on three published evaluations in Appendix 2. The authors claim that the framework supports transparency, comparability, and interpretability across diverse evaluations without imposing new taxonomies or rigid formats, while explicitly excluding impact evaluations.
Significance. If the claimed coverage holds, this framework is a useful synthesis: it integrates widely used evaluation concepts, connects them to documented challenges such as data contamination, prompt sensitivity, and statistical analysis, and makes evaluation structure accessible to non-specialists including policymakers. The paper's strengths are its careful reuse of existing terminology, its explicit integration of evaluation challenges into the element structure, and the worked examples in Appendix 2. However, the central promise—that the seven elements capture the full range of aspects an evaluation method may encompass—is asserted rather than demonstrated. Because transparency and comparability across evaluations depend on this coverage, the framework's significance is conditional on a validation step that the paper itself defers to future work.
major comments (3)
- [Section 1] The claim that the framework is 'intended to be comprehensive—capturing the full range of aspects that an evaluation method may encompass' is load-bearing for the paper's stated benefits, but the only evidence offered is a qualitative reading of a non-exhaustive set of evaluations. Section 4 explicitly postpones systematic validation by proposing to 'analyze a range of evaluations using this framework as a set of use cases and to iteratively refine it.' As written, the comprehensiveness claim is a plausible hypothesis, not an established result; the paper should either weaken the claim to reflect the current evidence or provide a systematic methodology, such as a structured corpus analysis with coverage statistics and failure cases.
- [Appendix 2 and Section 4] The three use cases in Appendix 2 are self-admittedly 'not exhaustive' and 'simplified,' and they are all described by the authors in terms that fit the framework without tension. There is no discussion of borderline or adversarial cases, such as simulator-based environments, dynamic multi-agent protocols, or elicitation procedures that cut across System Inputs and Operational Context. Without such cases, the paper does not establish that the framework can represent diverse evaluations without forcing aspects into ill-fitting categories or silently dropping them. A concrete test would be to apply the framework to a balanced corpus of evaluations—including agent evaluations, red-teaming, meta-evaluations, and benchmark-based evaluations—and report inter-annotator agreement, coverage failures, and revisions made in response.
- [Section 3 (element boundaries)] The boundaries between several elements are not operationally defined. For example, prompt techniques are placed under System Inputs (Section 3.4), while access methods, configuration settings, and auxiliary tools are placed under Operational Context (Section 3.3), and multi-step task structure is placed under Task (Section 3.2). The paper does not provide decision rules for assigning a given evaluation feature to one element rather than another, and it does not discuss how to handle features that plausibly span multiple elements. Because comparability across evaluations requires that different analysts map the same evaluation onto the same structure, the absence of an annotation protocol or inter-rater reliability evidence is a gap in the central claim.
minor comments (6)
- [Appendix 2 table] The row for 'Evaluated Subject' in Paper 1 contains the typo 'descentralized' (should be 'decentralized'), and the term appears several times in the table.
- [Appendix 2 table, Measurement] The phrase 'Liket Scale' should be 'Likert Scale,' and 'the evaluator was also asked' should be 'the evaluators were also asked.'
- [Appendix A.7] The text contains 'ANOV A' with a spurious space; this should read 'ANOVA.'
- [Appendix B] The table cites 'Vidgen et al., 2021' for the RoBERTa hate speech classifier, but this reference is missing from the reference list.
- [Section 3.6] The phrase 'diverity' appears in the Appendix 2 table under the 'Diversity' metric; this should be corrected to 'diversity.'
- [Global] The paper uses footnotes for broad definitions of capability evaluations, but some footnotes extend to multiple lines and could be moved to the main text or made more concise for readability.
Circularity Check
No circularity: the framework is a descriptive synthesis of external evaluation literature, with no fitted inputs, derived predictions, or load-bearing self-citations.
full rationale
The paper does not derive a quantitative result from fitted parameters, and it contains no equations whose outputs are their own inputs. Its central product is a descriptive conceptual framework whose seven elements and transversal components are assembled from externally published work and from the authors' qualitative reading of existing evaluations; Appendix 2 then applies that framework to three published papers as illustrative use cases. The claimed benefit of the framework—supporting transparency, comparability, and interpretability—is presented as a design goal and as an assertion about the usefulness of the framework, not as a prediction inferred from the same framework. The paper explicitly acknowledges that its survey of evaluations is non-exhaustive and that future validation through additional use cases is needed, which is a limitation in evidence for comprehensiveness rather than a circular derivation. No step in the paper reduces a stated result to its own definition, no parameter is fitted and then renamed as a prediction, and no load-bearing premise is justified solely by a self-citation. Accordingly, the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption The seven-element decomposition is comprehensive enough to capture the full range of evaluation aspects.
- domain assumption A descriptive, non-sequential framework is more useful for analysis than a prescriptive sequential one.
- domain assumption Existing terminology from cited works can be coherently integrated without loss of meaning.
- domain assumption The three Appendix 2 use cases are representative enough to illustrate applicability.
Cite this review
Pith. "Pith review of A Conceptual Framework for AI Capability Evaluations." pith.science (2026). https://pith.science/paper/7EUAVMOO
@misc{pith2026250618213,
author = {Pith},
title = {Pith review of: A Conceptual Framework for AI Capability Evaluations},
year = {2026},
howpublished = {\url{https://pith.science/paper/7EUAVMOO}},
note = {Machine review of arXiv:2506.18213}
}
read the original abstract
As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence about system capabilities and risks. Yet there remains a lack of clarity on how to perform these assessments both comprehensively and reliably. To address this gap, we propose a conceptual framework for analyzing AI capability evaluations, offering a structured, descriptive approach that systematizes the analysis of widely used methods and terminology without imposing new taxonomies or rigid formats. This framework supports transparency, comparability, and interpretability across diverse evaluations. It also enables researchers to identify methodological weaknesses, assists practitioners in designing evaluations, and provides policymakers with an accessible tool to scrutinize, compare, and navigate complex evaluation landscapes.
Figures
Reference graph
Works this paper leans on
-
[1]
Early insights from developing question-answer evaluations for frontier AI , 2024
AISI. Early insights from developing question-answer evaluations for frontier AI , 2024. Accessed May 10, 2025
2024
-
[2]
Benchmarking foundation models with language-model-as-an-examiner
Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al. Benchmarking foundation models with language-model-as-an-examiner. Advances in Neural Information Processing Systems, 36: 0 78142--78167, 2023
2023
-
[3]
Declare and Justify: Explicit assumptions in AI evaluations are necessary for effective regulation
Barnett, P. and Thiergart, L. Declare and justify: Explicit assumptions in ai evaluations are necessary for effective regulation. arXiv preprint arXiv:2411.12820, 2024
work page Pith review arXiv 2024
-
[4]
A quantitative study of nlp approaches to question difficulty estimation
Benedetto, L. A quantitative study of nlp approaches to question difficulty estimation. In International Conference on Artificial Intelligence in Education, pp.\ 428--434. Springer, 2023
2023
-
[5]
Evaluating ai for law: Bridging the gap with open-source solutions
Bhambhoria, R., Dahan, S., Li, J., and Zhu, X. Evaluating ai for law: Bridging the gap with open-source solutions. arXiv preprint arXiv:2404.12349, 2024
arXiv 2024
-
[6]
F., Ammanamanchi, P
Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., et al. Lessons from the trenches on reproducible evaluation of language models. CoRR, 2024
2024
-
[7]
R., Steunebrink, B
Bieger, J., Th \'o risson, K. R., Steunebrink, B. R., and Thorarensen, T. Evaluation of general-purpose artificial intelligence: Why, what & how. In Proceedings of the IJCAI Workshop on Evaluating General-Purpose Artificial Intelligence (EGPAI 2016), 2016. URL https://alumni.media.mit.edu/ kris/ftp/EGPAI_2016_paper_9.pdf. Accessed May 10, 2025
2016
-
[8]
T., Li, Y., Lundberg, S., et al
Bubeck, S., Chadrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
2023
Show all 104 references
-
[9]
Evaluating ai evaluation: Perils and prospects
Burden, J. Evaluating ai evaluation: Perils and prospects. arXiv preprint arXiv:2407.09221, 2024
2024 arXiv
-
[10]
Paradigms of ai evaluation: Mapping goals, methodologies and culture
Burden, J., Te s i \'c , M., Pacchiardi, L., and Hern \'a ndez-Orallo, J. Paradigms of ai evaluation: Mapping goals, methodologies and culture. arXiv preprint arXiv:2502.15620, 2025
2025 arXiv
-
[11]
R., and Cheung, S.-C
Cao, J., Chan, Y.-K., Ling, Z., Wang, W., Li, S., Liu, M., Qiao, R., Han, Y., Wang, C., Yu, B., He, P., Wang, S., Zheng, Z., Lyu, M. R., and Cheung, S.-C. How should we build a benchmark? revisiting 274 code-related benchmarks for llms, 2025. URL https://arxiv.org/abs/2501.10711
2025 arXiv
-
[12]
L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al. Black-box access is insufficient for rigorous ai audits. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, ...
2024
-
[13]
Cheng, Y., Georgopoulos, M., Cevher, V., and Chrysos, G. G. Leveraging the context through multi-round interactions for jailbreaking attacks. arXiv preprint arXiv:2402.09177, 2024
2024 arXiv
-
[14]
N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M., Gonzalez, J. E., et al. Chatbot arena: An open platform for evaluating llms by human preference. In Forty-first International Conference on Machine Learning, 2024
2024
-
[15]
On the limitations of reference-free evaluations of generated text
Deutsch, D., Dror, R., and Roth, D. On the limitations of reference-free evaluations of generated text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp.\ 10960--10977, 2022
2022
-
[16]
R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M
Didolkar, A., Goyal, A., Ke, N. R., Guo, S., Valko, M., Lillicrap, T., Jimenez Rezende, D., Bengio, Y., Mozer, M. C., and Arora, S. Metacognitive capabilities of llms: An exploration in mathematical problem solving. Advances in Neural Information Processing Systems, 37: 0 1978...
2024
-
[17]
Generalization or memorization: Data contamination and trustworthy evaluation for large language models
Dong, Y., Jiang, X., Liu, H., Jin, Z., Gu, B., Yang, M., and Li, G. Generalization or memorization: Data contamination and trustworthy evaluation for large language models. In Findings of the Association for Computational Linguistics ACL 2024, pp.\ 12039--12050, 2024
2024
-
[18]
W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H
Dow, A., Vaughan, J. W., Barocas, S., Atalla, C., Chouldechova, A., and Wallach, H. Dimensions of generative ai evaluation design. In Proceedings of the NeurIPS 2024 Workshop on Evaluating Evaluations: Examining Best Practices for Measuring Broader Impacts of Generative AI, 20...
2024
-
[19]
Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation
Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation. arXiv preprint arXiv:2502.06559, 2025
2025 arXiv
-
[20]
Second draft of the general purpose AI code of practice, April 2024
European Commission . Second draft of the general purpose AI code of practice, April 2024. Written by independent experts. Accessed May 10, 2025
2024
-
[21]
Issue brief: Early best practices for frontier AI safety evaluations, 2024
Frontier Model Forum . Issue brief: Early best practices for frontier AI safety evaluations, 2024. Accessed May 10, 2025
2024
-
[22]
Llm-based nlg evaluation: Current status and challenges
Gao, M., Hu, X., Yin, X., Ruan, J., Pu, X., and Wan, X. Llm-based nlg evaluation: Current status and challenges. Computational Linguistics, pp.\ 1--28, 2025
2025
-
[23]
A case for better evaluation standards in nlg
Gehrmann, S., Clark, E., and Sellam, T. A case for better evaluation standards in nlg. In Workshop on Setting up ML Evaluation Standards to Accelerate Progress at ICLR 2022, 2022. URL https://iclr.cc/virtual/2022/7328. Poster presentation
2022
-
[24]
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Gehrmann, S., Clark, E., and Sellam, T. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. Journal of Artificial Intelligence Research, 77: 0 103--166, 2023
2023
-
[25]
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Guha, N., Nyarko, J., Ho, D., R \'e , C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models. Advances in Neural Information Processing S...
2023
-
[26]
R., Hullman, J., and Subramonyam, H
Gupta, N. R., Hullman, J., and Subramonyam, H. A conceptual framework for ethical evaluation of machine learning systems. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp.\ 534--546, 2024
2024
-
[27]
Deception abilities emerged in large language models
Hagendorff, T. Deception abilities emerged in large language models. Proceedings of the National Academy of Sciences, 121 0 (24): 0 e2317967121, 2024
2024
-
[28]
Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., and Schulz, E. Machine psychology, 2024. URL https://arxiv.org/abs/2303.13988
2024 arXiv
-
[29]
a m \"a l \
H \"a m \"a l \"a inen, M. and Alnajjar, K. Human evaluation of creative NLG systems: An interdisciplinary survey on recent papers. In Bosselut, A., Durmus, E., Gangal, V. P., Gehrmann, S., Jernite, Y., Perez-Beltrachini, L., Shaikh, S., and Xu, W. (eds.), Proceedings of the 1...
2021 doi
-
[30]
and Sharadin, N
Harding, J. and Sharadin, N. What is it for a machine learning model to have a capability? The British Journal for the Philosophy of Science, 2024. Advance online publication. Available at https://doi.org/10.1086/732153
2024 doi
-
[31]
Hofst \"a tter, F., Teoh, J., van der Weij, T., and Ward, F. R. The elicitation game: Stress-testing capability elicitation techniques. In Workshop on Socially Responsible Language Modelling Research, 2024. URL https://openreview.net/forum?id=zy6LB5t62f
2024
-
[32]
R., Srivastava, A., and Agrawal, P
Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J. R., Srivastava, A., and Agrawal, P. Curiosity-driven red-teaming for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?...
2024
-
[33]
and Zhou, X.-H
Hu, T. and Zhou, X.-H. Unveiling llm evaluation focused on metrics: Challenges and solutions. arXiv preprint arXiv:2404.09135, 2024
2024 arXiv
-
[34]
On the limitations of fine-tuned judge models for llm evaluation
Huang, H., Qu, Y., Zhou, H., Liu, J., Yang, M., Xu, B., and Zhao, T. On the limitations of fine-tuned judge models for llm evaluation. arXiv preprint arXiv:2403.02839, 2024
2024 arXiv
-
[35]
M ath P rompter: Mathematical reasoning using large language models
Imani, S., Du, L., and Shrivastava, H. M ath P rompter: Mathematical reasoning using large language models. In Sitaram, S., Beigman Klebanov, B., and Williams, J. D. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 5: Indu...
2023 doi
-
[36]
Reference-free evaluation metrics for text generation: A survey
Ito, T., van Deemter, K., and Suzuki, J. Reference-free evaluation metrics for text generation: A survey. arXiv preprint arXiv:2501.12011, 2025
2025 arXiv
-
[37]
Ivanova, A. A. Running cognitive evaluations on large language models: The do's and the don'ts. arXiv preprint arXiv:2312.01276, 2023
2023 arXiv
-
[38]
Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Jacovi, A., Caciularu, A., Goldman, O., and Goldberg, Y. Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp.\ 507...
2023
-
[39]
Cladder: assessing causal reasoning in language models
Jin, Z., Chen, Y., Leeb, F., Gresele, L., Kamal, O., Lyu, Z., Blin, K., Gonzalez, F., Kleiman-Weiner, M., Sachan, M., and Sch\" o lkopf, B. Cladder: assessing causal reasoning in language models. In Proceedings of the 37th International Conference on Neural Information Process...
2023
-
[40]
T., and Sch \"o lkopf, B
Jin, Z., Liu, J., LYU, Z., Poff, S., Sachan, M., Mihalcea, R., Diab, M. T., and Sch \"o lkopf, B. Can large language models infer causation from correlation? In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=vqIH0ObdqL
2024
-
[41]
R., Rockt \"a schel, T., and Perez, E
Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., Grefenstette, E., Bowman, S. R., Rockt \"a schel, T., and Perez, E. Debating with more persuasive llms leads to more truthful answers. In Proceedings of the 41st International Conference on Machine L...
2024
-
[42]
Causal reasoning and large language models: Opening a new frontier for causality
Kiciman, E., Ness, R., Sharma, A., and Tan, C. Causal reasoning and large language models: Opening a new frontier for causality. Transactions on Machine Learning Research, 2023
2023
-
[43]
Ai agent governance: A field guide
Kraprayoon, J., Williams, Z., and Fayyaz, R. Ai agent governance: A field guide. arXiv preprint arXiv:2505.21808, 2025
2025 arXiv
-
[44]
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S. Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=VD-AYtP0dve
2023
-
[45]
P., Wu, H., and Yu, H
Lalor, J. P., Wu, H., and Yu, H. Building an evaluation scale using item response theory. In Proceedings of the conference on empirical methods in natural language processing. Conference on empirical methods in natural language processing, volume 2016, pp.\ 648, 2016
2016
-
[46]
Laskar, M. T. R., Alqahtani, S., Bari, M. S., Rahman, M., Khan, M. A. M., Khan, H., Jahan, I., Bhuiyan, A., Tan, C. W., Parvez, M. R., et al. A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations. In Proceedin...
2024
-
[47]
J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M
Lee, S., Kim, M., Cherif, L., Dobre, D., Lee, J., Hwang, S. J., Kawaguchi, K., Gidel, G., Bengio, Y., Malkin, N., and Jain, M. Learning diverse attacks on large language models for robust red-teaming and safety tuning. In Red Teaming GenAI: What Can We Learn from Adversaries?,...
2025
-
[48]
Leveraging large language models for nlg evaluation: Advances and challenges
Li, Z., Xu, X., Shen, T., Xu, C., Gu, J.-C., Lai, Y., Tao, C., and Ma, S. Leveraging large language models for nlg evaluation: Advances and challenges. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp.\ 16028--16045, 2024
2024
-
[49]
D., Re, C., Acosta-Navas, D., Hudson, D
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F....
2023
-
[50]
Liao, Q. V. and Xiao, Z. Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100, 2023
2023 arXiv
-
[51]
D., and Schmidt, L
Liao, T., Taori, R., Raji, I. D., and Schmidt, L. Are we learning yet? a meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/fo...
2021
-
[52]
Against the achilles' heel: A survey on red teaming for generative models
Lin, L., Mu, H., Zhai, Z., Wang, M., Wang, Y., Wang, R., Gao, J., Zhang, Y., Che, W., Baldwin, T., et al. Against the achilles' heel: A survey on red teaming for generative models. Journal of Artificial Intelligence Research, 82: 0 687--775, 2025
2025
-
[53]
Datasets for large language models: A comprehensive survey
Liu, Y., Cao, J., Liu, C., Ding, K., and Jin, L. Datasets for large language models: A comprehensive survey. arXiv preprint arXiv:2402.18041, 2024
2024 arXiv
-
[54]
R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M
McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M. N. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. arXiv preprint arXiv:2402.09880, 2024
2024 arXiv
-
[55]
W., and Meisen, T
Meyes, R., Lu, M., de Puiseau, C. W., and Meisen, T. Ablation studies in artificial neural networks. arXiv preprint arXiv:1901.08644, 2019
1901 arXiv
-
[56]
Adding error bars to evals: A statistical approach to language model evaluations
Miller, E. Adding error bars to evals: A statistical approach to language model evaluations. arXiv preprint arXiv:2411.00640, 2024
2024 arXiv
-
[57]
Auditing large language models: a three-layered approach
M \"o kander, J., Schuett, J., Kirk, H., and Floridi, L. Auditing large language models: a three-layered approach. AI and Ethics, 4 0 (4), 2023
2023
-
[58]
Evaluating the performance of large language models via debates
Moniri, B., Hassani, H., and Dobriban, E. Evaluating the performance of large language models via debates. In Chiruzzo, L., Ritter, A., and Wang, L. (eds.), Findings of the Association for Computational Linguistics: NAACL 2025, pp.\ 2040--2075, Albuquerque, New Mexico, April 2...
2025
-
[59]
and Kapoor, S
Narayanan, A. and Kapoor, S. Gpt-4 and professional benchmarks: The wrong answer to the right question. AI Snake Oil (Substack), March 2023. URL https://www.aisnakeoil.com/p/gpt-4-and-professional-benchmarks. Accessed May 10, 2025
2023
-
[60]
Oecd framework for the classification of ai systems
OECD. Oecd framework for the classification of ai systems. OECD Digital Economy Papers, No. 323, 2022
2022
-
[61]
and Kang, E
Orr, W. and Kang, E. B. Ai as a sport: On the competitive epistemologies of benchmarking. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp.\ 1875--1884, 2024
2024
-
[62]
Llm evaluators recognize and favor their own generations
Panickssery, A., Bowman, S., and Feng, S. Llm evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems, 37: 0 68772--68802, 2024
2024
-
[63]
T., and Soder, L
Paskov, P., Berglund, L., Smith, E. T., and Soder, L. Gpai evaluations standards taskforce: towards effective ai governance. In Workshop on Socially Responsible Language Modelling Research, 2024
2024
-
[64]
Preliminary suggestions for rigorous gpai model evaluations
Paskov, P., Byun, M., Wei, K., and Webster, T. Preliminary suggestions for rigorous gpai model evaluations. Technical Report PEA3971-1, RAND Corporation, April 2025. URL https://www.rand.org/pubs/perspectives/PEA3971-1.html. Accessed May 10, 2025
2025
-
[65]
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukosiute, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pp.\ 133...
2023
-
[66]
and Jud, H
Pfister, R. and Jud, H. Understanding and benchmarking artificial intelligence: Openai's o3 is not agi. arXiv preprint arXiv:2501.07458, 2025
2025 arXiv
-
[67]
The roots search tool: Data transparency for llms
Piktus, A., Akiki, C., Villegas, P., Lauren c on, H., Dupont, G., Luccioni, S., Jernite, Y., and Rogers, A. The roots search tool: Data transparency for llms. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstra...
2023
-
[68]
D., Denton, E., Bender, E
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. AI and the everything in the whole wide world benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=j...
2021
-
[69]
Large language model evaluation via multi ai agents: Preliminary results
Rasheed, Z., Waseem, M., Syst \"a , K., and Abrahamsson, P. Large language model evaluation via multi ai agents: Preliminary results. In International Conference on Learning Representations, pp.\ 1--12, 2024
2024
-
[70]
A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al
Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al. Gaps in the safety evaluation of generative ai. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pp.\...
2024
-
[71]
Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices
Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. Betterbench: Assessing AI benchmarks, uncovering issues, and establishing best practices. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024 ...
2024
-
[72]
Reuel, A., Soder, L., Bucknall, B., and Undheim, T. A. Position: technical research and talent is needed for effective ai governance. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org, 2024 b
2024
-
[73]
S., Rajkumar, N., Moës, N., Ladish, J., Bau, D., Bricman, P., Guha, N., Newman, J., Bengio, Y., South, T., Pentland, A., Koyejo, S., Kochenderfer, M
Reuel, A., Bucknall, B., Casper, S., Fist, T., Soder, L., Aarne, O., Hammond, L., Ibrahim, L., Chan, A., Wills, P., Anderljung, M., Garfinkel, B., Heim, L., Trask, A., Mukobi, G., Schaeffer, R., Baker, M., Hooker, S., Solaiman, I., Luccioni, A. S., Rajkumar, N., Moës, N., Ladi...
2025 arXiv
-
[74]
Better than random: reliable nlg human evaluation with constrained active sampling
Ruan, J., Pu, X., Gao, M., Wan, X., and Zhu, Y. Better than random: reliable nlg human evaluation with constrained active sampling. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp.\ 18915--18923, 2024
2024
-
[75]
K., Saha, S., Jain, V., Mondal, S., and Chadha, A
Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. A systematic survey of prompt engineering in large language models: Techniques and applications. CoRR, abs/2402.07927, 2024. URL https://doi.org/10.48550/arXiv.2402.07927
-
[76]
L., and Agirre, E
Sainz, O., Campos, J., Garc \' a-Ferrero, I., Etxaniz, J., de Lacalle, O. L., and Agirre, E. Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp.\ 10776--10787, 2023
2023
-
[77]
Targeting the benchmark: On methodology in current natural language processing research
Schlangen, D. Targeting the benchmark: On methodology in current natural language processing research. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume ...
2021
-
[78]
S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G
Schulhoff, S., Ilie, M., Balepur, N., Kahadze, K., Liu, A., Si, C., Li, Y., Gupta, A., Han, H., Schulhoff, S., Dulepet, P. S., Vidyadhara, S., Ki, D., Agrawal, S., Pham, C., Kroiz, G. C., Li, F., Tao, H., Srivastava, A., Costa, H. D., Gupta, S., Rogers, M. L., Goncearenco, I.,...
-
[79]
Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Sclar, M., Choi, Y., Tsvetkov, Y., and Suhr, A. Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations, 2024. URL https://op...
2024
-
[80]
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023
2023 arXiv
-
[81]
CHOPS : CH at with customer profile systems for customer service with LLM s
Shi, J., Li, J., Ma, Q., Yang, Z., Ma, H., and Li, L. CHOPS : CH at with customer profile systems for customer service with LLM s. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=9Wmdk94oKF
2024
-
[82]
Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms
Sirdeshmukh, V., Deshpande, K., Mols, J., Jin, L., Cardona, E.-Y., Lee, D., Kritz, J., Primack, W., Yue, S., and Xing, C. Multichallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier llms. arXiv preprint arXiv:2501.17399, 2025
2025 arXiv
-
[83]
K., Grundy, E
Slattery, P., Saeri, A. K., Grundy, E. A., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., and Thompson, N. The ai risk repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. CoRR, 2024
2024
-
[84]
A study of translation edit rate with targeted human annotation
Snover, M., Dorr, B., Schwartz, R., Micciulla, L., and Makhoul, J. A study of translation edit rate with targeted human annotation. In Proceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers, pp.\ 223--231, 2006
2006
-
[85]
Audit cards: Contextualizing ai evaluations
Staufer, L., Yang, M., Reuel, A., and Casper, S. Audit cards: Contextualizing ai evaluations. arXiv preprint arXiv:2504.13839, 2025
2025 arXiv
-
[86]
Comprehensive reassessment of large-scale evaluation outcomes in llms: A multifaceted statistical approach
Sun, K., Wang, R., Liu, H., and Søgaard, A. Comprehensive reassessment of large-scale evaluation outcomes in llms: A multifaceted statistical approach. CoRR, abs/2403.15250, 2024. URL https://doi.org/10.48550/arXiv.2403.15250
-
[87]
Measuring data science automation: A survey of evaluation tools for ai assistants and agents
Testini, I., Hern \'a ndez-Orallo, J., and Pacchiardi, L. Measuring data science automation: A survey of evaluation tools for ai assistants and agents. arXiv preprint arXiv:2506.08800, 2025
2025
-
[88]
Thurnherr, B. C. Who should develop which ai evaluations?, April 2024. Accessed May 10, 2025
2024
-
[89]
Best practices for the human evaluation of automatically generated text
Van Der Lee, C., Gatt, A., Van Miltenburg, E., Wubben, S., and Krahmer, E. Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation, pp.\ 355--368, 2019
2019
-
[90]
M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S
Vania, C., Htut, P. M., Huang, W., Mungra, D., Yuanzhe Pang, R., Phang, J., Liu, H., Cho, K., and Bowman, S. R. Comparing test sets with item response theory. In Annual Meeting of the Association for Computational Linguistics, 2021
2021
-
[91]
Mint: Evaluating llms in multi-turn interaction with tools and language feedback
Wang, X., Wang, Z., Liu, J., Chen, Y., Yuan, L., Peng, H., and Ji, H. Mint: Evaluating llms in multi-turn interaction with tools and language feedback. In 12th International Conference on Learning Representations, ICLR 2024, 2024
2024
-
[92]
A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al
Weidinger, L., Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., et al. Sociotechnical safety evaluation of generative ai systems. arXiv preprint arXiv:2310.11986, 2023
-
[93]
D., Wallach, H., Mitchell, M., Wang, A., Salaudeen, O., Bommasani, R., Ganguli, D., Koyejo, S., and Isaac, W
Weidinger, L., Raji, I. D., Wallach, H., Mitchell, M., Wang, A., Salaudeen, O., Bommasani, R., Ganguli, D., Koyejo, S., and Isaac, W. Toward an evaluation science for generative ai systems. arXiv preprint arXiv:2503.05336, 2025
2025 arXiv
-
[94]
An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping
Xia, B., Lu, Q., Zhu, L., and Xing, Z. An ai system evaluation framework for advancing ai safety: Terminology, taxonomy, lifecycle mapping. In Proceedings of the 1st ACM International Conference on AI-Powered Software, pp.\ 74--78, 2024
2024
-
[95]
A critical review of causal inference benchmarks for large language models
Yang, L., Clivio, O., Shirvaikar, V., and Falck, F. A critical review of causal inference benchmarks for large language models. In AAAI 2024 Workshop on ''Are Large Language Models Simply Causal Parrots?'', 2023. URL https://openreview.net/forum?id=mRwgczYZFJ
2024
-
[96]
Evaluatology: The science and engineering of evaluation
Zhan, J., Wang, L., Gao, W., Li, H., Wang, C., Huang, Y., Li, Y., Yang, Z., Kang, G., Luo, C., Ye, H., Dai, S., and Zhang, Z. Evaluatology: The science and engineering of evaluation. BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 4 0 (1): 0 100162, 2024. I...
2024
-
[97]
K., Klyman, K., Mai, Y., Levine, Y., Zhang, Y., Bommasani, R., and Liang, P
Zhang, A. K., Klyman, K., Mai, Y., Levine, Y., Zhang, Y., Bommasani, R., and Liang, P. Language model developers should report train-test overlap. arXiv preprint arXiv:2410.08385, 2024 a
2024 arXiv
-
[98]
Q., Shaw, R., Anthis, J
Zhang, A. Q., Shaw, R., Anthis, J. R., Milton, A., Tseng, E., Suh, J., Ahmad, L., Kumar, R. S. S., Posada, J., Shestakofsky, B., et al. The human factor in ai red teaming: Perspectives from social and collaborative computing. In Companion Publication of the 2024 Conference on ...
2024
-
[99]
Pacost: Paired confidence significance testing for benchmark contamination detection in large language models
Zhang, H., Lin, Y., and Wan, X. Pacost: Paired confidence significance testing for benchmark contamination detection in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pp.\ 1794--1809, 2024 c
2024
-
[100]
and Kanayet, F
Zhang, Y. and Kanayet, F. Genai evaluation maturity framework (gemf) to assess and improve genai evaluations. In Workshop on Evaluating Evaluations: Examining Best Practices for Measuring Broader Impacts of Generative AI at NeurIPS 2024 , 2024. URL https://neurips.cc/virtual/2...
2024
-
[101]
Llmeval: A preliminary study on how to evaluate large language models
Zhang, Y., Zhang, M., Yuan, H., Liu, S., Shi, Y., Gui, T., Zhang, Q., and Huang, X. Llmeval: A preliminary study on how to evaluate large language models. Proceedings of the AAAI Conference on Artificial Intelligence, 38 0 (17): 0 19615--19622, Mar. 2024 d . doi:10.1609/aaai.v...
2024 doi
-
[102]
L., Trischler, A., Daum \'e III, H., Suleman, K., and Olteanu, A
Zhou, K., Blodgett, S. L., Trischler, A., Daum \'e III, H., Suleman, K., and Olteanu, A. Deconstructing nlg evaluation: Evaluation practices, assumptions, and their implications. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computa...
2022
-
[103]
X., Chen, X., Lin, Y., Wen, J.-R., and Han, J
Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J. Don't make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023
2023 arXiv
-
[104]
Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks
Zhou, X., Weyssow, M., Widyasari, R., Zhang, T., He, J., Lyu, Y., Chang, J., Zhang, B., Huang, D., and Lo, D. Lessleak-bench: A first investigation of data leakage in llms across 83 software engineering benchmarks. arXiv preprint arXiv:2502.06215, 2025
2025 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.