REVIEW 3 major objections 4 minor 1 cited by
Deprecating Benchmarks: Criteria and Framework
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that outdated or flawed AI benchmarks should be actively deprecated, and provides seven criteria plus a three-phase assessment-reporting-notification framework to do it.
desk verdict A solid, honest position paper that adapts dataset deprecation to benchmarks with useful criteria and sample reports; the lack of operational thresholds is a real limitation but not a fatal one. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a deprecation cycle built from seven heuristic criteria—saturation, contamination, statistical bias, high annotation error rate, task obsolescence, invalidated assumptions, and semantic drift—and a three-phase workflow: assessment, reporting, and notification. The criteria are deliberately non-binary, meant to function like case law rather than thresholds; the workflow produces a deprecation report that states the decision, the affected version, the timeline, alternatives, and how past results should be read. Version control (via DOI) and an appeals process are part of the machinery, because without them a deprecation notice cannot be tracked or contested.
What would settle it
Run a controlled field trial in which a governance body publishes a deprecation list for ten widely used, demonstrably contaminated benchmarks, then track citation and leaderboard usage of those benchmarks for a year; if usage does not decline relative to ten matched non-deprecated benchmarks, the notification and enforcement phases do not change practice.
Extended reading notes
Core claim
The paper claims that benchmark deprecation should be an explicit, documented, and governable act rather than an informal consequence of decay. It treats benchmarks as versioned artifacts that can be fully deprecated or partially upgraded, and argues that the reason they are rarely retired is not technical but institutional: developers lose reputational capital, labs game leaderboards, and no regulator is mandated to reassess validity. Against that background, the paper's proposal is that seven soft criteria plus a three-phase process of assessment, reporting, and notification give developers and governance actors a workable way to retire benchmarks and to say, in public, why.
Load-bearing premise
The framework depends on the seven deprecation criteria working as soft heuristics without concrete thresholds or validated measurement procedures; if that premise fails, deprecation decisions become arbitrary and inconsistent across governance bodies.
Editorial extensions
If this is right
- Benchmark developers and governance bodies can retire a benchmark without waiting for consensus, because the criteria give a shared vocabulary for why a benchmark has stopped measuring capabilities.
- Partial deprecation means a flawed benchmark can be upgraded by removing contaminated or erroneous subsets instead of being discarded wholesale, preserving the parts that still work.
- Published deprecation lists and notices would make it harder for labs to claim state-of-the-art results on benchmarks known to be saturated or leaked.
- Regulators that already require benchmarks, such as those implementing the EU AI Act, can attach deprecation lists to their evaluation duties and set compliance timelines for switching.
- Version control and deprecation reports make historical leaderboard scores interpretable, because past results are explicitly tied to the benchmark version that produced them.
Reading between the lines
- If deprecation becomes routine, benchmark authors may begin building private holdouts or dynamic items from the start, since a static public benchmark now carries retirement risk from release day.
- The paper stops short of assigning weights to the seven criteria; an editorial inference is that different governance contexts will need different priority orders, e.g. contamination may matter more for safety-critical capability tests than for general knowledge tests.
- A deprecation list could itself become a political instrument if inclusion is opaque; the paper's transparency and appeals provisions are therefore not optional extras but the load-bearing guardrails of the proposal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that outdated or flawed benchmarks should be actively deprecated to prevent distorted capability assessments, wasted evaluation resources, and safety-washing. It proposes a non-exhaustive list of seven deprecation criteria (saturation, contamination, statistical bias, high annotation error rate, task obsolescence, invalidated assumptions, and semantic drift) and a three-phase framework consisting of assessment, reporting, and notification, with an explicit concept of partial deprecation (upgrading). It also provides recommendations for benchmark developers and governance actors, an EU-specific adaptation involving the AI Office, and sample deprecation reports for SWE-bench. The authors state that concrete thresholds are beyond the scope of the paper and that the criteria are meant as evolving, non-rigid heuristics.
Significance. If operationalized, the framework would address a real gap: there is currently no standard process for retiring or revising benchmarks, and the literature documents many stale or flawed benchmarks still in use. The paper is well grounded in prior work, especially Luccioni et al. (2022) and Staufer et al. (2025), gives concrete illustrative examples, and includes detailed sample deprecation reports that make the workflow tangible. It also honestly acknowledges its own scope limits: the Impact Statement says deprecation may not be sufficient on its own, and footnote 2 defers all concrete thresholds. The criteria set and the partial-deprecation concept are the main substantive contributions; the main shortcoming is that the criteria are not yet operational enough to guarantee reproducible or consistent decisions across governance actors.
major comments (3)
- [Section 3, footnote 2; Criteria 1-7] The decision procedure is underdetermined. Section 3 states that the criteria are 'soft' guidance and footnote 2 says that 'determining concrete thresholds for these criteria is beyond the scope of this paper.' This is not merely a presentation issue: with no cutoff for saturation, no standard for what counts as a 'high' annotation error rate, and no detection protocol for contamination, two governance bodies applying the same criteria to the same benchmark can reach opposite conclusions. The 'case law' analogy in Section 3 does not resolve this, because Section 4.1's appeals process is internal to a single actor and there is no cross-actor precedent system. Since the Introduction's central claim is that deprecation prevents distorted capability assessments, the framework must at least specify how decisions are to be made consistently. The authors should add operational definitions, measurement protocols, or a validation plan (e.g., pilot applications with inter-rater reliability statistics), and discuss how conflicting deprecation lists from different actors would be reconciled.
- [Section 4.1, partial deprecation] The suggested corrective measures for partial deprecation are not operational. 'Resampling to rebalance class distributions' and 'removing leaked data' are only possible when the flaw is in the benchmark's sample selection; if contamination has already occurred through model pretraining, removing leaked items from the benchmark does not repair the evaluation. 'Incorporating random variables' is vague and could change the construct the benchmark is meant to measure. Section 4.2 asks for reports to 'delineate which benchmark components remain valid,' but the paper does not provide a component taxonomy or a versioning schema for partial deprecation, despite Recommendation 1 calling for DOIs. Without such a schema, partial deprecation cannot be implemented consistently across benchmark developers and governance actors.
- [Appendix B; Abstract] The paper's empirical support is limited to illustrative and partially fictional case studies. The abstract claims the framework will 'prevent distorted capability assessments, wasted resources, and safety-washing,' but no retrospective application or simulation is provided to show that the criteria would have identified well-known flawed benchmarks (e.g., MMLU, GSM8K, or SWE-bench), or that different raters would agree on the assessments. Adding a retrospective analysis of two or three documented cases, including an assessment of inter-rater agreement, would materially strengthen the central claim and give governance actors concrete guidance on how to apply the criteria.
minor comments (4)
- [Section 2, paragraph 3] The claim that benchmarks 'have transformed into mainly test-time artifacts' is too strong for benchmarks that are still used in fine-tuning or instruction tuning; suggest qualifying this to 'for many frontier-model evaluations'.
- [Section 3, criterion 8 (task obsolescence)] There is a typo: 'the BIG benchmark(Srivastava et al., 2022)' is missing a space before the parenthetical citation.
- [Section 5] The EU example assumes that the AI Office can mandate model-card and technical-report updates after notification, but the cited legal provisions (Articles 15.2 and 51.3) concern managing evaluations and supplementing benchmarks, not directly imposing model-card obligations; the paper should clarify the legal basis for this enforcement assumption.
- [Recommendation 7] The recommendation to 'prevent models from being accessed because they are tested on deprecated benchmarks' could have unintended consequences for research and for historical comparisons; the paper should discuss explicit exemptions beyond the brief mention of 'authorized use' in Section 4.2.
Circularity Check
No significant circularity: the paper proposes heuristic deprecation criteria and a governance framework without deriving predictions from fitted inputs.
full rationale
This is a position/framework paper with no derivation chain to walk. It proposes seven soft deprecation criteria justified by a literature review, plus a three-phase assessment, reporting, and notification framework. The criteria are explicitly heuristic and non-exhaustive; the paper states that determining concrete thresholds is beyond scope, so there is no fitted parameter later renamed as a prediction and no quantity defined in terms of another and then claimed as an output. The only self-citation (Staufer et al., 2025, by co-author Staufer) supports the motivation that evaluations should state when they become deprecated; it is not load-bearing for the criteria, which instead rest on external work such as Luccioni et al. (2022), Reuel et al. (2024), and Eriksson et al. (2025). No benchmark predictions are made, no uniqueness theorem is invoked, and no existing result is merely renamed. Concerns about the lack of operational thresholds are correctness and reproducibility risks, not circularity. The central claim is therefore self-contained with respect to its cited inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption Benchmarks function primarily as test-time artifacts for frontier models, making deprecation cheaper and faster than for training datasets.
- domain assumption Governance actors (e.g., the EU AI Office) have the authority, capacity, and legitimacy to create and enforce deprecation lists.
- ad hoc to paper Deprecation criteria can be applied as flexible heuristics without fixed thresholds and still yield consistent, fair decisions.
Cite this review
Pith. "Pith review of Deprecating Benchmarks: Criteria and Framework." pith.science (2026). https://pith.science/paper/AMQKACNM
@misc{pith2026250706434,
author = {Pith},
title = {Pith review of: Deprecating Benchmarks: Criteria and Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/AMQKACNM}},
note = {Machine review of arXiv:2507.06434}
}
read the original abstract
As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.
Forward citations
Cited by 1 Pith paper
-
Unsteady Metrics and Benchmarking Cultures of AI Model Builders
AI model builders mostly highlight unique benchmarks that act as flexible narrative tools for market positioning rather than standardized scientific measurements.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Alonso, O. and Church, K. Evaluating the Evaluations: A Perspective on Benchmarks . In ACM SIGIR Forum, volume 58, pp.\ 1--27. ACM New York, NY, USA, 2025
work page 2025
-
[3]
Frontier AI regulation: Managing emerging risks to public safety
Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., et al. Frontier AI regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2307.03718, 2023
arXiv 2023
-
[4]
S., Jenner, E., Casper, S., Sourbut, O., et al
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024
arXiv 2024
-
[5]
Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Barbosa-Silva, A., Ott, S., Blagec, K., Brauner, J., and Samwald, M. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. arXiv preprint arXiv:2203.04592, 2022
arXiv 2022
-
[6]
Guidelines for Retracting Articles
Barbour, V., Kleinert, S., Wager, E., and Yentis, S. Guidelines for Retracting Articles . Technical report, Committee on Publication Ethics, September 2009
work page 2009
-
[7]
Barrett, A. M., Jackson, K., Murphy, E. R., Madkour, N., and Newman, J. Benchmark early and red team often: A framework for assessing and managing dual-use hazards of AI foundation models. arXiv preprint arXiv:2405.10986, 2024
work page Pith review arXiv 2024
-
[8]
Reliable benchmarking: requirements and solutions
Beyer, D., L \"o we, S., and Wendler, P. Reliable benchmarking: requirements and solutions. International Journal on Software Tools for Technology Transfer, 21 0 (1): 0 1--29, 2019
work page 2019
Show all 48 references
-
[9]
J., Kolesnikov, A., Zhai, X., and Oord, A
Beyer, L., H \'e naff, O. J., Kolesnikov, A., Zhai, X., and Oord, A. v. d. Are we done with ImageNet ? arXiv preprint arXiv:2006.07159, 2020
2006 arXiv
-
[10]
Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals
Blagec, K., Kraiger, J., Fr \"u hwirt, W., and Samwald, M. Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals. Journal of Biomedical Informatics, 137: 0 104274, 2023
2023
-
[11]
State-of-the-art: The temporal order of benchmarking culture
Campolo, A. State-of-the-art: The temporal order of benchmarking culture. 4 0 (2): 0 35, 2025. ISSN 2731-4650, 2731-4669. doi:10.1007/s44206-025-00190-x. URL https://link.springer.com/10.1007/s44206-025-00190-x
2025 doi
-
[12]
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021
2021 arXiv
-
[13]
S., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C
Chowdhury, N., Aung, J., Chan, J. S., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C. E., Yang, J., Ho, L., Patwardhan, T., Liu, K., and Madry, A. Introducing SWE-bench Verified . https://openai.com/index/introducing-swe-bench-ver...
2024
-
[14]
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021
2021 arXiv
-
[15]
Benchmarks for automated commonsense reasoning: A survey
Davis, E. Benchmarks for automated commonsense reasoning: A survey. ACM Computing Surveys, 56 0 (4): 0 1--41, 2023
2023
-
[16]
A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O
Dehghani, M., Tay, Y., Gritsenko, A. A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O. The benchmark lottery. arXiv preprint arXiv:2107.07002, 2021
2021 arXiv
-
[17]
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation . arXiv preprint arXiv:2502.06559, 2025
2025 arXiv
-
[18]
European Parliament, C. o. t. E. U. Regulation ( EU ) 2024/1689 of the European Parliament and of the Council of 13 June 2024 . https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689, 2024. [Accessed 26-09-2024]
2024
-
[19]
W., Wallach, H., Iii, H
Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K. Datasheets for Datasets , December 2021
2021
-
[20]
P., Leang, J
Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., He, X., Zhao, Y., Du, X., Madani, M. R. G., et al. Are We Done with MMLU? arXiv preprint arXiv:2406.04127, 2024
2024 arXiv
-
[21]
AILuminate: Introducing v1
Ghosh, S., Frase, H., Williams, A., Luger, S., R \"o ttger, P., Barez, F., McGregor, S., Fricklas, K., Kumar, M., Bollacker, K., et al. AILuminate: Introducing v1. 0 of the AI Risk and Reliability Benchmark from MLCommons . arXiv preprint arXiv:2503.05731, 2025
2025 arXiv
-
[22]
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020
2009 arXiv
-
[23]
Llmtest\_needleinahaystack, 2024
Kamradt, G. Llmtest\_needleinahaystack, 2024. URL https://github.com/gkamradt/LLMTest_NeedleInAHaystack. Accessed: 2025-05-12
2024
-
[24]
Diachronic word embeddings and semantic shifts: a survey
Kutuzov, A., vrelid, L., Szymanski, T., and Velldal, E. Diachronic word embeddings and semantic shifts: a survey. arXiv preprint arXiv:1806.03537, 2018
2018 arXiv
-
[25]
Multi needle in a haystack, March 2024
LangChain . Multi needle in a haystack, March 2024. URL https://blog.langchain.dev/multi-needle-in-a-haystack/. Accessed: 2025-05-12
2024
-
[26]
D., and Schmidt, L
Liao, T., Taori, R., Raji, I. D., and Schmidt, L. Are we learning yet? A meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021
2021
-
[27]
Artificial Intelligence Law of the People’s Republic of China (Draft for Suggestions from Scholars)
Linghan, Z., Jianjun, Y., Ying, C., Jingwu, Z., Xuzhi, H., Zhifeng, Z., and Xiaoben, X. Artificial Intelligence Law of the People’s Republic of China (Draft for Suggestions from Scholars) . Center for Security and Emerging Technology. Accessed: Aug, 16, 2024
2024
-
[28]
S., Corry, F., Sridharan, H., Ananny, M., Schultz, J., and Crawford, K
Luccioni, A. S., Corry, F., Sridharan, H., Ananny, M., Schultz, J., and Crawford, K. A framework for deprecating datasets: Standardizing documentation, identification, and communication . In Proceedings of the 2022 acm conference on fairness, accountability, and transparency, ...
2022
-
[29]
C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., Lotufo, J
Maslej, N., Fattorini, L., Perrault, R., Gil, Y., Parli, V., Kariuki, N., Capstick, E., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., Lotufo, J. B., Rome, A., Shi, ...
2025
-
[30]
R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M
McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M. N. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. arXiv preprint arXiv:2402.09880, 2024
2024 arXiv
-
[31]
Benchmarking the benchmarks
Miltenberger, M., Arzt, S., Holzinger, P., and N \"a umann, J. Benchmarking the benchmarks . In Proceedings of the 2023 ACM Asia Conference on Computer and Communications Security, pp.\ 387--400, 2023
2023
-
[32]
Digital object identifier (doi ) system
Paskin, N. Digital object identifier (doi ) system. Encyclopedia of library and information sciences, 3: 0 1586--1592, 2010
2010
-
[33]
M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M. tinyBenchmarks: evaluating LLMs with fewer examples . arXiv preprint arXiv:2402.14992, 2024
2024 arXiv
-
[34]
M., Hanna, A., and Paullada, A
Raji, D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. AI and the Everything in the Whole Wide World Benchmark . In Vanschoren, J. and Yeung, S. (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , volume 1, 2021 a
2021
-
[35]
D., Bender, E
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A. Ai and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366, 2021 b
2021 arXiv
-
[36]
A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L
Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L. Gaps in the Safety Evaluation of Generative AI . Proceeding...
2024 doi
-
[37]
How much are LLM s contaminated? A comprehensive survey and the llmsanitize library
Ravaut, M., Ding, B., Jiao, F., Chen, H., Li, X., Zhao, R., Qin, C., Xiong, C., and Joty, S. How much are LLM s contaminated? A comprehensive survey and the llmsanitize library. arXiv preprint arXiv:2404.00699, 2024
2024 arXiv
-
[38]
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? Advances in Neural Information Processing Systems, 37: 0 68559--68594, 2024
Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R., et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? Advances in Neural Information Processing Systems, 37: 0 68559--68594, 2024
2024
-
[39]
Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. J. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices , 2024. URL https://arxiv.org/abs/2411.12990
2024 arXiv
-
[40]
Meta got caught gaming AI benchmarks
Robison, K. Meta got caught gaming AI benchmarks. 2025. URL https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming
2025
-
[41]
The Leaderboard Illusion
Singh, S., Nan, Y., Wang, A., D'Souza, D., Kapoor, S., \"U st \"u n, A., Koyejo, S., Deng, Y., Longpre, S., Smith, N., et al. The Leaderboard Illusion . arXiv preprint arXiv:2504.20879, 2025
2025 arXiv
-
[42]
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022
2022 arXiv
-
[43]
Audit Cards: Contextualizing AI Evaluations
Staufer, L., Yang, M., Reuel, A., and Casper, S. Audit Cards: Contextualizing AI Evaluations . arXiv preprint arXiv:2504.13839, April 2025
2025 arXiv
-
[44]
S., and Whittaker, M
Varoquaux, G., Luccioni, A. S., and Whittaker, M. Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI . arXiv preprint arXiv:2409.14160, 2024
2024 arXiv
-
[45]
S., Deng, Y., Dunn, S., and Zhang, L
Xia, C. S., Deng, Y., Dunn, S., and Zhang, L. Agentless: Demystifying LLM-based Software Engineering Agents , 2024. URL https://arxiv.org/abs/2407.01489
2024 arXiv
-
[46]
Benchmark Data Contamination of Large Language Models: A Survey
Xu, C., Guan, S., Greene, D., Kechadi, M., et al. Benchmark Data Contamination of Large Language Models: A Survey . arXiv preprint arXiv:2406.04244, 2024 a
2024 arXiv
-
[47]
Benchmarking benchmark leakage in large language models
Xu, R., Wang, Z., Fan, R.-Z., and Liu, P. Benchmarking benchmark leakage in large language models. arXiv preprint arXiv:2404.18824, 2024 b
2024 arXiv
-
[48]
X., Chen, X., Lin, Y., Wen, J.-R., and Han, J
Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J. Don't make your LLM an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023
2023 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.