Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Deprecating Benchmarks: Criteria and Framework

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper argues that outdated or flawed AI benchmarks should be actively deprecated, and provides seven criteria plus a three-phase assessment-reporting-notification framework to do it.

desk verdict A solid, honest position paper that adapts dataset deprecation to benchmarks with useful criteria and sample reports; the lack of operational thresholds is a real limitation but not a fatal one. read the letter →

arxiv 2507.06434 v1 pith:AMQKACNM submitted 2025-07-08 cs.CY cs.AIcs.LG

classification cs.CYcs.AIcs.LG
keywords benchmarkdeprecationAIevaluationsaturationdatacontaminationgovernancesafety-washinglifecyclepartial
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that AI benchmarks, once outdated or flawed, must be actively deprecated rather than left in use out of inertia. It proposes seven criteria—saturation, contamination, statistical bias, annotation errors, task obsolescence, invalidated assumptions, and semantic drift—for deciding when a benchmark should be fully retired or partially upgraded. It then lays out a three-phase process: assess whether and how much deprecation is needed, write a deprecation report, and notify users, with version control and appeals built in. The goal is to stop benchmark scores from inflating model capabilities, wasting evaluation effort, and enabling safety-washing.

What carries the argument

The central mechanism is a deprecation cycle built from seven heuristic criteria—saturation, contamination, statistical bias, high annotation error rate, task obsolescence, invalidated assumptions, and semantic drift—and a three-phase workflow: assessment, reporting, and notification. The criteria are deliberately non-binary, meant to function like case law rather than thresholds; the workflow produces a deprecation report that states the decision, the affected version, the timeline, alternatives, and how past results should be read. Version control (via DOI) and an appeals process are part of the machinery, because without them a deprecation notice cannot be tracked or contested.

What would settle it

Run a controlled field trial in which a governance body publishes a deprecation list for ten widely used, demonstrably contaminated benchmarks, then track citation and leaderboard usage of those benchmarks for a year; if usage does not decline relative to ten matched non-deprecated benchmarks, the notification and enforcement phases do not change practice.

Watch

Extended reading notes

Core claim

The paper claims that benchmark deprecation should be an explicit, documented, and governable act rather than an informal consequence of decay. It treats benchmarks as versioned artifacts that can be fully deprecated or partially upgraded, and argues that the reason they are rarely retired is not technical but institutional: developers lose reputational capital, labs game leaderboards, and no regulator is mandated to reassess validity. Against that background, the paper's proposal is that seven soft criteria plus a three-phase process of assessment, reporting, and notification give developers and governance actors a workable way to retire benchmarks and to say, in public, why.

Load-bearing premise

The framework depends on the seven deprecation criteria working as soft heuristics without concrete thresholds or validated measurement procedures; if that premise fails, deprecation decisions become arbitrary and inconsistent across governance bodies.

Editorial extensions

If this is right

  • Benchmark developers and governance bodies can retire a benchmark without waiting for consensus, because the criteria give a shared vocabulary for why a benchmark has stopped measuring capabilities.
  • Partial deprecation means a flawed benchmark can be upgraded by removing contaminated or erroneous subsets instead of being discarded wholesale, preserving the parts that still work.
  • Published deprecation lists and notices would make it harder for labs to claim state-of-the-art results on benchmarks known to be saturated or leaked.
  • Regulators that already require benchmarks, such as those implementing the EU AI Act, can attach deprecation lists to their evaluation duties and set compliance timelines for switching.
  • Version control and deprecation reports make historical leaderboard scores interpretable, because past results are explicitly tied to the benchmark version that produced them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If deprecation becomes routine, benchmark authors may begin building private holdouts or dynamic items from the start, since a static public benchmark now carries retirement risk from release day.
  • The paper stops short of assigning weights to the seven criteria; an editorial inference is that different governance contexts will need different priority orders, e.g. contamination may matter more for safety-critical capability tests than for general knowledge tests.
  • A deprecation list could itself become a political instrument if inclusion is opaque; the paper's transparency and appeals provisions are therefore not optional extras but the load-bearing guardrails of the proposal.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that outdated or flawed benchmarks should be actively deprecated to prevent distorted capability assessments, wasted evaluation resources, and safety-washing. It proposes a non-exhaustive list of seven deprecation criteria (saturation, contamination, statistical bias, high annotation error rate, task obsolescence, invalidated assumptions, and semantic drift) and a three-phase framework consisting of assessment, reporting, and notification, with an explicit concept of partial deprecation (upgrading). It also provides recommendations for benchmark developers and governance actors, an EU-specific adaptation involving the AI Office, and sample deprecation reports for SWE-bench. The authors state that concrete thresholds are beyond the scope of the paper and that the criteria are meant as evolving, non-rigid heuristics.

Significance. If operationalized, the framework would address a real gap: there is currently no standard process for retiring or revising benchmarks, and the literature documents many stale or flawed benchmarks still in use. The paper is well grounded in prior work, especially Luccioni et al. (2022) and Staufer et al. (2025), gives concrete illustrative examples, and includes detailed sample deprecation reports that make the workflow tangible. It also honestly acknowledges its own scope limits: the Impact Statement says deprecation may not be sufficient on its own, and footnote 2 defers all concrete thresholds. The criteria set and the partial-deprecation concept are the main substantive contributions; the main shortcoming is that the criteria are not yet operational enough to guarantee reproducible or consistent decisions across governance actors.

major comments (3)
  1. [Section 3, footnote 2; Criteria 1-7] The decision procedure is underdetermined. Section 3 states that the criteria are 'soft' guidance and footnote 2 says that 'determining concrete thresholds for these criteria is beyond the scope of this paper.' This is not merely a presentation issue: with no cutoff for saturation, no standard for what counts as a 'high' annotation error rate, and no detection protocol for contamination, two governance bodies applying the same criteria to the same benchmark can reach opposite conclusions. The 'case law' analogy in Section 3 does not resolve this, because Section 4.1's appeals process is internal to a single actor and there is no cross-actor precedent system. Since the Introduction's central claim is that deprecation prevents distorted capability assessments, the framework must at least specify how decisions are to be made consistently. The authors should add operational definitions, measurement protocols, or a validation plan (e.g., pilot applications with inter-rater reliability statistics), and discuss how conflicting deprecation lists from different actors would be reconciled.
  2. [Section 4.1, partial deprecation] The suggested corrective measures for partial deprecation are not operational. 'Resampling to rebalance class distributions' and 'removing leaked data' are only possible when the flaw is in the benchmark's sample selection; if contamination has already occurred through model pretraining, removing leaked items from the benchmark does not repair the evaluation. 'Incorporating random variables' is vague and could change the construct the benchmark is meant to measure. Section 4.2 asks for reports to 'delineate which benchmark components remain valid,' but the paper does not provide a component taxonomy or a versioning schema for partial deprecation, despite Recommendation 1 calling for DOIs. Without such a schema, partial deprecation cannot be implemented consistently across benchmark developers and governance actors.
  3. [Appendix B; Abstract] The paper's empirical support is limited to illustrative and partially fictional case studies. The abstract claims the framework will 'prevent distorted capability assessments, wasted resources, and safety-washing,' but no retrospective application or simulation is provided to show that the criteria would have identified well-known flawed benchmarks (e.g., MMLU, GSM8K, or SWE-bench), or that different raters would agree on the assessments. Adding a retrospective analysis of two or three documented cases, including an assessment of inter-rater agreement, would materially strengthen the central claim and give governance actors concrete guidance on how to apply the criteria.
minor comments (4)
  1. [Section 2, paragraph 3] The claim that benchmarks 'have transformed into mainly test-time artifacts' is too strong for benchmarks that are still used in fine-tuning or instruction tuning; suggest qualifying this to 'for many frontier-model evaluations'.
  2. [Section 3, criterion 8 (task obsolescence)] There is a typo: 'the BIG benchmark(Srivastava et al., 2022)' is missing a space before the parenthetical citation.
  3. [Section 5] The EU example assumes that the AI Office can mandate model-card and technical-report updates after notification, but the cited legal provisions (Articles 15.2 and 51.3) concern managing evaluations and supplementing benchmarks, not directly imposing model-card obligations; the paper should clarify the legal basis for this enforcement assumption.
  4. [Recommendation 7] The recommendation to 'prevent models from being accessed because they are tested on deprecated benchmarks' could have unintended consequences for research and for historical comparisons; the paper should discuss explicit exemptions beyond the brief mention of 'authorized use' in Section 4.2.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper proposes heuristic deprecation criteria and a governance framework without deriving predictions from fitted inputs.

full rationale

This is a position/framework paper with no derivation chain to walk. It proposes seven soft deprecation criteria justified by a literature review, plus a three-phase assessment, reporting, and notification framework. The criteria are explicitly heuristic and non-exhaustive; the paper states that determining concrete thresholds is beyond scope, so there is no fitted parameter later renamed as a prediction and no quantity defined in terms of another and then claimed as an output. The only self-citation (Staufer et al., 2025, by co-author Staufer) supports the motivation that evaluations should state when they become deprecated; it is not load-bearing for the criteria, which instead rest on external work such as Luccioni et al. (2022), Reuel et al. (2024), and Eriksson et al. (2025). No benchmark predictions are made, no uniqueness theorem is invoked, and no existing result is merely renamed. Concerns about the lack of operational thresholds are correctness and reproducibility risks, not circularity. The central claim is therefore self-contained with respect to its cited inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities. The framework relies on domain assumptions about governance authority and the sufficiency of heuristic criteria, both acknowledged limitations.

assumptions (3)
  • domain assumption Benchmarks function primarily as test-time artifacts for frontier models, making deprecation cheaper and faster than for training datasets.
    Used in Section 2 to argue that benchmark deprecation is easier than dataset deprecation and enables rapid governance response.
  • domain assumption Governance actors (e.g., the EU AI Office) have the authority, capacity, and legitimacy to create and enforce deprecation lists.
    Section 5 and Recommendation 4 assume that such bodies can compel model developers to update evaluations; the paper does not demonstrate this authority.
  • ad hoc to paper Deprecation criteria can be applied as flexible heuristics without fixed thresholds and still yield consistent, fair decisions.
    Section 3 states thresholds are out of scope but the framework depends on treating the criteria as 'soft' guidance similar to case law.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deprecating Benchmarks: Criteria and Framework." pith.science (2026). https://pith.science/paper/AMQKACNM

@misc{pith2026250706434,
  author       = {Pith},
  title        = {Pith review of: Deprecating Benchmarks: Criteria and Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AMQKACNM}},
  note         = {Machine review of arXiv:2507.06434}
}
read the original abstract

As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domains. However, there is a lack of guidance on when and how benchmarks should be deprecated once they cease to effectively perform their purpose. This risks benchmark scores over-valuing model capabilities, or worse, obscuring capabilities and safety-washing. Based on a review of benchmarking practices, we propose criteria to decide when to fully or partially deprecate benchmarks, and a framework for deprecating benchmarks. Our work aims to advance the state of benchmarking towards rigorous and quality evaluations, especially for frontier models, and our recommendations are aimed to benefit benchmark developers, benchmark users, AI governance actors (across governments, academia, and industry panels), and policy makers.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unsteady Metrics and Benchmarking Cultures of AI Model Builders

    cs.AI 2026-05 accept novelty 8.0 of 10

    AI model builders mostly highlight unique benchmarks that act as flexible narrative tools for market positioning rather than standardized scientific measurements.

Reference graph

Works this paper leans on

48 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    and Church, K

    Alonso, O. and Church, K. Evaluating the Evaluations: A Perspective on Benchmarks . In ACM SIGIR Forum, volume 58, pp.\ 1--27. ACM New York, NY, USA, 2025

  3. [3]

    Frontier AI regulation: Managing emerging risks to public safety

    Anderljung, M., Barnhart, J., Korinek, A., Leung, J., O'Keefe, C., Whittlestone, J., Avin, S., Brundage, M., Bullock, J., Cass-Beggs, D., et al. Frontier AI regulation: Managing emerging risks to public safety. arXiv preprint arXiv:2307.03718, 2023

  4. [4]

    S., Jenner, E., Casper, S., Sourbut, O., et al

    Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024

  5. [5]

    Mapping global dynamics of benchmark creation and saturation in artificial intelligence

    Barbosa-Silva, A., Ott, S., Blagec, K., Brauner, J., and Samwald, M. Mapping global dynamics of benchmark creation and saturation in artificial intelligence. arXiv preprint arXiv:2203.04592, 2022

  6. [6]

    Guidelines for Retracting Articles

    Barbour, V., Kleinert, S., Wager, E., and Yentis, S. Guidelines for Retracting Articles . Technical report, Committee on Publication Ethics, September 2009

  7. [7]

    Benchmark Early and Red Team Often: A Framework for Assessing and Managing Dual-Use Hazards of AI Foundation Models

    Barrett, A. M., Jackson, K., Murphy, E. R., Madkour, N., and Newman, J. Benchmark early and red team often: A framework for assessing and managing dual-use hazards of AI foundation models. arXiv preprint arXiv:2405.10986, 2024

  8. [8]

    Reliable benchmarking: requirements and solutions

    Beyer, D., L \"o we, S., and Wendler, P. Reliable benchmarking: requirements and solutions. International Journal on Software Tools for Technology Transfer, 21 0 (1): 0 1--29, 2019

Show all 48 references
  1. [9]

    J., Kolesnikov, A., Zhai, X., and Oord, A

    Beyer, L., H \'e naff, O. J., Kolesnikov, A., Zhai, X., and Oord, A. v. d. Are we done with ImageNet ? arXiv preprint arXiv:2006.07159, 2020

  2. [10]

    Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals

    Blagec, K., Kraiger, J., Fr \"u hwirt, W., and Samwald, M. Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals. Journal of Biomedical Informatics, 137: 0 104274, 2023

  3. [11]

    State-of-the-art: The temporal order of benchmarking culture

    Campolo, A. State-of-the-art: The temporal order of benchmarking culture. 4 0 (2): 0 35, 2025. ISSN 2731-4650, 2731-4669. doi:10.1007/s44206-025-00190-x. URL https://link.springer.com/10.1007/s44206-025-00190-x

  4. [12]

    Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021

  5. [13]

    S., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C

    Chowdhury, N., Aung, J., Chan, J. S., Jaffe, O., Sherburn, D., Starace, G., Mays, E., Dias, R., Aljubeh, M., Glaese, M., Jimenez, C. E., Yang, J., Ho, L., Patwardhan, T., Liu, K., and Madry, A. Introducing SWE-bench Verified . https://openai.com/index/introducing-swe-bench-ver...

  6. [14]

    Training verifiers to solve math word problems

    Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021

  7. [15]

    Benchmarks for automated commonsense reasoning: A survey

    Davis, E. Benchmarks for automated commonsense reasoning: A survey. ACM Computing Surveys, 56 0 (4): 0 1--41, 2023

  8. [16]

    A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O

    Dehghani, M., Tay, Y., Gritsenko, A. A., Zhao, Z., Houlsby, N., Diaz, F., Metzler, D., and Vinyals, O. The benchmark lottery. arXiv preprint arXiv:2107.07002, 2021

  9. [17]

    Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

    Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation . arXiv preprint arXiv:2502.06559, 2025

  10. [18]

    European Parliament, C. o. t. E. U. Regulation ( EU ) 2024/1689 of the European Parliament and of the Council of 13 June 2024 . https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689, 2024. [Accessed 26-09-2024]

  11. [19]

    W., Wallach, H., Iii, H

    Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Iii, H. D., and Crawford, K. Datasheets for Datasets , December 2021

  12. [20]

    P., Leang, J

    Gema, A. P., Leang, J. O. J., Hong, G., Devoto, A., Mancino, A. C. M., Saxena, R., He, X., Zhao, Y., Du, X., Madani, M. R. G., et al. Are We Done with MMLU? arXiv preprint arXiv:2406.04127, 2024

  13. [21]

    AILuminate: Introducing v1

    Ghosh, S., Frase, H., Williams, A., Luger, S., R \"o ttger, P., Barez, F., McGregor, S., Fricklas, K., Kumar, M., Bollacker, K., et al. AILuminate: Introducing v1. 0 of the AI Risk and Reliability Benchmark from MLCommons . arXiv preprint arXiv:2503.05731, 2025

  14. [22]

    Measuring massive multitask language understanding

    Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  15. [23]

    Llmtest\_needleinahaystack, 2024

    Kamradt, G. Llmtest\_needleinahaystack, 2024. URL https://github.com/gkamradt/LLMTest_NeedleInAHaystack. Accessed: 2025-05-12

  16. [24]

    Diachronic word embeddings and semantic shifts: a survey

    Kutuzov, A., vrelid, L., Szymanski, T., and Velldal, E. Diachronic word embeddings and semantic shifts: a survey. arXiv preprint arXiv:1806.03537, 2018

  17. [25]

    Multi needle in a haystack, March 2024

    LangChain . Multi needle in a haystack, March 2024. URL https://blog.langchain.dev/multi-needle-in-a-haystack/. Accessed: 2025-05-12

  18. [26]

    D., and Schmidt, L

    Liao, T., Taori, R., Raji, I. D., and Schmidt, L. Are we learning yet? A meta review of evaluation failures across machine learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021

  19. [27]

    Artificial Intelligence Law of the People’s Republic of China (Draft for Suggestions from Scholars)

    Linghan, Z., Jianjun, Y., Ying, C., Jingwu, Z., Xuzhi, H., Zhifeng, Z., and Xiaoben, X. Artificial Intelligence Law of the People’s Republic of China (Draft for Suggestions from Scholars) . Center for Security and Emerging Technology. Accessed: Aug, 16, 2024

  20. [28]

    S., Corry, F., Sridharan, H., Ananny, M., Schultz, J., and Crawford, K

    Luccioni, A. S., Corry, F., Sridharan, H., Ananny, M., Schultz, J., and Crawford, K. A framework for deprecating datasets: Standardizing documentation, identification, and communication . In Proceedings of the 2022 acm conference on fairness, accountability, and transparency, ...

  21. [29]

    C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., Lotufo, J

    Maslej, N., Fattorini, L., Perrault, R., Gil, Y., Parli, V., Kariuki, N., Capstick, E., Reuel, A., Brynjolfsson, E., Etchemendy, J., Ligett, K., Lyons, T., Manyika, J., Niebles, J. C., Shoham, Y., Wald, R., Walsh, T., Hamrah, A., Santarlasci, L., Lotufo, J. B., Rome, A., Shi, ...

  22. [30]

    R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M

    McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Watters, P., and Halgamuge, M. N. Inadequacies of large language model benchmarks in the era of generative artificial intelligence. arXiv preprint arXiv:2402.09880, 2024

  23. [31]

    Benchmarking the benchmarks

    Miltenberger, M., Arzt, S., Holzinger, P., and N \"a umann, J. Benchmarking the benchmarks . In Proceedings of the 2023 ACM Asia Conference on Computer and Communications Security, pp.\ 387--400, 2023

  24. [32]

    Digital object identifier (doi ) system

    Paskin, N. Digital object identifier (doi ) system. Encyclopedia of library and information sciences, 3: 0 1586--1592, 2010

  25. [33]

    M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M

    Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M. tinyBenchmarks: evaluating LLMs with fewer examples . arXiv preprint arXiv:2402.14992, 2024

  26. [34]

    M., Hanna, A., and Paullada, A

    Raji, D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. AI and the Everything in the Whole Wide World Benchmark . In Vanschoren, J. and Yeung, S. (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , volume 1, 2021 a

  27. [35]

    D., Bender, E

    Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A. Ai and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366, 2021 b

  28. [36]

    A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L

    Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., Griffin, C., Bariach, B., Gabriel, I., Rieser, V., Isaac, W., and Weidinger, L. Gaps in the Safety Evaluation of Generative AI . Proceeding...

  29. [37]

    How much are LLM s contaminated? A comprehensive survey and the llmsanitize library

    Ravaut, M., Ding, B., Jiao, F., Chen, H., Li, X., Zhao, R., Qin, C., Xiong, C., and Joty, S. How much are LLM s contaminated? A comprehensive survey and the llmsanitize library. arXiv preprint arXiv:2404.00699, 2024

  30. [38]

    Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? Advances in Neural Information Processing Systems, 37: 0 68559--68594, 2024

    Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R., et al. Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? Advances in Neural Information Processing Systems, 37: 0 68559--68594, 2024

  31. [39]

    Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. J. BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices , 2024. URL https://arxiv.org/abs/2411.12990

  32. [40]

    Meta got caught gaming AI benchmarks

    Robison, K. Meta got caught gaming AI benchmarks. 2025. URL https://www.theverge.com/meta/645012/meta-llama-4-maverick-benchmarks-gaming

  33. [41]

    The Leaderboard Illusion

    Singh, S., Nan, Y., Wang, A., D'Souza, D., Kapoor, S., \"U st \"u n, A., Koyejo, S., Deng, Y., Longpre, S., Smith, N., et al. The Leaderboard Illusion . arXiv preprint arXiv:2504.20879, 2025

  34. [42]

    Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022

  35. [43]

    Audit Cards: Contextualizing AI Evaluations

    Staufer, L., Yang, M., Reuel, A., and Casper, S. Audit Cards: Contextualizing AI Evaluations . arXiv preprint arXiv:2504.13839, April 2025

  36. [44]

    S., and Whittaker, M

    Varoquaux, G., Luccioni, A. S., and Whittaker, M. Hype, Sustainability, and the Price of the Bigger-is-Better Paradigm in AI . arXiv preprint arXiv:2409.14160, 2024

  37. [45]

    S., Deng, Y., Dunn, S., and Zhang, L

    Xia, C. S., Deng, Y., Dunn, S., and Zhang, L. Agentless: Demystifying LLM-based Software Engineering Agents , 2024. URL https://arxiv.org/abs/2407.01489

  38. [46]

    Benchmark Data Contamination of Large Language Models: A Survey

    Xu, C., Guan, S., Greene, D., Kechadi, M., et al. Benchmark Data Contamination of Large Language Models: A Survey . arXiv preprint arXiv:2406.04244, 2024 a

  39. [47]

    Benchmarking benchmark leakage in large language models

    Xu, R., Wang, Z., Fan, R.-Z., and Liu, P. Benchmarking benchmark leakage in large language models. arXiv preprint arXiv:2404.18824, 2024 b

  40. [48]

    X., Chen, X., Lin, Y., Wen, J.-R., and Han, J

    Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J. Don't make your LLM an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.