Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that current LLM benchmarks, from GLUE and MMLU to human and LLM judges, can be gamed, so top scores do not prove genuine language understanding.

desk verdict A readable but rough survey of known benchmark vulnerabilities; no new evidence or systematic method, so the advertised 'systematic analysis' overclaims. read the letter →

arxiv 2412.03597 v1 pith:WE255WAS submitted 2024-12-02 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords largelanguagemodelsbenchmarkhackingdatacontaminationLLM-as-judgehumanevaluationleaderboardreliabilityGoodhart'slawbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper seeks to establish that leaderboard-style evaluation of large language models is systematically unreliable: models can post high scores on GLUE, MMLU, and similar benchmarks while lacking genuine language understanding or adaptability. It argues that benchmark overfitting, public test-set contamination, selective reporting, human-judge inconsistency, and LLM-as-judge biases all inflate apparent progress and can be exploited deliberately. If the paper is right, widely quoted benchmark numbers should be read as weak evidence of real capability, and evaluation must move toward dynamic, contamination-resistant, domain-specific protocols. The paper is a narrative review rather than a new experiment; its conclusion rests on assembling existing studies that document each vulnerability.

What carries the argument

The load-bearing object is the benchmark score treated as an estimator of true performance, written as $B = P + \varepsilon$, together with the generalization gap $G(M) = E[S(M,B_{\text{test}})] - E[S(M,B_{\text{real}})]$. The paper uses these formalizations, plus the contamination rate and the mutual-information measure of leakage, to make explicit the ways a score can rise while capability does not; the named principle underneath is Goodhart's law, that a measure ceases to be good once it becomes a target. The formalisms do the argumentative work of showing that benchmark optimization, contamination, and judge bias are not hypothetical but quantifiable distortions.

What would settle it

Take the current top-ranked models on GLUE, MMLU, and a popular LLM-judge arena, run a contamination audit to remove any leaked examples, then test the same models on a freshly written, never-public set of analogous tasks in the same domains; if the rankings and scores stay nearly the same and the generalization gap $G(M)$ is near zero for most models, the paper's claim of pervasive benchmark vulnerability would be contradicted for those benchmarks.

Watch

Extended reading notes

Core claim

The paper's central claim is that the evaluation ecosystem for LLMs contains a structural mismatch: the measures being optimized—static benchmark scores and judge ratings—are not the capabilities the field actually wants, namely robust understanding and adaptation to novel tasks. It identifies three families of exploitation: benchmark hacking (overfitting and task-specific optimization), data contamination (training/test overlap that inflates scores), and evaluator bias (humans and LLM judges whose judgments are noisy, format-sensitive, self-preferring, or otherwise gameable). The authors formalize these distortions with simple equations: benchmark score $B$ decomposes as $B = P + \varepsilon$; contamination rate is $CR = |D_{\text{train}} \cap D_{\text{test}}| / |D_{\text{test}}|$; mutual information $I(\theta; D_{\text{test}} | D_{\text{train}})$ signals leakage; and a generalization gap $G(M) = E[S(M,B_{\text{test}})] - E[S(M,B_{\text{real}})]$ quantifies the distance between benchmark success and real-world performance. They conclude that no single existing evaluation method is trustworthy on its own, that near-perfect benchmark scores deserve skepticism, and that future frameworks should be zero-day, zero-shot, domain-specific, and governed so they can be iterated before models overfit to them.

Load-bearing premise

The argument that these vulnerabilities are pervasive depends on the unstated assumption that the studies the paper selects are a representative sample of LLM evaluation practice and that no significant counter-evidence was left out; the paper gives no systematic search strategy, inclusion criteria, or quantitative synthesis to support that.

Editorial extensions

If this is right

  • Leaderboard rankings should be interpreted as upper-bound marketing claims rather than measurements of understanding, because reported scores can be inflated by contamination and overfitting.
  • Near-perfect or saturated scores on static benchmarks are weak evidence of progress; they may reflect exploitation of dataset artifacts rather than model competence.
  • Human evaluation cannot serve as an unbiased gold standard: annotators disagree on the same outputs and can be fooled by superficial fluency.
  • LLM-as-judge scores need debiasing: judges show self-preference, length bias, and sensitivity to prompt phrasing, so single-judge numbers are not stable.
  • Evaluation practice should shift toward dynamic, zero-day, domain-specific tasks with contamination audits and transparent reporting of methodology.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the authors leave implicit: new benchmarks should ship with a contamination audit and a preregistered refresh schedule before their numbers are treated as evidence of capability.
  • The same mechanisms would likely apply to code and mathematics benchmarks, where public solutions make contamination easier; the paper does not test this, but its logic predicts similar inflation.
  • The generalization-gap formalism suggests a cheap diagnostic the authors do not propose: report $G(M)$ alongside every published benchmark score so brittle leaders are visible.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a narrative review of vulnerabilities in large language model (LLM) evaluation, arguing that current benchmarks such as GLUE and MMLU are susceptible to gaming, data contamination, and evaluator bias, and that high leaderboard scores therefore create a false impression of genuine language understanding. After a brief history of NLP benchmarks, the paper surveys benchmark overfitting, public dataset leakage, test set contamination, task-specific optimization, adversarial benchmarking, human evaluation bias, and LLM-as-judge biases, with illustrative mathematical notation. It concludes that no single evaluation method is reliable and proposes future work on zero-day, zero-shot evaluation protocols and consortium-based dynamic benchmarking.

Significance. If the paper's central claim were established—that current benchmarks pervasively misrepresent LLM capabilities—it would have substantial implications for how the field measures progress and allocates trust in leaderboard results. The manuscript does collect a useful set of known critique themes, including contamination detection, annotation artifacts, self-preference in LLM judges, and human evaluator inconsistency. However, the paper provides no new experiments, no quantitative synthesis of existing findings, and no methodologically grounded framework for its 'systematic analysis' claim. Its value as a survey is further weakened by multiple misattributed citations and internal textual errors. The significance of the topic is high, but the contribution as written is a checklist of known issues rather than a substantiated systematic assessment.

major comments (4)
  1. [Abstract; Sections 2-3] The paper repeatedly claims a 'systematic analysis' and 'pervasive vulnerabilities across the evaluation spectrum', but no methodology is described anywhere in the manuscript. There is no search strategy, no inclusion or exclusion criteria, no corpus description, and no quantitative synthesis. The examples discussed (GPT-3 on LAMBADA, BERT on SQuAD, human evaluator and LLM-judge studies) are a curated selection, so the generalization to 'pervasive' vulnerabilities is unsupported. This is load-bearing because the central conclusion that benchmarks give a false perception of progress depends on the representativeness of the cited studies; without a defined method, the claim is an overgeneralization from an anecdotal sample.
  2. [Section 3.1, Eq. (1); Section 3.2; Section 3.3] The mathematical formulations (B = P + ε, the contamination rate CR, the mutual information I(θ; D_test | D_train), and the exposure metric) are purely definitional and do not constitute evidence that any specific benchmark is inaccurate. The paper never estimates these quantities on actual models or datasets, so the equations provide no support for the claim that benchmarks misrepresent true performance. If the intent is to formalize known failure modes, this should be stated explicitly; as written, the formalism may give the misleading impression of a quantitative proof.
  3. [Section 3.11; Abstract] The central claim that benchmarks create 'a false perception of progress' requires a benchmark-independent operationalization of 'true performance,' which the paper never provides. Without such a yardstick, the claim is unfalsifiable as stated. A concrete test would be to compare leaderboard rankings with held-out dynamic evaluations or with performance on well-defined real-world tasks, and to measure the rank correlation; the paper contains no such test. Similarly, the assertion in Section 3.11 that 'models approach or achieve perfect scores on established benchmarks' is an empirical claim presented without data or citation.
  4. [Sections 3.4 and 3.5; References [6], [20], [56]] Several citations do not support the claims to which they are attached, undermining the reliability of the survey. In Section 3.5, universal adversarial triggers are attributed to Wallace et al. with citation [6], but reference [6] is Raji and Buolamwini (2019), not Wallace et al. In Section 3.4, the finding that BERT models use shallow heuristics on MNLI is attributed to 'McCoy' with citation [56], but reference [56] is the MNLI dataset paper by Williams et al. (2018), not McCoy et al. In Section 3.1, the claims about GPT-3 on LAMBADA and BERT on SQuAD are supported by citations [25] and [20], which are the LAMBADA and SQuAD 2.0 dataset papers, not analyses of overfitting or pattern exploitation. These misattributions mean the cited evidence cannot be checked by the reader and weaken the manuscript's authority.
minor comments (5)
  1. [Section 3.6] This section contains a duplicated paragraph: the text beginning 'One key limitation of human evaluations is the inherent inconsistency...' and ending '...poorly constructed [45]' appears twice verbatim, interrupting the flow of the argument.
  2. [Section 3.7] The section ends mid-sentence with 'as well as the exploration of entire' and no continuation; the sentence and the section appear to be truncated.
  3. [References] The reference list has inconsistencies: reference [45] is used for two different papers (Chiang et al., 'Chatbot Arena', and Zheng et al., 'Judging LLM-as-a-Judge'), and references [18] and [41] both cite the same paper by Dubois et al. (2024). These duplicate entries need to be resolved.
  4. [Section 3.4] The text refers to 'McCoy demonstrated' without a corresponding reference entry; the actual citation [56] points to the MNLI dataset paper rather than to any work by McCoy, so the intended source should be identified and cited correctly.
  5. [Section 1.1] The terminology 'Human-as-Judge' and 'LLM-as-Judge' is used with inconsistent capitalization and hyphenation; the authors should adopt a single consistent form throughout.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a narrative survey whose claims rest on external citations and illustrative definitions, not on a derivation that reduces to its own inputs.

full rationale

This paper is a narrative survey rather than an empirical derivation, and I find no circular step that reduces a claimed result to its own inputs. The formal notation in Sections 3.1-3.4 (e.g., B = P + epsilon, CR = |Dtrain intersect Dtest| / |Dtest|, and I(theta; Dtest | Dtrain)) is definitional scaffolding used to restate cited findings; no parameter is fitted and no quantity is predicted from a prior that already contains it. The load-bearing evidence for 'pervasive vulnerabilities' comes from external studies (GPT-3 on LAMBADA, BERT on SQuAD, human-evaluator inconsistency in Clark et al., LLM-judge biases in Zheng et al.), not from the authors' own prior work, and no uniqueness theorem or ansatz is imported from self-citations. The central weakness -- generalizing from a curated selection of mostly older model/task observations to 'from basic metrics to ... GLUE and MMLU' without a systematic search or an independent operationalization of 'true performance' -- is a representativeness and falsifiability problem, not a circularity problem. Accordingly the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no new entities or fitted parameters. Its conclusions rest on the accuracy and representativeness of the cited literature, plus the applicability of Goodhart's law to benchmark-driven model development.

assumptions (2)
  • domain assumption Goodhart's law applies to LLM benchmarks: once a score becomes a target, it ceases to be a good measure.
    Invoked in Section 2 (citing [51]) as the rationale for why benchmark scores are unreliable. This is a plausible but unproved social-scientific assertion about model development incentives.
  • domain assumption The cited studies' findings are accurate and transferable to current LLMs.
    The paper's overview of vulnerabilities (overfitting, contamination, judge bias) relies entirely on prior work such as [39], [45], [46], and [47] without revalidation or replication. If those studies are flawed or outdated, the survey's conclusions lose support.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?." pith.science (2026). https://pith.science/paper/WE255WAS

@misc{pith2026241203597,
  author       = {Pith},
  title        = {Pith review of: The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WE255WAS}},
  note         = {Machine review of arXiv:2412.03597}
}
read the original abstract

The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine language understanding and adaptability. Our systematic analysis of NLP evaluation frameworks reveals pervasive vulnerabilities across the evaluation spectrum, from basic metrics to complex benchmarks like GLUE and MMLU. These vulnerabilities manifest through benchmark exploitation, dataset contamination, and evaluation bias, creating a false perception of progress in language understanding capabilities. Through extensive review of contemporary evaluation approaches, we identify significant limitations in static benchmark designs, human evaluation protocols, and LLM-as-judge frameworks, all of which compromise the reliability of current performance assessments. As LLM capabilities evolve and existing benchmarks become redundant, we lay the groundwork for new evaluation methods that resist manipulation, minimize data contamination, and assess domain-specific tasks. This requires frameworks that are adapted dynamically, addressing current limitations and providing a more accurate reflection of LLM performance.

Figures

Figures reproduced from arXiv: 2412.03597 by the authors.

Figure 1
Figure 1. Benchmarks released by year till Aug 2024 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. LLM-Leaderboard [13] [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Open LLM Leaderboard by Hugging Face due to ”benchmark hacking,” where benchmarks are manipulated to exaggerate model performance. 2 3.1 Benchmark Overfitting Benchmark overfitting, or ”overtuning,” occurs when models are excessively optimized for specific bench￾marks without genuinely improving their general ca￾pabilities. This is akin to ”p-hacking” in empirical studies, where data is manipulated to achieve mislea… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Data leakage distribution [40] dataset. In an ideal scenario, these sets should be dis￾joint: Dtrain ∩ Dtest = ∅ However, due to the public nature of many bench￾mark datasets, we often encounter a situation where: Dtrain ∩ Dtest ̸= ∅ This overlap can be quantified usin…
Figure 5
Figure 5. Figure 5: Detecting dataset contamination via log [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Excerpts from human evaluators’ explana [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

    cs.HC 2026-08 conditional novelty 6.0 of 10

    Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.

  2. Rethinking the Understanding Ability across LLMs through Mutual Information

    cs.CL 2025-05 conditional novelty 4.0 of 10

    The paper uses token-level recoverability as a computable lower bound on mutual information to compare LLMs and to fine-tune them, finding encoder-only models preserve information better than decoder-only models.

Reference graph

Works this paper leans on

65 extracted references · 47 canonical work pages · cited by 2 Pith papers

  1. [45]

    Gonzalez, et al

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, et al. 2024. Chat- bot Arena: An Open Platform for Evaluat- ing LLMs by Human Preference. arXiv preprint arXiv:2403.04132

  2. [6]

    Inioluwa Raji and Joy Buolamwini. 2019. Ac- tionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’19) , pages 429–435. ACM. DOI: 10.1145/3306618.3314244

  3. [56]

    GPT-4 Technical Report

  4. [25]

    Denis Paperno, Germ ´an Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern ´andez. 2016. The LAMBADA dataset: Word prediction requiring a broad dis- course context

  5. [20]

    Pranav Rajpurkar, Robin Jia, and Percy Liang

  6. [1]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert- V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...

  7. [2]

    Nicholas Carlini, Florian Tram `er, Eric Wallace, Matthew Jagielski, Ariel Herbert-V oss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, ´Ulfar Erlingsson, Alina Oprea, and Colin Raffel

  8. [3]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , pages 4171–4186. As...

Show all 65 references
  1. [4]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qim- ing Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott...

  2. [5]

    Alec Radford, Karthik Narasimhan, Tim Sali- mans, Ilya Sutskever, and others. 2018. Improv- ing Language Understanding by Generative Pre- training. OpenAI. Preprint, pages 1–12

  3. [7]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Atten- tion is All You Need. In Advances in Neural Infor- mation Processing Systems (NeurIPS 2017). Curran Associates, Inc

  4. [8]

    Eliezer Yudkowsky. 2008. Artificial Intelligence as a Positive and Negative Factor in Global Risk. In Nick Bostrom and Milan M. Cirkovic (eds.), Global Catastrophic Risks, pages 308–345. Oxford University Press

  5. [9]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017. Men Also Like Shopping: Reducing Gender Bias Amplifica- tion using Corpus-level Constraints. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , pages 2979–2...

  6. [10]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman

  7. [11]

    Alex Wang, Yada Pruksachatkun, Nikita Nan- gia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020. Super- GLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

  8. [12]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Ja- cob Steinhardt. 2021. Measuring Massive Multitask Language Understanding

  9. [13]

    Open LLM Leaderboard. n.d. Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard-old. Retrieved from https://huggingface.co/spaces/ open-llm-leaderboard-old/open_ llm_leaderboard

  10. [14]

    Manning, Christopher R ´e, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Ya- sunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Bin- hang Yuan, Bobby Yan, Ce Zhang, Christian Cos- grove, Christopher D. Manning, Christopher R ´e, Diana Acos...

  11. [15]

    March 2024

    Ziyou Yan. March 2024. Task-Specific LLM Evals that Do & Don’t Work. eugeneyan.com. Retrieved from https://eugeneyan.com/ writing/evals/

  12. [16]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?

  13. [17]

    Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...

  14. [18]

    Hashimoto

    Yann Dubois, Bal ´azs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length- Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

  15. [19]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Au- tomatic Evaluation of Machine Translation. InPro- ceedings of the 40th Annual Meeting on Associa- tion for Computational Linguistics, pages 311–318. Association for Computational Linguistics

  16. [21]

    Saiful Bari, and Haidar Khan

    Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Al- subaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, M. Saiful Bari, and Haidar Khan. 2024. When Bench- marks are Targets: Revealing the Sensitivity of ...

  17. [22]

    Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal Ad- versarial Triggers for Attacking and Analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Process- ing and the 9th International Joint Conferenc...

  18. [23]

    Bowman and George E

    Samuel R. Bowman and George E. Dahl. 2021. What Will it Take to Fix Benchmarking in Natural Language Understanding?

  19. [24]

    Ido Dagan. 2006. The PASCAL Recognising Textual Entailment Challenge. In Machine Learn- ing Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tex- tual Entailment, pages 177–190. Springer Berlin Heidelberg

  20. [26]

    Nicholas Carlini, Chang Liu, ´Ulfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The Secret Sharer: Evaluating and Testing Unintended Memo- rization in Neural Networks

  21. [27]

    Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Ad- versarial NLI: A New Benchmark for Natural Lan- guage Understanding

  22. [28]

    Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. 2020. Learning the Difference that Makes a Difference with Counterfactually- Augmented Data

  23. [29]

    Bowman, and Noah A

    Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation Artifacts in Natural Language Inference Data

  24. [30]

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Sum- marization Branches Out, pages 74–81. Associa- tion for Computational Linguistics

  25. [31]

    Satanjeev Banerjee and Alon Lavie. 2005. ME- TEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrin- sic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–7...

  26. [32]

    Jesse Dodge, Maarten Sap, Ana Marasovi ´c, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Docu- menting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

  27. [33]

    Raphael Olivier and Bhiksha Raj. 2023. How many perturbations break this model? Evaluating robustness beyond adversarial accuracy

  28. [34]

    Yogesh Balaji, Tom Goldstein, and Judy Hoff- man. 2019. Instance Adaptive Adversarial Train- ing: Improved Accuracy Tradeoffs in Neural Nets

  29. [35]

    Sophie Xhonneux, Alessandro Sordoni, Stephan G¨unnemann, Gauthier Gidel, and Leo Schwinn

  30. [36]

    Moustafa Alzantot, Yash Sharma, Ahmed Elgo- hary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating Natural Language Adver- sarial Examples

  31. [37]

    Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Alek- sander Madry. 2019. Adversarial Examples Are Not Bugs, They Are Features

  32. [38]

    Aviral Kale, Truc Nguyen, Jack Harris, Chang Li, Jiayi Zhang, and Xiaoli Ma. 2023. Provenance Documentation to Enable Explainable and Trust- worthy AI: A Literature Review. Data Intelligence, 5(1):139–162

  33. [39]

    Hashimoto

    Yanai Oren, Nicolas Meister, Niladri Chatterji, Firoj Ladhak, and Tatsunori B. Hashimoto. 2023. Proving Test Set Contamination in Black Box Language Models. arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2310.17623

  34. [40]

    Silvio Balloccu, Petra Schmidtov ´a, Mat ´uˇs Lango, and Ondˇrej Duˇsek. 2024. Leak, Cheat, Re- peat: Data Contamination and Evaluation Malprac- tices in Closed-Source LLMs. arXiv [Cs.CL]. Re- trieved from http://arxiv.org/abs/2402.03927

  35. [41]

    Hashimoto

    Yann Dubois, Bal ´azs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length- Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. arXiv preprint arXiv:2404.04475

  36. [42]

    Gonzalez, and Ion Stoica

    Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline

  37. [43]

    Hashimoto

    Xuechen Li, Tianyi Zhang, Yann Dubois, Ro- han Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Al- pacaEval: An Automatic Evaluator of Instruction- Following Models

  38. [44]

    Colin White, Samuel Dooley, Matthew Roberts, Ameya Pal, Benjamin Feuer, Sarthak Jain, et al. 2024. LiveBench: A Challenging, Contamination-Free LLM Benchmark. arXiv preprint arXiv:2406.19314

  39. [46]

    Elizabeth Clark, Tal August, Sarah Serrano, Nate Haduong, Suchin Gururangan, and Noah A. Smith

  40. [47]

    Chung-Hsuan Chiang and Hung-Yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2305.01937

  41. [48]

    Hugo Touvron, et al. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971

  42. [49]

    OpenAI, Josh Achiam, Steven Adler, Sand- hini Agarwal, Lama Ahmad, Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al

  43. [50]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Informa- tion Processing Systems, 36

  44. [51]

    Marilyn Strathern. 1997. ‘Improving ratings’: audit in the British University system. European Review, 5(3), 305–321. doi:10.1002/(SICI)1234- 981X(199707)5:3¡305::AID-EURO184¿3.0.CO;2- 4

  45. [52]

    arXiv preprint arXiv:2107.00061

    All That’s ’Human’ Is Not Gold: Evaluat- ing Human Evaluation of Generated Text. arXiv preprint arXiv:2107.00061

  46. [53]

    Stephanie Lin, Jacob Hilton, and Owain Evans

  47. [54]

    Brown, Adam Santoro, Aditya Gupta, Adri `a Garriga-Alonso, et al

    Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri `a Garriga-Alonso, et al. 2023. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. Availab...

  48. [55]

    Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020. FreeLB: En- hanced Adversarial Training for Natural Language Understanding. Available at https://arxiv. org/abs/1909.11764

  49. [57]

    Inze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sin...

  50. [59]

    Marina Sokolova, Nathalie Japkowicz, and Stan Szpakowicz. 2006. Beyond Accuracy, F-Score and ROC: A Family of Discriminant Measures for Performance Evaluation. In AI 2006: Ad- vances in Artificial Intelligence, Lecture Notes in Computer Science , V ol. 4304, pages 1015-1021. d...

  51. [64]

    Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Infer- ence. Available athttps://arxiv.org/abs/ 1704.05426

  52. [65]

    Mingchen Zhuge, Changsheng Zhao, Dy- lan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, J ¨urgen Schmid- huber. 2024. Agent-as-a-Judge: Evaluate Agents with Agents. Availabl...

  53. [2018]

    In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)

    Know What You Don’t Know: Unanswer- able Questions for SQuAD. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Retrieved from https://arxiv.org/abs/1806.03822

  54. [2019]

    GLUE: A Multi-Task Benchmark and Anal- ysis Platform for Natural Language Understanding

  55. [2021]

    In 30th USENIX Security Sympo- sium (USENIX Security 21) , pages 2633–2650

    Extracting Training Data from Large Lan- guage Models. In 30th USENIX Security Sympo- sium (USENIX Security 21) , pages 2633–2650. USENIX Association

  56. [2022]

    Available at https:// arxiv.org/abs/2109.07958

    TruthfulQA: Measuring How Models Mimic Human Falsehoods. Available at https:// arxiv.org/abs/2109.07958

  57. [2024]

    Efficient Adversarial Training in LLMs with Continuous Attacks

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.