REVIEW 4 major objections 5 minor 2 cited by
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that current LLM benchmarks, from GLUE and MMLU to human and LLM judges, can be gamed, so top scores do not prove genuine language understanding.
desk verdict A readable but rough survey of known benchmark vulnerabilities; no new evidence or systematic method, so the advertised 'systematic analysis' overclaims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the benchmark score treated as an estimator of true performance, written as $B = P + \varepsilon$, together with the generalization gap $G(M) = E[S(M,B_{\text{test}})] - E[S(M,B_{\text{real}})]$. The paper uses these formalizations, plus the contamination rate and the mutual-information measure of leakage, to make explicit the ways a score can rise while capability does not; the named principle underneath is Goodhart's law, that a measure ceases to be good once it becomes a target. The formalisms do the argumentative work of showing that benchmark optimization, contamination, and judge bias are not hypothetical but quantifiable distortions.
What would settle it
Take the current top-ranked models on GLUE, MMLU, and a popular LLM-judge arena, run a contamination audit to remove any leaked examples, then test the same models on a freshly written, never-public set of analogous tasks in the same domains; if the rankings and scores stay nearly the same and the generalization gap $G(M)$ is near zero for most models, the paper's claim of pervasive benchmark vulnerability would be contradicted for those benchmarks.
Extended reading notes
Core claim
The paper's central claim is that the evaluation ecosystem for LLMs contains a structural mismatch: the measures being optimized—static benchmark scores and judge ratings—are not the capabilities the field actually wants, namely robust understanding and adaptation to novel tasks. It identifies three families of exploitation: benchmark hacking (overfitting and task-specific optimization), data contamination (training/test overlap that inflates scores), and evaluator bias (humans and LLM judges whose judgments are noisy, format-sensitive, self-preferring, or otherwise gameable). The authors formalize these distortions with simple equations: benchmark score $B$ decomposes as $B = P + \varepsilon$; contamination rate is $CR = |D_{\text{train}} \cap D_{\text{test}}| / |D_{\text{test}}|$; mutual information $I(\theta; D_{\text{test}} | D_{\text{train}})$ signals leakage; and a generalization gap $G(M) = E[S(M,B_{\text{test}})] - E[S(M,B_{\text{real}})]$ quantifies the distance between benchmark success and real-world performance. They conclude that no single existing evaluation method is trustworthy on its own, that near-perfect benchmark scores deserve skepticism, and that future frameworks should be zero-day, zero-shot, domain-specific, and governed so they can be iterated before models overfit to them.
Load-bearing premise
The argument that these vulnerabilities are pervasive depends on the unstated assumption that the studies the paper selects are a representative sample of LLM evaluation practice and that no significant counter-evidence was left out; the paper gives no systematic search strategy, inclusion criteria, or quantitative synthesis to support that.
Editorial extensions
If this is right
- Leaderboard rankings should be interpreted as upper-bound marketing claims rather than measurements of understanding, because reported scores can be inflated by contamination and overfitting.
- Near-perfect or saturated scores on static benchmarks are weak evidence of progress; they may reflect exploitation of dataset artifacts rather than model competence.
- Human evaluation cannot serve as an unbiased gold standard: annotators disagree on the same outputs and can be fooled by superficial fluency.
- LLM-as-judge scores need debiasing: judges show self-preference, length bias, and sensitivity to prompt phrasing, so single-judge numbers are not stable.
- Evaluation practice should shift toward dynamic, zero-day, domain-specific tasks with contamination audits and transparent reporting of methodology.
Reading between the lines
- A consequence the authors leave implicit: new benchmarks should ship with a contamination audit and a preregistered refresh schedule before their numbers are treated as evidence of capability.
- The same mechanisms would likely apply to code and mathematics benchmarks, where public solutions make contamination easier; the paper does not test this, but its logic predicts similar inflation.
- The generalization-gap formalism suggests a cheap diagnostic the authors do not propose: report $G(M)$ alongside every published benchmark score so brittle leaders are visible.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a narrative review of vulnerabilities in large language model (LLM) evaluation, arguing that current benchmarks such as GLUE and MMLU are susceptible to gaming, data contamination, and evaluator bias, and that high leaderboard scores therefore create a false impression of genuine language understanding. After a brief history of NLP benchmarks, the paper surveys benchmark overfitting, public dataset leakage, test set contamination, task-specific optimization, adversarial benchmarking, human evaluation bias, and LLM-as-judge biases, with illustrative mathematical notation. It concludes that no single evaluation method is reliable and proposes future work on zero-day, zero-shot evaluation protocols and consortium-based dynamic benchmarking.
Significance. If the paper's central claim were established—that current benchmarks pervasively misrepresent LLM capabilities—it would have substantial implications for how the field measures progress and allocates trust in leaderboard results. The manuscript does collect a useful set of known critique themes, including contamination detection, annotation artifacts, self-preference in LLM judges, and human evaluator inconsistency. However, the paper provides no new experiments, no quantitative synthesis of existing findings, and no methodologically grounded framework for its 'systematic analysis' claim. Its value as a survey is further weakened by multiple misattributed citations and internal textual errors. The significance of the topic is high, but the contribution as written is a checklist of known issues rather than a substantiated systematic assessment.
major comments (4)
- [Abstract; Sections 2-3] The paper repeatedly claims a 'systematic analysis' and 'pervasive vulnerabilities across the evaluation spectrum', but no methodology is described anywhere in the manuscript. There is no search strategy, no inclusion or exclusion criteria, no corpus description, and no quantitative synthesis. The examples discussed (GPT-3 on LAMBADA, BERT on SQuAD, human evaluator and LLM-judge studies) are a curated selection, so the generalization to 'pervasive' vulnerabilities is unsupported. This is load-bearing because the central conclusion that benchmarks give a false perception of progress depends on the representativeness of the cited studies; without a defined method, the claim is an overgeneralization from an anecdotal sample.
- [Section 3.1, Eq. (1); Section 3.2; Section 3.3] The mathematical formulations (B = P + ε, the contamination rate CR, the mutual information I(θ; D_test | D_train), and the exposure metric) are purely definitional and do not constitute evidence that any specific benchmark is inaccurate. The paper never estimates these quantities on actual models or datasets, so the equations provide no support for the claim that benchmarks misrepresent true performance. If the intent is to formalize known failure modes, this should be stated explicitly; as written, the formalism may give the misleading impression of a quantitative proof.
- [Section 3.11; Abstract] The central claim that benchmarks create 'a false perception of progress' requires a benchmark-independent operationalization of 'true performance,' which the paper never provides. Without such a yardstick, the claim is unfalsifiable as stated. A concrete test would be to compare leaderboard rankings with held-out dynamic evaluations or with performance on well-defined real-world tasks, and to measure the rank correlation; the paper contains no such test. Similarly, the assertion in Section 3.11 that 'models approach or achieve perfect scores on established benchmarks' is an empirical claim presented without data or citation.
- [Sections 3.4 and 3.5; References [6], [20], [56]] Several citations do not support the claims to which they are attached, undermining the reliability of the survey. In Section 3.5, universal adversarial triggers are attributed to Wallace et al. with citation [6], but reference [6] is Raji and Buolamwini (2019), not Wallace et al. In Section 3.4, the finding that BERT models use shallow heuristics on MNLI is attributed to 'McCoy' with citation [56], but reference [56] is the MNLI dataset paper by Williams et al. (2018), not McCoy et al. In Section 3.1, the claims about GPT-3 on LAMBADA and BERT on SQuAD are supported by citations [25] and [20], which are the LAMBADA and SQuAD 2.0 dataset papers, not analyses of overfitting or pattern exploitation. These misattributions mean the cited evidence cannot be checked by the reader and weaken the manuscript's authority.
minor comments (5)
- [Section 3.6] This section contains a duplicated paragraph: the text beginning 'One key limitation of human evaluations is the inherent inconsistency...' and ending '...poorly constructed [45]' appears twice verbatim, interrupting the flow of the argument.
- [Section 3.7] The section ends mid-sentence with 'as well as the exploration of entire' and no continuation; the sentence and the section appear to be truncated.
- [References] The reference list has inconsistencies: reference [45] is used for two different papers (Chiang et al., 'Chatbot Arena', and Zheng et al., 'Judging LLM-as-a-Judge'), and references [18] and [41] both cite the same paper by Dubois et al. (2024). These duplicate entries need to be resolved.
- [Section 3.4] The text refers to 'McCoy demonstrated' without a corresponding reference entry; the actual citation [56] points to the MNLI dataset paper rather than to any work by McCoy, so the intended source should be identified and cited correctly.
- [Section 1.1] The terminology 'Human-as-Judge' and 'LLM-as-Judge' is used with inconsistent capitalization and hyphenation; the authors should adopt a single consistent form throughout.
Circularity Check
No circularity: the paper is a narrative survey whose claims rest on external citations and illustrative definitions, not on a derivation that reduces to its own inputs.
full rationale
This paper is a narrative survey rather than an empirical derivation, and I find no circular step that reduces a claimed result to its own inputs. The formal notation in Sections 3.1-3.4 (e.g., B = P + epsilon, CR = |Dtrain intersect Dtest| / |Dtest|, and I(theta; Dtest | Dtrain)) is definitional scaffolding used to restate cited findings; no parameter is fitted and no quantity is predicted from a prior that already contains it. The load-bearing evidence for 'pervasive vulnerabilities' comes from external studies (GPT-3 on LAMBADA, BERT on SQuAD, human-evaluator inconsistency in Clark et al., LLM-judge biases in Zheng et al.), not from the authors' own prior work, and no uniqueness theorem or ansatz is imported from self-citations. The central weakness -- generalizing from a curated selection of mostly older model/task observations to 'from basic metrics to ... GLUE and MMLU' without a systematic search or an independent operationalization of 'true performance' -- is a representativeness and falsifiability problem, not a circularity problem. Accordingly the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Goodhart's law applies to LLM benchmarks: once a score becomes a target, it ceases to be a good measure.
- domain assumption The cited studies' findings are accurate and transferable to current LLMs.
Cite this review
Pith. "Pith review of The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?." pith.science (2026). https://pith.science/paper/WE255WAS
@misc{pith2026241203597,
author = {Pith},
title = {Pith review of: The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?},
year = {2026},
howpublished = {\url{https://pith.science/paper/WE255WAS}},
note = {Machine review of arXiv:2412.03597}
}
read the original abstract
The pursuit of leaderboard rankings in Large Language Models (LLMs) has created a fundamental paradox: models excel at standardized tests while failing to demonstrate genuine language understanding and adaptability. Our systematic analysis of NLP evaluation frameworks reveals pervasive vulnerabilities across the evaluation spectrum, from basic metrics to complex benchmarks like GLUE and MMLU. These vulnerabilities manifest through benchmark exploitation, dataset contamination, and evaluation bias, creating a false perception of progress in language understanding capabilities. Through extensive review of contemporary evaluation approaches, we identify significant limitations in static benchmark designs, human evaluation protocols, and LLM-as-judge frameworks, all of which compromise the reliability of current performance assessments. As LLM capabilities evolve and existing benchmarks become redundant, we lay the groundwork for new evaluation methods that resist manipulation, minimize data contamination, and assess domain-specific tasks. This requires frameworks that are adapted dynamically, addressing current limitations and providing a more accurate reflection of LLM performance.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 2 Pith papers
-
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Chat UI and API access to the same chatbot produce different accuracy, consistency, citation, and refusal behaviors on safety benchmarks, and web search changes these patterns further.
-
Rethinking the Understanding Ability across LLMs through Mutual Information
The paper uses token-level recoverability as a computable lower bound on mutual information to compare LLMs and to fine-tune them, finding encoder-only models preserve information better than decoder-only models.
Reference graph
Works this paper leans on
-
[45]
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, et al. 2024. Chat- bot Arena: An Open Platform for Evaluat- ing LLMs by Human Preference. arXiv preprint arXiv:2403.04132
arXiv 2024
-
[6]
Inioluwa Raji and Joy Buolamwini. 2019. Ac- tionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (AIES ’19) , pages 429–435. ACM. DOI: 10.1145/3306618.3314244
arXiv 2019
-
[56]
GPT-4 Technical Report
-
[25]
Denis Paperno, Germ ´an Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fern ´andez. 2016. The LAMBADA dataset: Word prediction requiring a broad dis- course context
work page 2016
-
[20]
Pranav Rajpurkar, Robin Jia, and Percy Liang
-
[1]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert- V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litw...
work page 2020
-
[2]
Nicholas Carlini, Florian Tram `er, Eric Wallace, Matthew Jagielski, Ariel Herbert-V oss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, ´Ulfar Erlingsson, Alina Oprea, and Colin Raffel
-
[3]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Un- derstanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , pages 4171–4186. As...
work page 2019
Show all 65 references
-
[4]
Mark Chen, Jerry Tworek, Heewoo Jun, Qim- ing Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott...
2021
-
[5]
Alec Radford, Karthik Narasimhan, Tim Sali- mans, Ilya Sutskever, and others. 2018. Improv- ing Language Understanding by Generative Pre- training. OpenAI. Preprint, pages 1–12
2018
-
[7]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Atten- tion is All You Need. In Advances in Neural Infor- mation Processing Systems (NeurIPS 2017). Curran Associates, Inc
2017
-
[8]
Eliezer Yudkowsky. 2008. Artificial Intelligence as a Positive and Negative Factor in Global Risk. In Nick Bostrom and Milan M. Cirkovic (eds.), Global Catastrophic Risks, pages 308–345. Oxford University Press
2008
-
[9]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2017. Men Also Like Shopping: Reducing Gender Bias Amplifica- tion using Corpus-level Constraints. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , pages 2979–2...
2017
-
[10]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman
-
[11]
Alex Wang, Yada Pruksachatkun, Nikita Nan- gia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2020. Super- GLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
2020
-
[12]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Ja- cob Steinhardt. 2021. Measuring Massive Multitask Language Understanding
2021
-
[13]
Open LLM Leaderboard. n.d. Open LLM Leaderboard - a Hugging Face Space by open-llm-leaderboard-old. Retrieved from https://huggingface.co/spaces/ open-llm-leaderboard-old/open_ llm_leaderboard
-
[14]
Manning, Christopher R ´e, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Ya- sunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Bin- hang Yuan, Bobby Yan, Ce Zhang, Christian Cos- grove, Christopher D. Manning, Christopher R ´e, Diana Acos...
2023
-
[15]
March 2024
Ziyou Yan. March 2024. Task-Specific LLM Evals that Do & Don’t Work. eugeneyan.com. Retrieved from https://eugeneyan.com/ writing/evals/
2024
-
[16]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?
2019
-
[17]
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...
2021
-
[18]
Hashimoto
Yann Dubois, Bal ´azs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length- Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
2024
-
[19]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: A Method for Au- tomatic Evaluation of Machine Translation. InPro- ceedings of the 40th Annual Meeting on Associa- tion for Computational Linguistics, pages 311–318. Association for Computational Linguistics
2002
-
[21]
Saiful Bari, and Haidar Khan
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Al- subaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, M. Saiful Bari, and Haidar Khan. 2024. When Bench- marks are Targets: Revealing the Sensitivity of ...
2024
-
[22]
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019. Universal Ad- versarial Triggers for Attacking and Analyzing NLP. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Process- ing and the 9th International Joint Conferenc...
2019
-
[23]
Bowman and George E
Samuel R. Bowman and George E. Dahl. 2021. What Will it Take to Fix Benchmarking in Natural Language Understanding?
2021
-
[24]
Ido Dagan. 2006. The PASCAL Recognising Textual Entailment Challenge. In Machine Learn- ing Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tex- tual Entailment, pages 177–190. Springer Berlin Heidelberg
2006
-
[26]
Nicholas Carlini, Chang Liu, ´Ulfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The Secret Sharer: Evaluating and Testing Unintended Memo- rization in Neural Networks
2019
-
[27]
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Ad- versarial NLI: A New Benchmark for Natural Lan- guage Understanding
2020
-
[28]
Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. 2020. Learning the Difference that Makes a Difference with Counterfactually- Augmented Data
2020
-
[29]
Bowman, and Noah A
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith. 2018. Annotation Artifacts in Natural Language Inference Data
2018
-
[30]
Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Sum- marization Branches Out, pages 74–81. Associa- tion for Computational Linguistics
2004
-
[31]
Satanjeev Banerjee and Alon Lavie. 2005. ME- TEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrin- sic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–7...
2005
-
[32]
Jesse Dodge, Maarten Sap, Ana Marasovi ´c, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Docu- menting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
2021
-
[33]
Raphael Olivier and Bhiksha Raj. 2023. How many perturbations break this model? Evaluating robustness beyond adversarial accuracy
2023
-
[34]
Yogesh Balaji, Tom Goldstein, and Judy Hoff- man. 2019. Instance Adaptive Adversarial Train- ing: Improved Accuracy Tradeoffs in Neural Nets
2019
-
[35]
Sophie Xhonneux, Alessandro Sordoni, Stephan G¨unnemann, Gauthier Gidel, and Leo Schwinn
-
[36]
Moustafa Alzantot, Yash Sharma, Ahmed Elgo- hary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating Natural Language Adver- sarial Examples
2018
-
[37]
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Alek- sander Madry. 2019. Adversarial Examples Are Not Bugs, They Are Features
2019
-
[38]
Aviral Kale, Truc Nguyen, Jack Harris, Chang Li, Jiayi Zhang, and Xiaoli Ma. 2023. Provenance Documentation to Enable Explainable and Trust- worthy AI: A Literature Review. Data Intelligence, 5(1):139–162
2023
-
[39]
Hashimoto
Yanai Oren, Nicolas Meister, Niladri Chatterji, Firoj Ladhak, and Tatsunori B. Hashimoto. 2023. Proving Test Set Contamination in Black Box Language Models. arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2310.17623
2023 arXiv
-
[40]
Silvio Balloccu, Petra Schmidtov ´a, Mat ´uˇs Lango, and Ondˇrej Duˇsek. 2024. Leak, Cheat, Re- peat: Data Contamination and Evaluation Malprac- tices in Closed-Source LLMs. arXiv [Cs.CL]. Re- trieved from http://arxiv.org/abs/2402.03927
2024 arXiv
-
[41]
Hashimoto
Yann Dubois, Bal ´azs Galambosi, Percy Liang, and Tatsunori B. Hashimoto. 2024. Length- Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. arXiv preprint arXiv:2404.04475
2024 arXiv
-
[42]
Gonzalez, and Ion Stoica
Tianle Li, Wei-Lin Chiang, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline
2024
-
[43]
Hashimoto
Xuechen Li, Tianyi Zhang, Yann Dubois, Ro- han Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Al- pacaEval: An Automatic Evaluator of Instruction- Following Models
2023
-
[44]
Colin White, Samuel Dooley, Matthew Roberts, Ameya Pal, Benjamin Feuer, Sarthak Jain, et al. 2024. LiveBench: A Challenging, Contamination-Free LLM Benchmark. arXiv preprint arXiv:2406.19314
2024 arXiv
-
[46]
Elizabeth Clark, Tal August, Sarah Serrano, Nate Haduong, Suchin Gururangan, and Noah A. Smith
-
[47]
Chung-Hsuan Chiang and Hung-Yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations? arXiv [Cs.CL]. Retrieved from http://arxiv.org/abs/2305.01937
2023 arXiv
-
[48]
Hugo Touvron, et al. 2023. LLaMA: Open and Efficient Foundation Language Models. arXiv preprint arXiv:2302.13971
2023 arXiv
-
[49]
OpenAI, Josh Achiam, Steven Adler, Sand- hini Agarwal, Lama Ahmad, Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Al- tenschmidt, Sam Altman, Shyamal Anadkat, et al
-
[50]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Informa- tion Processing Systems, 36
2024
-
[51]
Marilyn Strathern. 1997. ‘Improving ratings’: audit in the British University system. European Review, 5(3), 305–321. doi:10.1002/(SICI)1234- 981X(199707)5:3¡305::AID-EURO184¿3.0.CO;2- 4
1997 doi
-
[52]
arXiv preprint arXiv:2107.00061
All That’s ’Human’ Is Not Gold: Evaluat- ing Human Evaluation of Generated Text. arXiv preprint arXiv:2107.00061
-
[53]
Stephanie Lin, Jacob Hilton, and Owain Evans
-
[54]
Brown, Adam Santoro, Aditya Gupta, Adri `a Garriga-Alonso, et al
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adri `a Garriga-Alonso, et al. 2023. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. Availab...
2023 arXiv
-
[55]
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020. FreeLB: En- hanced Adversarial Training for Natural Language Understanding. Available at https://arxiv. org/abs/1909.11764
2020 arXiv
-
[57]
Inze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sin...
2023
-
[59]
Marina Sokolova, Nathalie Japkowicz, and Stan Szpakowicz. 2006. Beyond Accuracy, F-Score and ROC: A Family of Discriminant Measures for Performance Evaluation. In AI 2006: Ad- vances in Artificial Intelligence, Lecture Notes in Computer Science , V ol. 4304, pages 1015-1021. d...
2006 doi
-
[64]
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. A Broad-Coverage Challenge Corpus for Sentence Understanding through Infer- ence. Available athttps://arxiv.org/abs/ 1704.05426
2018 arXiv
-
[65]
Mingchen Zhuge, Changsheng Zhao, Dy- lan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, J ¨urgen Schmid- huber. 2024. Agent-as-a-Judge: Evaluate Agents with Agents. Availabl...
2024 arXiv
-
[2018]
In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Know What You Don’t Know: Unanswer- able Questions for SQuAD. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP) . Retrieved from https://arxiv.org/abs/1806.03822
2018 arXiv
-
[2019]
GLUE: A Multi-Task Benchmark and Anal- ysis Platform for Natural Language Understanding
-
[2021]
In 30th USENIX Security Sympo- sium (USENIX Security 21) , pages 2633–2650
Extracting Training Data from Large Lan- guage Models. In 30th USENIX Security Sympo- sium (USENIX Security 21) , pages 2633–2650. USENIX Association
-
[2022]
Available at https:// arxiv.org/abs/2109.07958
TruthfulQA: Measuring How Models Mimic Human Falsehoods. Available at https:// arxiv.org/abs/2109.07958
-
[2024]
Efficient Adversarial Training in LLMs with Continuous Attacks
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.