REVIEW 3 major objections 3 minor 54 references
AI Benchmarks and Datasets for LLM Evaluation
T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper presents a catalog that maps 41 AI benchmarks and datasets to the seven EU Trustworthy AI requirements, letting practitioners choose evaluation tools aligned with the EU AI Act.
desk verdict A useful catalog whose only original contribution—the mapping to EU requirements—is asserted without method and contains internal inconsistencies; fixable, but not yet a reliable resource. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the crosswalk between two existing taxonomies: the seven EU requirements for Trustworthy AI and the list of 41 benchmarks and datasets. Table 2 is the crosswalk, with rows for benchmarks and columns for the seven requirements, marking each benchmark with the requirements it addresses. The paper also uses in-text tag lists such as '# Technical Robustness and safety' and '# Transparency' as a second, descriptive layer that records the same mapping. The crosswalk is what carries the argument: it turns a vague need for compliance-oriented evaluation into a concrete selection problem, and it reveals which requirements have sparse coverage.
What would settle it
A concrete check: compare each benchmark's in-text tag list with its Table 2 row, and observe whether entries such as ANLI, tagged '# Technical Robustness and safety' in Section 3.1 but placed under a different requirement in Table 2, can be reproduced; if even one such mismatch exists, the mapping is not self-consistent and the catalog cannot serve as a reliable compliance guide.
Extended reading notes
Core claim
The central claim is that the catalog, summarized in Table 2, enables practitioners to identify and use AI benchmarks throughout the AI system lifecycle by showing which of the seven EU Trustworthy AI requirements each benchmark addresses. The table assigns each of 41 benchmarks and datasets to one or more requirements, for example linking MMLU and HellaSwag to transparency, RobustBench to technical robustness and safety, and the AI Safety Benchmark v0.5 to technical robustness, accountability, and diversity. The authors connect this to the EU AI Act's obligations for general-purpose AI providers, including model evaluation and adversarial testing for models presenting systemic risk. The stated purpose is to enrich an existing holistic audit methodology with practical benchmarks, so that compliance assessments can point to concrete quantitative tests.
Load-bearing premise
The catalog's usefulness depends on the assignment of each benchmark to one or more of the seven EU requirements being correct, consistent, and reproducible from the benchmark descriptions.
Editorial extensions
If this is right
- A practitioner facing a specific EU AI Act concern can scan Table 2 and pick a benchmark for that requirement without reading each original benchmark paper.
- The table makes gaps visible: requirements with few or no rows, such as societal and environmental well-being or accountability, are places where evaluation tooling is still missing.
- Because the table is presented as part of an ongoing project, new benchmarks can be added to it as they appear, extending the same tagging scheme.
- The categorization gives compliance discussions a shared vocabulary, so that an audit or a conformity assessment can refer to concrete quantitative tests per Trustworthy AI requirement.
Reading between the lines
- A natural extension the paper does not perform is validating the tags: an expert panel or a user study could check each assignment, since the paper gives no method for deciding which requirement a benchmark tests.
- The same seven-column tagging could be refined to map benchmarks to specific obligations in the EU AI Act, such as the systemic-risk duties of general-purpose AI providers, rather than to the high-level requirements alone.
- The tagging scheme could also be applied to non-LLM AI benchmarks and to the Act's risk tiers, turning the catalog into a general compliance-oriented index.
- Since the table is a snapshot, it could be paired with a submission mechanism so the community keeps the assignments current; until then, Table 2 ages as new benchmarks appear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a catalog of 41 AI benchmarks and datasets, each accompanied by a short summary and a set of tags, with the stated goal of helping practitioners identify and utilize benchmarks for evaluating AI systems against the EU AI Act. The paper reviews the Z-Inspection methodology, the EU AI Act, and the COMPL-AI framework, then introduces the catalog and a mapping table (Table 2) that assigns each benchmark to one or more of the seven EU Trustworthy AI requirements. The paper makes no experimental or mathematical claims; its concrete contribution is the catalog and the benchmark-to-requirement mapping.
Significance. If the mapping in Table 2 were accurate and procedurally grounded, the catalog could serve as a useful practical resource for aligning LLM evaluation with the EU AI Act. The original benchmark summaries are mostly faithful to their cited sources, and the paper provides a concise comparative view of many popular benchmarks. However, the central value of the paper rests on the benchmark-to-requirement mapping, which is asserted without a stated method, contains internal inconsistencies, and includes at least one misidentified benchmark. Since the abstract promises that practitioners can use the catalog to select benchmarks for EU AI Act-related evaluation, these defects bear directly on the paper's core contribution. The paper offers no quantitative validation or machine-checked artifacts; the catalog is a manually assembled table that would need substantial rework to become reliable.
major comments (3)
- [Section 3.1 vs. Table 2] The tag assigned to ANLI in the text (§3.1, "Tags: # Dataset # Technical Robustness and safety") is inconsistent with Table 2, where the only mark for ANLI falls in the HAO (Human Agency and Oversight) column. Because Table 2 is the sole concrete deliverable of the paper, this discrepancy directly undermines the claim that the table faithfully summarizes the catalog's per-entry tags.
- [Sections 3.2, 3.5, 3.39 vs. Section 2.1] The mapping from benchmarks to the seven EU requirements is asserted without any stated rubric or decision procedure. For example, HellaSwag (§3.2), MMLU (§3.5), and ARC (§3.39) are each tagged # Transparency in Table 2, but Section 2.1 defines Transparency as traceability, explanation, disclosure of AI presence, and informing users of capabilities and limitations; it says nothing about general knowledge, commonsense reasoning, or question answering. No bridging argument explains why these benchmarks evaluate transparency, so a practitioner cannot infer which EU requirement a given benchmark addresses. Since the catalog's purpose is precisely to enable such inference, this unsupported mapping is the load-bearing weakness of the paper.
- [Section 3.41 and Table 2] Section 3.41 is titled "Visual Question Answering" but the text describes OK-VQA [17], a benchmark requiring external knowledge for visual question answering. Table 2 lists the entry as "VQA [17]", which does not match the described benchmark. This misidentification is a factual error in the catalog content itself, not merely a typographical issue, and it raises doubts about the accuracy of the other entries.
minor comments (3)
- [Throughout] There are several typos that should be corrected, including "Kewords" in §3.23, "Langugage" in §3.24, "CasulaBench(2)" in §3.17 and Table 2, and the doubled word "could could" in Section 2.1.
- [Table 2] For entries tagged only as datasets (e.g., CommonsenseQA, CORD-19, OpenAssistant Conversations), Table 2 leaves all requirement columns empty; the paper should state explicitly that empty cells mean the benchmark is not mapped to any of the seven requirements, since this is a meaningful choice rather than an omission.
- [Abstract and Section 1] The paper announces a project but does not provide a URL, repository, or other pointer to the project's outputs beyond Table 2; adding a link or describing where the catalog is maintained would strengthen the actionable value of the paper.
Circularity Check
No load-bearing circularity: the catalog is an asserted taxonomy; the only self-citation is motivational and does not force the benchmark-to-requirement mapping.
full rationale
The paper contains no derivation chain, no fitted parameters, and no predictive claim that could be equivalent to its own inputs by construction. Its contribution is an enumerated catalog of benchmarks/datasets in Sections 3.1-3.41 and a mapping matrix in Table 2 from those benchmarks to the seven EU Trustworthy AI requirements defined in Section 2.1. The abstract frames the deliverable as 'collecting and categorizing AI benchmarks,' and Table 2 is presented as a 'comprehensive list of AI benchmarks and datasets, along with the seven EU requirements for Trustworthy AI that each addresses.' No equation, algorithm, or theorem connects a benchmark's properties to its tags, so the tags are assertions rather than derivations; this is an under-justification/correctness concern, not circularity. For example, Section 3.10 tags IFEval with '# Technical Robustness and Safety # Transparency' without showing how the Section 2.1 definitions of those requirements entail those labels, and Section 3.1 tags ANLI as '# Technical Robustness and safety' while Table 2 places ANLI's mark in a different column. These are evidentiary and consistency problems for the catalog's stated purpose, not self-referential reductions. On self-citation: Section 2.1 describes 'The Z-Inspection® [45, 53, 35] is a novel process grounded in applied ethics,' and reference [53] includes author Todor Ivanov; the introduction also says the project aims to 'enrich this methodology with practical benchmarks.' This is a genuine but minor self-citation. It is not load-bearing, however, because the benchmarks listed in Section 3 and the tags in Table 2 are not derived from Z-Inspection; they are presented as independent descriptions of the cited benchmark papers, and the seven-requirement framework comes from the EU guidelines, not from the authors' prior work. Thus the central content of the paper retains independent evidentiary status, and no circular step can be exhibited.
Assumptions & free parameters
assumptions (3)
- domain assumption The seven EU requirements for Trustworthy AI are a suitable and sufficient categorization scheme for LLM evaluation benchmarks.
- domain assumption Each listed benchmark genuinely tests the EU requirement or requirements marked in Table 2.
- domain assumption Z-Inspection and COMPL-AI are valid foundations for linking benchmarks to EU AI Act compliance.
Cite this review
Pith. "Pith review of AI Benchmarks and Datasets for LLM Evaluation." pith.science (2026). https://pith.science/paper/SILNRJJ3
@misc{pith2026241201020,
author = {Pith},
title = {Pith review of: AI Benchmarks and Datasets for LLM Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/SILNRJJ3}},
note = {Machine review of arXiv:2412.01020}
}
read the original abstract
LLMs demand significant computational resources for both pre-training and fine-tuning, requiring distributed computing capabilities due to their large model sizes \cite{sastry2024computing}. Their complex architecture poses challenges throughout the entire AI lifecycle, from data collection to deployment and monitoring \cite{OECD_AIlifecycle}. Addressing critical AI system challenges, such as explainability, corrigibility, interpretability, and hallucination, necessitates a systematic methodology and rigorous benchmarking \cite{guldimann2024complai}. To effectively improve AI systems, we must precisely identify systemic vulnerabilities through quantitative evaluation, bolstering system trustworthiness. The enactment of the EU AI Act \cite{EUAIAct} by the European Parliament on March 13, 2024, establishing the first comprehensive EU-wide requirements for the development, deployment, and use of AI systems, further underscores the importance of tools and methodologies such as Z-Inspection. It highlights the need to enrich this methodology with practical benchmarks to effectively address the technical challenges posed by AI systems. To this end, we have launched a project that is part of the AI Safety Bulgaria initiatives \cite{AI_Safety_Bulgaria}, aimed at collecting and categorizing AI benchmarks. This will enable practitioners to identify and utilize these benchmarks throughout the AI system lifecycle.
Reference graph
Works this paper leans on
-
[17]
Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and R oozbeh Mot- taghi. Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019
work page 2019
-
[1]
https://aisafetybulgaria.c om/
Ai safety bulgaria, 2024. https://aisafetybulgaria.c om/
work page 2024
-
[2]
Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems
Christoph Brücke, Philipp Härtling, Rodrigo Escobar Pa lacios, Hamesh Patel, and Tilmann Rabl. Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems. Proc. VLDB Endow. , 16(12):3649–3661, 2023
work page 2023
- [3]
-
[4]
Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance
Paulo Henrique Couto, Quang Phuoc Ho, Nageeta Kumari, Be nedic- tus Kent Rachmat, Thanh Gia Hieu Khuong, Ihsan Ullah, and Lis heng Sun-Hosoya. Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance. CoRR, abs/2406.10294, 2024
arXiv 2024
-
[5]
Robustbench: a standardized adversarial ro bustness bench- mark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag , Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mit tal, and Matthias Hein. Robustbench: a standardized adversarial ro bustness bench- mark. In Joaquin Vanschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Processing Systems Track on Datasets a nd Bench- marks 1, ...
work page 2021
-
[6]
https://artificialintelligenceact.eu /the-act/
Eu ai act, 2024. https://artificialintelligenceact.eu /the-act/
work page 2024
-
[7]
https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
Ethics guidelines for trustworthy ai, 2019. https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
work page 2019
Show all 54 references
-
[8]
Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024
Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna Gueorguieva, Mislav Balunovi ć, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, and Martin Vech ev. Compl-ai framework: A technical interpretation and llm benchmarkin g suit...
2024
-
[9]
Measuring massive multita sk language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Ma ntas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multita sk language understanding. In 9th International Conference on Learning Representa- tions, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenRe...
2021
-
[10]
Measuri ng mathe- matical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Aro ra, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuri ng mathe- matical problem solving with the MATH dataset. In Joaquin Va nschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Proce...
2021
-
[11]
Weld, and Luke Zett lemoyer
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zett lemoyer. Trivi- aqa: A large scale distantly supervised challenge dataset f or reading com- prehension. In Regina Barzilay and Min-Yen Kan, editors, Proceedings of the 55th Annual Meeting of the Association for Computati...
2017
-
[12]
Openassistant conversations - democra tizing large lan- guage model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotir is Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh D uc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glush kov, Ar- nav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Ng uyen, ...
2023 arXiv
-
[13]
Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun S un. Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024
2024
-
[14]
Andrew Liu, Hongjian Zhou, Yining Hua, Omid Rohanian, A nshul Thakur, Lei Clifton, and David A. Clifton. Large language models in t he clinic: A comprehensive benchmark, 2024
2024
-
[15]
Glore: Evaluating logical reasoning of large langua ge models
Hanmeng Liu, Zhiyang Teng, Ruoxi Ning, Jian Liu, Qiji Zh ou, and Yue Zhang. Glore: Evaluating logical reasoning of large langua ge models. CoRR, abs/2310.09107, 2023
2023 arXiv
-
[16]
Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning
Zeyuan Ma, Hongshu Guo, Jiacheng Chen, Zhenrui Li, Guoj un Peng, Yue- Jiao Gong, Yining Ma, and Zhiguang Cao. Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Ha rdt,...
2023
-
[18]
Abstractive text summarization u sing sequence- to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Sant os, Çaglar Gülçehre, and Bing Xiang. Abstractive text summarization u sing sequence- to-sequence rnns and beyond. In Yoav Goldberg and Stefan Rie zler, edi- tors, Proceedings of the 20th SIGNLL Conference on Computational ...
2016
-
[19]
Adversarial NLI: A new benchmark for natura l language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, J ason Weston, and Douwe Kiela. Adversarial NLI: A new benchmark for natura l language understanding. In Dan Jurafsky, Joyce Chai, Natalie Schlut er, and Joel R. 23 Tetreault, editors, Proceedings of the 58th Annual Meeting...
2020
-
[20]
https://oecd.ai/en/ai-pr inciples
Ai system lifecycle, 2024. https://oecd.ai/en/ai-pr inciples
2024
-
[21]
The LAMBADA dataset: Word prediction requir- ing a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou , Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Ge mma Boleda, and Raquel Fernández. The LAMBADA dataset: Word prediction requir- ing a broad discourse context. In Proceedings of the 54th Annual Meeting o...
2016
-
[22]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Sa muel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark . CoRR, abs/2311.12022, 2023
2023 arXiv
-
[23]
Winogrande: An adversarial winograd schema challenge at sc ale
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula , and Yejin Choi. Winogrande: An adversarial winograd schema challenge at sc ale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intellige n...
2020
-
[24]
Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F
Girish Sastry, Lennart Heim, Haydn Belfield, Markus And erljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F. Trager, Shahar A vin, Adrian Weller, Yoshua Bengio ,...
2024
-
[25]
Sur- vey of different large language model architectures: Trends , benchmarks, and challenges
Minghao Shao, Abdul Basit, Ramesh Karri, and Muhammad S hafique. Sur- vey of different large language model architectures: Trends , benchmarks, and challenges. IEEE Access, 2024
2024
-
[26]
Concep tnet 5.5: An open multilingual graph of general knowledge
Robyn Speer, Joshua Chin, and Catherine Havasi. Concep tnet 5.5: An open multilingual graph of general knowledge. In Satinder S ingh and Shaul Markovitch, editors, Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, Cal...
2017
-
[27]
Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, an d Greg Durrett. Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing. In The Twelfth International Conference on Learning Representat ions, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview....
2024
-
[28]
Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024
Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu, Yanshu Li, J iashuo Sun, Juntao Li, Min Zhang, and Yu Cheng. Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024
2024
-
[29]
A corpus for reasoning about natural language gr ounded in photographs, 2019
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Hua jun Bai, and Yoav Artzi. A corpus for reasoning about natural language gr ounded in photographs, 2019
2019
-
[30]
Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024
Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongme i Zhang. Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024
2024
-
[31]
Le, Ed H
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebas tian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H . Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan L. B oyd-Graber...
2023
-
[32]
Commonsenseqa: A question answering challenge targeting c ommonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jon athan Berant. Commonsenseqa: A question answering challenge targeting c ommonsense knowledge. In Jill Burstein, Christy Doran, and Thamar Solo rio, editors, Proceedings of the 2019 Conference of the North American Chapt er...
2019
-
[33]
Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls
Weizhi Tang and Vaishak Belle. Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls. CoRR, abs/2407.05434, 2024
2024
-
[34]
Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C
Jeyan Thiyagalingam, Gregor von Laszewski, Junqi Yin, Murali Emani, Juri Papay, Gregg Barrett, Piotr Luszczek, Aristeidis Tsar is, Christine R. Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C. Fox, and Tony Hey. AI benchmarking for sc i...
2022
-
[35]
Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V
Dennis Vetter, Julia Amann, Frédérick Bruneault, Mega n Coffee, Boris Düdder, Alessio Gallucci, Thomas Krendl Gilbert, Thilo Hag endorff, Irmhild van Halem, Eleanore Hickman, Elisabeth Hildt, Sune Holm, Geor- gios Kararigas, Pedro Kringen, Vince I. Madai, Emilie Wiinb lad Mathez...
2023
-
[36]
Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, Kurt D
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, Kurt D. Bollacker, Rish i Bomassani, Marisa Ferrara Boston, Siméon Campos, Kal Chakra, Canyu Che n, Cody Colema...
2024 arXiv
-
[37]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpre et Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superg lue: A stickier benchmark for general-purpose language understa nding systems, 2020
2020
-
[38]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill , Omer Levy, and Samuel R. Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
2019
-
[39]
Murdick, Devvret Ris hi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D
Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Ris hi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuans an Wang, Chris ...
2004 arXiv
-
[40]
Mmlu-pro: A more robust and challenging multi-task la nguage un- derstanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhrani l Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jia ng, Tianle 26 Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and W enhu Chen. Mmlu-pro: A more robust and challenging multi-task la nguage...
2024 arXiv
-
[41]
CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models
Zeyu Wang. CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models. In Kam-Fai Wong, Min Zhang, Ruifeng Xu, Jing Li, Zhongyu Wei, Lin Gui, Bin Lian g, and Runcong Zhao, editors, Proceedings of the 10th SIGHAN Workshop on Chi...
2024
-
[42]
Logicv ista: Mul- timodal LLM logical reasoning benchmark in visual contexts
Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. Logicv ista: Mul- timodal LLM logical reasoning benchmark in visual contexts . CoRR, abs/2407.04973, 2024
2024 arXiv
-
[43]
Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering
Vikas Yadav, Steven Bethard, and Mihai Surdeanu. Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering. In Proceedings of the 2019 Conference on Empirical Meth- ods in Natural Language Processing and the 9th Internationa...
2019
-
[44]
Eval- uating the quality of hallucination benchmarks for large vi sion-language models
Bei Yan, Jie Zhang, Zheng Yuan, Shiguang Shan, and Xilin Chen. Eval- uating the quality of hallucination benchmarks for large vi sion-language models. CoRR, abs/2406.17115, 2024
2024
-
[45]
https://z-inspection.org/
Z-inspection, 2024. https://z-inspection.org/
2024
-
[46]
Zebralogic: Benchmarking the logical reasoning abili ty of language models,
-
[47]
Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi , and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguisti cs, A...
2019
-
[48]
Benchmarking trustworthiness of multimodal large language models: A comprehensive study
Yichi Zhang, Yao Huang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Yifan Wang, Huanran Chen, Xiao Yang, Xingxing Wei, Han g Su, Yinpeng Dong, and Jun Zhu. Benchmarking trustworthiness of multimodal large language models: A comprehensive study. CoRR, abs/2406.07057, 2024
2024 arXiv
-
[49]
Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024
Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, and Xuming Hu. Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024. 27
2024
-
[50]
Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024
Yihang Zheng, Bo Li, Zhenghao Lin, Yi Luo, Xuanhe Zhou, C hen Lin, Jinsong Su, Guoliang Li, and Shifu Li. Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024
2024
-
[51]
Instruction-followin g evaluation for large language models
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha B rahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-followin g evaluation for large language models. CoRR, abs/2311.07911, 2023
2023 arXiv
-
[52]
Causalbench: A comprehensive benchmark for ca usal learn- ing capability of large language models
Yu Zhou, Xingyu Wu, Beicheng Huang, Jibin Wu, Liang Feng , and Kay Chen Tan. Causalbench: A comprehensive benchmark for ca usal learn- ing capability of large language models. CoRR, abs/2404.06349, 2024
2024 arXiv
-
[53]
Roberto V. Zicari, John Brodersen, James Brusseau, Bor is Düdder, Timo Eichhorn, Todor Ivanov, Georgios Kararigas, Pedro Kringen , Melissa Mc- Cullough, Florian Möslein, Naveed Mushtaq, Gemma Roig, Nor man Stürtz, Karsten Tolle, Jesmin Jahan Tithi, Irmhild van Halem, and Ma gn...
2021
-
[2024]
https://huggingface.co/blog/yuchenlin/zebra-lo gic
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.