Pith. sign in

REVIEW 3 major objections 3 minor 54 references

AI Benchmarks and Datasets for LLM Evaluation

T0 review · 3 major / 3 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper presents a catalog that maps 41 AI benchmarks and datasets to the seven EU Trustworthy AI requirements, letting practitioners choose evaluation tools aligned with the EU AI Act.

desk verdict A useful catalog whose only original contribution—the mapping to EU requirements—is asserted without method and contains internal inconsistencies; fixable, but not yet a reliable resource. read the letter →

arxiv 2412.01020 v1 pith:SILNRJJ3 submitted 2024-12-02 cs.DC cs.AI

classification cs.DCcs.AI
keywords LLMevaluationAIbenchmarksEUActTrustworthyrequirementsbenchmarkcategorizationdatasetcatalogsystemlifecycle
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the EU AI Act creates a practical need: practitioners must choose benchmarks that test whether an LLM satisfies the Act's requirements, but no single list organizes the available benchmarks that way. To meet that need, the authors launch a project that collects and categorizes AI benchmarks and datasets, with the concrete deliverable being a table that maps 41 of them to the seven EU requirements for Trustworthy AI: human agency and oversight, technical robustness and safety, privacy and data governance, transparency, diversity and fairness, societal and environmental well-being, and accountability. If the mapping is trustworthy, the table gives a quick way to see which benchmarks exist for which compliance concern, and where evaluation tools are missing. The paper does not introduce a new benchmark or a new evaluation result; its contribution is the categorization itself.

What carries the argument

The machinery is the crosswalk between two existing taxonomies: the seven EU requirements for Trustworthy AI and the list of 41 benchmarks and datasets. Table 2 is the crosswalk, with rows for benchmarks and columns for the seven requirements, marking each benchmark with the requirements it addresses. The paper also uses in-text tag lists such as '# Technical Robustness and safety' and '# Transparency' as a second, descriptive layer that records the same mapping. The crosswalk is what carries the argument: it turns a vague need for compliance-oriented evaluation into a concrete selection problem, and it reveals which requirements have sparse coverage.

What would settle it

A concrete check: compare each benchmark's in-text tag list with its Table 2 row, and observe whether entries such as ANLI, tagged '# Technical Robustness and safety' in Section 3.1 but placed under a different requirement in Table 2, can be reproduced; if even one such mismatch exists, the mapping is not self-consistent and the catalog cannot serve as a reliable compliance guide.

Watch

Extended reading notes

Core claim

The central claim is that the catalog, summarized in Table 2, enables practitioners to identify and use AI benchmarks throughout the AI system lifecycle by showing which of the seven EU Trustworthy AI requirements each benchmark addresses. The table assigns each of 41 benchmarks and datasets to one or more requirements, for example linking MMLU and HellaSwag to transparency, RobustBench to technical robustness and safety, and the AI Safety Benchmark v0.5 to technical robustness, accountability, and diversity. The authors connect this to the EU AI Act's obligations for general-purpose AI providers, including model evaluation and adversarial testing for models presenting systemic risk. The stated purpose is to enrich an existing holistic audit methodology with practical benchmarks, so that compliance assessments can point to concrete quantitative tests.

Load-bearing premise

The catalog's usefulness depends on the assignment of each benchmark to one or more of the seven EU requirements being correct, consistent, and reproducible from the benchmark descriptions.

Editorial extensions

If this is right

  • A practitioner facing a specific EU AI Act concern can scan Table 2 and pick a benchmark for that requirement without reading each original benchmark paper.
  • The table makes gaps visible: requirements with few or no rows, such as societal and environmental well-being or accountability, are places where evaluation tooling is still missing.
  • Because the table is presented as part of an ongoing project, new benchmarks can be added to it as they appear, extending the same tagging scheme.
  • The categorization gives compliance discussions a shared vocabulary, so that an audit or a conformity assessment can refer to concrete quantitative tests per Trustworthy AI requirement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not perform is validating the tags: an expert panel or a user study could check each assignment, since the paper gives no method for deciding which requirement a benchmark tests.
  • The same seven-column tagging could be refined to map benchmarks to specific obligations in the EU AI Act, such as the systemic-risk duties of general-purpose AI providers, rather than to the high-level requirements alone.
  • The tagging scheme could also be applied to non-LLM AI benchmarks and to the Act's risk tiers, turning the catalog into a general compliance-oriented index.
  • Since the table is a snapshot, it could be paired with a submission mechanism so the community keeps the assignments current; until then, Table 2 ages as new benchmarks appear.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. This paper presents a catalog of 41 AI benchmarks and datasets, each accompanied by a short summary and a set of tags, with the stated goal of helping practitioners identify and utilize benchmarks for evaluating AI systems against the EU AI Act. The paper reviews the Z-Inspection methodology, the EU AI Act, and the COMPL-AI framework, then introduces the catalog and a mapping table (Table 2) that assigns each benchmark to one or more of the seven EU Trustworthy AI requirements. The paper makes no experimental or mathematical claims; its concrete contribution is the catalog and the benchmark-to-requirement mapping.

Significance. If the mapping in Table 2 were accurate and procedurally grounded, the catalog could serve as a useful practical resource for aligning LLM evaluation with the EU AI Act. The original benchmark summaries are mostly faithful to their cited sources, and the paper provides a concise comparative view of many popular benchmarks. However, the central value of the paper rests on the benchmark-to-requirement mapping, which is asserted without a stated method, contains internal inconsistencies, and includes at least one misidentified benchmark. Since the abstract promises that practitioners can use the catalog to select benchmarks for EU AI Act-related evaluation, these defects bear directly on the paper's core contribution. The paper offers no quantitative validation or machine-checked artifacts; the catalog is a manually assembled table that would need substantial rework to become reliable.

major comments (3)
  1. [Section 3.1 vs. Table 2] The tag assigned to ANLI in the text (§3.1, "Tags: # Dataset # Technical Robustness and safety") is inconsistent with Table 2, where the only mark for ANLI falls in the HAO (Human Agency and Oversight) column. Because Table 2 is the sole concrete deliverable of the paper, this discrepancy directly undermines the claim that the table faithfully summarizes the catalog's per-entry tags.
  2. [Sections 3.2, 3.5, 3.39 vs. Section 2.1] The mapping from benchmarks to the seven EU requirements is asserted without any stated rubric or decision procedure. For example, HellaSwag (§3.2), MMLU (§3.5), and ARC (§3.39) are each tagged # Transparency in Table 2, but Section 2.1 defines Transparency as traceability, explanation, disclosure of AI presence, and informing users of capabilities and limitations; it says nothing about general knowledge, commonsense reasoning, or question answering. No bridging argument explains why these benchmarks evaluate transparency, so a practitioner cannot infer which EU requirement a given benchmark addresses. Since the catalog's purpose is precisely to enable such inference, this unsupported mapping is the load-bearing weakness of the paper.
  3. [Section 3.41 and Table 2] Section 3.41 is titled "Visual Question Answering" but the text describes OK-VQA [17], a benchmark requiring external knowledge for visual question answering. Table 2 lists the entry as "VQA [17]", which does not match the described benchmark. This misidentification is a factual error in the catalog content itself, not merely a typographical issue, and it raises doubts about the accuracy of the other entries.
minor comments (3)
  1. [Throughout] There are several typos that should be corrected, including "Kewords" in §3.23, "Langugage" in §3.24, "CasulaBench(2)" in §3.17 and Table 2, and the doubled word "could could" in Section 2.1.
  2. [Table 2] For entries tagged only as datasets (e.g., CommonsenseQA, CORD-19, OpenAssistant Conversations), Table 2 leaves all requirement columns empty; the paper should state explicitly that empty cells mean the benchmark is not mapped to any of the seven requirements, since this is a meaningful choice rather than an omission.
  3. [Abstract and Section 1] The paper announces a project but does not provide a URL, repository, or other pointer to the project's outputs beyond Table 2; adding a link or describing where the catalog is maintained would strengthen the actionable value of the paper.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circularity: the catalog is an asserted taxonomy; the only self-citation is motivational and does not force the benchmark-to-requirement mapping.

full rationale

The paper contains no derivation chain, no fitted parameters, and no predictive claim that could be equivalent to its own inputs by construction. Its contribution is an enumerated catalog of benchmarks/datasets in Sections 3.1-3.41 and a mapping matrix in Table 2 from those benchmarks to the seven EU Trustworthy AI requirements defined in Section 2.1. The abstract frames the deliverable as 'collecting and categorizing AI benchmarks,' and Table 2 is presented as a 'comprehensive list of AI benchmarks and datasets, along with the seven EU requirements for Trustworthy AI that each addresses.' No equation, algorithm, or theorem connects a benchmark's properties to its tags, so the tags are assertions rather than derivations; this is an under-justification/correctness concern, not circularity. For example, Section 3.10 tags IFEval with '# Technical Robustness and Safety # Transparency' without showing how the Section 2.1 definitions of those requirements entail those labels, and Section 3.1 tags ANLI as '# Technical Robustness and safety' while Table 2 places ANLI's mark in a different column. These are evidentiary and consistency problems for the catalog's stated purpose, not self-referential reductions. On self-citation: Section 2.1 describes 'The Z-Inspection® [45, 53, 35] is a novel process grounded in applied ethics,' and reference [53] includes author Todor Ivanov; the introduction also says the project aims to 'enrich this methodology with practical benchmarks.' This is a genuine but minor self-citation. It is not load-bearing, however, because the benchmarks listed in Section 3 and the tags in Table 2 are not derived from Z-Inspection; they are presented as independent descriptions of the cited benchmark papers, and the seven-requirement framework comes from the EU guidelines, not from the authors' prior work. Thus the central content of the paper retains independent evidentiary status, and no circular step can be exhibited.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper's claims rest on the EU Trustworthy AI taxonomy, on the accuracy of the benchmark-to-requirement mappings, and on the authority of Z-Inspection and COMPL-AI. No free parameters or invented entities are involved because the paper is a catalog, not a derivation.

assumptions (3)
  • domain assumption The seven EU requirements for Trustworthy AI are a suitable and sufficient categorization scheme for LLM evaluation benchmarks.
    Invoked throughout Section 2 and Table 2; the entire tagging scheme depends on this without justification.
  • domain assumption Each listed benchmark genuinely tests the EU requirement or requirements marked in Table 2.
    The tags are asserted without evaluation; at least one assignment, ANLI, contradicts the paper's own text tag in Section 3.1.
  • domain assumption Z-Inspection and COMPL-AI are valid foundations for linking benchmarks to EU AI Act compliance.
    Sections 2.1 to 2.3 adopt these frameworks as given; author Ivanov is a co-author of Z-Inspection, so there is a self-citation element.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AI Benchmarks and Datasets for LLM Evaluation." pith.science (2026). https://pith.science/paper/SILNRJJ3

@misc{pith2026241201020,
  author       = {Pith},
  title        = {Pith review of: AI Benchmarks and Datasets for LLM Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SILNRJJ3}},
  note         = {Machine review of arXiv:2412.01020}
}
read the original abstract

LLMs demand significant computational resources for both pre-training and fine-tuning, requiring distributed computing capabilities due to their large model sizes \cite{sastry2024computing}. Their complex architecture poses challenges throughout the entire AI lifecycle, from data collection to deployment and monitoring \cite{OECD_AIlifecycle}. Addressing critical AI system challenges, such as explainability, corrigibility, interpretability, and hallucination, necessitates a systematic methodology and rigorous benchmarking \cite{guldimann2024complai}. To effectively improve AI systems, we must precisely identify systemic vulnerabilities through quantitative evaluation, bolstering system trustworthiness. The enactment of the EU AI Act \cite{EUAIAct} by the European Parliament on March 13, 2024, establishing the first comprehensive EU-wide requirements for the development, deployment, and use of AI systems, further underscores the importance of tools and methodologies such as Z-Inspection. It highlights the need to enrich this methodology with practical benchmarks to effectively address the technical challenges posed by AI systems. To this end, we have launched a project that is part of the AI Safety Bulgaria initiatives \cite{AI_Safety_Bulgaria}, aimed at collecting and categorizing AI benchmarks. This will enable practitioners to identify and utilize these benchmarks throughout the AI system lifecycle.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 41 canonical work pages

  1. [17]

    Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019

    Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and R oozbeh Mot- taghi. Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019

  2. [1]

    https://aisafetybulgaria.c om/

    Ai safety bulgaria, 2024. https://aisafetybulgaria.c om/

  3. [2]

    Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems

    Christoph Brücke, Philipp Härtling, Rodrigo Escobar Pa lacios, Hamesh Patel, and Tilmann Rabl. Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems. Proc. VLDB Endow. , 16(12):3649–3661, 2023

  4. [3]

    https://compl-ai.org/

    Compl-ai framework, 2024. https://compl-ai.org/

  5. [4]

    Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance

    Paulo Henrique Couto, Quang Phuoc Ho, Nageeta Kumari, Be nedic- tus Kent Rachmat, Thanh Gia Hieu Khuong, Ihsan Ullah, and Lis heng Sun-Hosoya. Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance. CoRR, abs/2406.10294, 2024

  6. [5]

    Robustbench: a standardized adversarial ro bustness bench- mark

    Francesco Croce, Maksym Andriushchenko, Vikash Sehwag , Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mit tal, and Matthias Hein. Robustbench: a standardized adversarial ro bustness bench- mark. In Joaquin Vanschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Processing Systems Track on Datasets a nd Bench- marks 1, ...

  7. [6]

    https://artificialintelligenceact.eu /the-act/

    Eu ai act, 2024. https://artificialintelligenceact.eu /the-act/

  8. [7]

    https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai

    Ethics guidelines for trustworthy ai, 2019. https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai

Show all 54 references
  1. [8]

    Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024

    Philipp Guldimann, Alexander Spiridonov, Robin Staab, Nikola Jovanović, Mark Vero, Velko Vechev, Anna Gueorguieva, Mislav Balunovi ć, Nikola Konstantinov, Pavol Bielik, Petar Tsankov, and Martin Vech ev. Compl-ai framework: A technical interpretation and llm benchmarkin g suit...

  2. [9]

    Measuring massive multita sk language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Ma ntas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multita sk language understanding. In 9th International Conference on Learning Representa- tions, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenRe...

  3. [10]

    Measuri ng mathe- matical problem solving with the MATH dataset

    Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Aro ra, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuri ng mathe- matical problem solving with the MATH dataset. In Joaquin Va nschoren and Sai-Kit Yeung, editors, Proceedings of the Neural Information Proce...

  4. [11]

    Weld, and Luke Zett lemoyer

    Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zett lemoyer. Trivi- aqa: A large scale distantly supervised challenge dataset f or reading com- prehension. In Regina Barzilay and Min-Yen Kan, editors, Proceedings of the 55th Annual Meeting of the Association for Computati...

  5. [12]

    Openassistant conversations - democra tizing large lan- guage model alignment

    Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotir is Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh D uc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glush kov, Ar- nav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Ng uyen, ...

  6. [13]

    Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024

    Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun S un. Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024

  7. [14]

    Andrew Liu, Hongjian Zhou, Yining Hua, Omid Rohanian, A nshul Thakur, Lei Clifton, and David A. Clifton. Large language models in t he clinic: A comprehensive benchmark, 2024

  8. [15]

    Glore: Evaluating logical reasoning of large langua ge models

    Hanmeng Liu, Zhiyang Teng, Ruoxi Ning, Jian Liu, Qiji Zh ou, and Yue Zhang. Glore: Evaluating logical reasoning of large langua ge models. CoRR, abs/2310.09107, 2023

  9. [16]

    Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning

    Zeyuan Ma, Hongshu Guo, Jiacheng Chen, Zhenrui Li, Guoj un Peng, Yue- Jiao Gong, Yining Ma, and Zhiguang Cao. Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Ha rdt,...

  10. [18]

    Abstractive text summarization u sing sequence- to-sequence rnns and beyond

    Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Sant os, Çaglar Gülçehre, and Bing Xiang. Abstractive text summarization u sing sequence- to-sequence rnns and beyond. In Yoav Goldberg and Stefan Rie zler, edi- tors, Proceedings of the 20th SIGNLL Conference on Computational ...

  11. [19]

    Adversarial NLI: A new benchmark for natura l language understanding

    Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, J ason Weston, and Douwe Kiela. Adversarial NLI: A new benchmark for natura l language understanding. In Dan Jurafsky, Joyce Chai, Natalie Schlut er, and Joel R. 23 Tetreault, editors, Proceedings of the 58th Annual Meeting...

  12. [20]

    https://oecd.ai/en/ai-pr inciples

    Ai system lifecycle, 2024. https://oecd.ai/en/ai-pr inciples

  13. [21]

    The LAMBADA dataset: Word prediction requir- ing a broad discourse context

    Denis Paperno, Germán Kruszewski, Angeliki Lazaridou , Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Ge mma Boleda, and Raquel Fernández. The LAMBADA dataset: Word prediction requir- ing a broad discourse context. In Proceedings of the 54th Annual Meeting o...

  14. [22]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Sa muel R. Bowman. GPQA: A graduate-level google-proof q&a benchmark . CoRR, abs/2311.12022, 2023

  15. [23]

    Winogrande: An adversarial winograd schema challenge at sc ale

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula , and Yejin Choi. Winogrande: An adversarial winograd schema challenge at sc ale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intellige n...

  16. [24]

    Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F

    Girish Sastry, Lennart Heim, Haydn Belfield, Markus And erljung, Miles Brundage, Julian Hazell, Cullen O’Keefe, Gillian K. Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F. Trager, Shahar A vin, Adrian Weller, Yoshua Bengio ,...

  17. [25]

    Sur- vey of different large language model architectures: Trends , benchmarks, and challenges

    Minghao Shao, Abdul Basit, Ramesh Karri, and Muhammad S hafique. Sur- vey of different large language model architectures: Trends , benchmarks, and challenges. IEEE Access, 2024

  18. [26]

    Concep tnet 5.5: An open multilingual graph of general knowledge

    Robyn Speer, Joshua Chin, and Catherine Havasi. Concep tnet 5.5: An open multilingual graph of general knowledge. In Satinder S ingh and Shaul Markovitch, editors, Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, February 4-9, 2017, San Francisco, Cal...

  19. [27]

    Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing

    Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, an d Greg Durrett. Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing. In The Twelfth International Conference on Learning Representat ions, ICLR 2024, Vienna, Austria, May 7-11, 2024 . OpenReview....

  20. [28]

    Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

    Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu, Yanshu Li, J iashuo Sun, Juntao Li, Min Zhang, and Yu Cheng. Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

  21. [29]

    A corpus for reasoning about natural language gr ounded in photographs, 2019

    Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Hua jun Bai, and Yoav Artzi. A corpus for reasoning about natural language gr ounded in photographs, 2019

  22. [30]

    Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024

    Yuan Sui, Mengyu Zhou, Mingjie Zhou, Shi Han, and Dongme i Zhang. Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024

  23. [31]

    Le, Ed H

    Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebas tian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H . Chi, Denny Zhou, and Jason Wei. Challenging big-bench tasks and whether chain-of-thought can solve them. In Anna Rogers, Jordan L. B oyd-Graber...

  24. [32]

    Commonsenseqa: A question answering challenge targeting c ommonsense knowledge

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jon athan Berant. Commonsenseqa: A question answering challenge targeting c ommonsense knowledge. In Jill Burstein, Christy Doran, and Thamar Solo rio, editors, Proceedings of the 2019 Conference of the North American Chapt er...

  25. [33]

    Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls

    Weizhi Tang and Vaishak Belle. Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls. CoRR, abs/2407.05434, 2024

  26. [34]

    Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C

    Jeyan Thiyagalingam, Gregor von Laszewski, Junqi Yin, Murali Emani, Juri Papay, Gregg Barrett, Piotr Luszczek, Aristeidis Tsar is, Christine R. Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C. Fox, and Tony Hey. AI benchmarking for sc i...

  27. [35]

    Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V

    Dennis Vetter, Julia Amann, Frédérick Bruneault, Mega n Coffee, Boris Düdder, Alessio Gallucci, Thomas Krendl Gilbert, Thilo Hag endorff, Irmhild van Halem, Eleanore Hickman, Elisabeth Hildt, Sune Holm, Geor- gios Kararigas, Pedro Kringen, Vince I. Madai, Emilie Wiinb lad Mathez...

  28. [36]

    Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, Kurt D

    Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed, Victor A kinwande, Namir Al-Nuaimi, Najla Alfaraj, Elie Alhajjar, Lora Aroyo, Trupti Bavalatti, Borhane Blili-Hamelin, Kurt D. Bollacker, Rish i Bomassani, Marisa Ferrara Boston, Siméon Campos, Kal Chakra, Canyu Che n, Cody Colema...

  29. [37]

    Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpre et Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superg lue: A stickier benchmark for general-purpose language understa nding systems, 2020

  30. [38]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill , Omer Levy, and Samuel R. Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019

  31. [39]

    Murdick, Devvret Ris hi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D

    Lucy Lu Wang, Kyle Lo, Yoganand Chandrasekhar, Russell Reas, Jiangjiang Yang, Darrin Eide, Kathryn Funk, Rodney Kinney, Ziyang Liu, William Merrill, Paul Mooney, Dewey A. Murdick, Devvret Ris hi, Jerry Sheehan, Zhihong Shen, Brandon Stilson, Alex D. Wade, Kuans an Wang, Chris ...

  32. [40]

    Mmlu-pro: A more robust and challenging multi-task la nguage un- derstanding benchmark

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhrani l Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jia ng, Tianle 26 Li, Max Ku, Kai Wang, Alex Zhuang, Rongqi Fan, Xiang Yue, and W enhu Chen. Mmlu-pro: A more robust and challenging multi-task la nguage...

  33. [41]

    CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models

    Zeyu Wang. CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models. In Kam-Fai Wong, Min Zhang, Ruifeng Xu, Jing Li, Zhongyu Wei, Lin Gui, Bin Lian g, and Runcong Zhao, editors, Proceedings of the 10th SIGHAN Workshop on Chi...

  34. [42]

    Logicv ista: Mul- timodal LLM logical reasoning benchmark in visual contexts

    Yijia Xiao, Edward Sun, Tianyu Liu, and Wei Wang. Logicv ista: Mul- timodal LLM logical reasoning benchmark in visual contexts . CoRR, abs/2407.04973, 2024

  35. [43]

    Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering

    Vikas Yadav, Steven Bethard, and Mihai Surdeanu. Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering. In Proceedings of the 2019 Conference on Empirical Meth- ods in Natural Language Processing and the 9th Internationa...

  36. [44]

    Eval- uating the quality of hallucination benchmarks for large vi sion-language models

    Bei Yan, Jie Zhang, Zheng Yuan, Shiguang Shan, and Xilin Chen. Eval- uating the quality of hallucination benchmarks for large vi sion-language models. CoRR, abs/2406.17115, 2024

  37. [45]

    https://z-inspection.org/

    Z-inspection, 2024. https://z-inspection.org/

  38. [46]

    Zebralogic: Benchmarking the logical reasoning abili ty of language models,

  39. [47]

    Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi , and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R. Traum, and Lluís Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguisti cs, A...

  40. [48]

    Benchmarking trustworthiness of multimodal large language models: A comprehensive study

    Yichi Zhang, Yao Huang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Yifan Wang, Huanran Chen, Xiao Yang, Xingxing Wei, Han g Su, Yinpeng Dong, and Jun Zhu. Benchmarking trustworthiness of multimodal large language models: A comprehensive study. CoRR, abs/2406.07057, 2024

  41. [49]

    Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024

    Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, and Xuming Hu. Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024. 27

  42. [50]

    Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024

    Yihang Zheng, Bo Li, Zhenghao Lin, Yi Luo, Xuanhe Zhou, C hen Lin, Jinsong Su, Guoliang Li, and Shifu Li. Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024

  43. [51]

    Instruction-followin g evaluation for large language models

    Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha B rahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. Instruction-followin g evaluation for large language models. CoRR, abs/2311.07911, 2023

  44. [52]

    Causalbench: A comprehensive benchmark for ca usal learn- ing capability of large language models

    Yu Zhou, Xingyu Wu, Beicheng Huang, Jibin Wu, Liang Feng , and Kay Chen Tan. Causalbench: A comprehensive benchmark for ca usal learn- ing capability of large language models. CoRR, abs/2404.06349, 2024

  45. [53]

    Roberto V. Zicari, John Brodersen, James Brusseau, Bor is Düdder, Timo Eichhorn, Todor Ivanov, Georgios Kararigas, Pedro Kringen , Melissa Mc- Cullough, Florian Möslein, Naveed Mushtaq, Gemma Roig, Nor man Stürtz, Karsten Tolle, Jesmin Jahan Tithi, Irmhild van Halem, and Ma gn...

  46. [2024]

    https://huggingface.co/blog/yuchenlin/zebra-lo gic

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.