Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:06:55.228037Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 135 outbound references and 33 inbound Pith citation observations for arXiv:2502.06559.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T15:06:55.228037Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:34:50.154288Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 135 outbound references displayed
External citation measurements
9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 0c37d026-2980-4063-a91d-9dbb3bb3df9a · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Thompson
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 484125d1-2ae7-4df3-b436-b1d77ad2717d · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Field-building and the epistemic culture of AI safety
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c1b2e4c-07fe-4278-963b-01609ebcdb51 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 696c2d53-8a8d-462a-9d7b-1ce5de68e947 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c8b2ce17-cafa-4bcf-bf17-9e10d658e37d · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20338bb-4173-4f0b-91bf-111af29a152d · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9809ec0c-2c0c-4853-bdb8-0cdf0367ef5d · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Experiences from using snowballing and database searches in systematic literature studies
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa8d07b-660c-4e6a-b008-0f38dba07b34 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a11a392-f6f5-4e89-8127-98f453d3d61e · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking in Optimization: Best Practice and Open Issues
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb336076-1ff3-44af-9a59-a08018a608af · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Death of the Static AI Benchmark , March 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80194e49-4368-4eca-9465-d551e6fe5822 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0890d5f-d21e-442e-8ac1-12a9e451f0ac · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI auditing: The Broken Bus on the Road to AI Accountability
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e25ed8-e2f8-4f44-b8d4-c27049d49e12 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f686d32-1905-479c-84be-66c6b6fb57bd · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Making Intelligence : Ethical Values in IQ and ML Benchmarks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b5eb97-07a8-4e52-93ca-cd48646f8309 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e708afa-7fdb-4a90-9a59-1a46fa1da1d1 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bowman and George Dahl
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3accfd96-c13e-4391-b0c0-8cbb94169648 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking, pages 363--368
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4ac40191-ef53-4541-a50d-b0b9a7814651 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating AI Evaluation: Perils and Prospects
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565f89de-82ff-4744-8bc2-0a111fc168ee · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ullman, Fernando Martinez-Plumed, Joshua B
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c3289a9-82a7-4ffc-ba2a-02466e0b23d9 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8740b4-a76e-4116-ba6a-bb63733cbcbc · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation A Survey on Evaluation of Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fb02ee-dad9-4db2-856e-9344e71bd96e · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Cheng, CS
Reference 23
Source-reported events for the cited work
correction dated 2022-08-24. Source: crossref record 10.1038/s41928-022-00839-2->10.1038/s41928-022-00798-8:correction, observed 2026-07-11T03:08:33.550039+00:00. This notice travels one citation hop only.
Observation 1306bf06-d65b-4073-887c-e60cfb9a90b6 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the Measure of Intelligence
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a2d6b0f-b196-4539-aadc-2ef477525f55 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation A survey of 25 years of evaluation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 06c69184-38f8-4c73-b390-76531cac405f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Benchmark Lottery
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7472ef13-715b-4005-84b2-d59542854a89 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the genealogy of machine learning datasets: A critical history of ImageNet
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf73567-e96b-45ca-91cc-b9c5e206b3b1 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bringing the People Back In: Contesting Benchmark Machine Learning Datasets
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c76f87d3-f4c7-450c-9d02-f550463139cc · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0017ebb3-a924-48e8-80d4-3ef1c890c58b · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Utility is in the Eye of the User: A Critique of NLP Leaderboards
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99ca6c3-cee7-472b-9441-7c5e05461bba · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2ccb33-28c5-487f-b0b4-149dcab46d89 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58203012-b7fd-4191-9b28-d2b0b6967556 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16135675-89cf-4cb4-af1a-1b4ab4decb6d · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e8c7efa-28f3-4950-8a14-77141566b3a6 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68abdec4-5a97-4cbb-a2c0-d07552426a14 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Datasheets for Datasets
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de85b09-ff27-4466-a91e-f2fd5e4525cd · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feab772e-8ce3-47b6-9c30-53d848da43fc · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Shortcut Learning in Deep Neural Networks
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc15d71d-32bd-4810-b3cf-8efe834a4e33 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Are We Done with MMLU?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd55a78b-da1c-4179-9bdc-c9f90c07035b · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Diversity in artificial intelligence conferences
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9ba3440b-6225-41f8-a049-d9647f762fcf · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Alignment faking in large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0efd86a3-2cc2-4eef-ab17-b0890489515e · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae31a2e-06bf-4a6d-8d4a-841d6cc8b96f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ai hype is built on high test scores
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d51f0fb-7731-40cb-a0dd-ddf5555d5eed · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Measuring Massive Multitask Language Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d81ad73e-70a8-4b69-bf92-e3653a813d0a · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0219469c-622c-4529-b07d-cefb2c727659 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the Limitations of Compute Thresholds as a Governance Strategy
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db1e91e-0e60-4dac-bea1-0c5291fd1b00 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluation Gaps in Machine Learning Practice
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4dca48b-bf4e-40c1-bfd1-d618aafcf675 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Systematic literature studies: database searches vs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation dacddb1d-b479-49e0-8b32-7e5bcece080f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c62eaa-2b20-4415-a428-c74b139ceab0 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Constitution of Algorithms : Ground - Truthing , Programming , Formulating
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 750eb3c3-0ff1-4a11-bc02-9ee57655392f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Under the radar? examining the evaluation of foundation models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0477496-d9a3-4df8-8ce3-56e75c20fccb · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ground truth tracings ( GTT ): On the epistemic limits of machine learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 672fc10c-0bb0-4269-b1d3-0aab8d3a844c · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI Agents That Matter
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff0e720a-aff5-4ce8-a171-f858ce4d4746 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Leakage in data mining: Formulation, detection, and avoidance
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1a8f690-3432-4c7c-8e35-71513d7a23af · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Everyone Is Judging AI by These Tests
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf245071-b30c-481b-a6aa-c0c0d7a5245c · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mulvehill, and Deborah L
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6a2a399-860f-4bfa-83e6-33a1cddae108 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Feeling fixes: Mess and emotion in algorithmic audits
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5e517da3-25ce-4e17-800d-f7f0a2ebb021 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e1c6beb-794c-4e4e-add0-2cead3d7f5c8 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d63b08-f881-4019-832e-4f1cb0049438 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Metaethical Perspectives on 'Benchmarking' AI Ethics
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc32110d-779b-4610-9e2a-3f26bda585e5 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Questionable practices in machine learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20a4b3b-91a0-4c83-aa53-351a8cccc922 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Question and answer test-train overlap in open-domain question answering datasets
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e73f1a-be3b-4574-869d-ebff3a8d7aaa · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Holistic Evaluation of Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8bf805d-b04c-4813-a1ca-07f82e07f1f8 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Rethinking Model Evaluation as Narrowing the Socio-Technical Gap
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf0f3579-3aa1-47b0-8995-0582f3e832a5 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Are we learning yet? a meta review of evaluation failures across machine learning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c509233e-65dc-496a-8ebd-1f0d0023ed99 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation ExplainaBoard: An Explainable Leaderboard for NLP
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9afcc5e-d8e6-4f7d-bc63-44ebe7af9679 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI competitions as infrastructures of power in medical imaging
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9566273-dcf3-4b57-9649-4fa3175e583f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63271f95-0e13-45ce-b248-98bbf8216ed6 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Data contamination: From memorization to exploitation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c701896-0398-47f9-8beb-827cb86465bb · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b10177e0-a44d-42b7-8306-ec8676ad9faf · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Put to the test: For a new sociology of testing
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 46985aa9-a16f-4070-a1c6-89fdbd217fb9 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mlperf training benchmark
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f87e9b-4d73-437f-9f10-3e85c932022c · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mlperf: An industry standard benchmark suite for machine learning performance
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58a662b0-38db-4b24-8cf1-35d1af3a5c68 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4c125f-88e8-42c1-9346-47ca4767b047 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Frontier Models are Capable of In-context Scheming
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23568cc2-de9c-48f6-aee7-377fb84f4349 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation What Do NLP Researchers Believe? Results of the NLP Community Metasurvey
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f504545f-ffa8-40bd-880e-6f80fcbc1095 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking the Benchmarks
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ae171c1f-3071-4919-9e8e-09b4bf06203f · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation How do we know how smart AI systems are? Science, 381 0 (6654), July 2023
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fff97b18-4589-4e30-a1df-0d7823c72029 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation State of What Art? A Call for Multi-Prompt LLM Evaluation
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4428494f-5418-4eff-80bd-e7bd448a556a · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Proxies: The Cultural Work of Standing In
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468fafed-c739-400b-972d-171bf0681ca3 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a64299a0-a34b-4db1-86d4-fb60e0642331 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating LLMs is a minefield, 2023 b
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de525309-a9a4-45cd-9c6f-56b9e6af60b0 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Feder Cooper, Daphne Ippolito, Christopher A
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc0240f9-4c51-43b3-be97-98ed3201c4e5 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Scalable Extraction of Training Data from (Production) Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fab166-1b30-442f-875e-7a55e2d94f62 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac601b34-b7c9-4d6d-835d-4340cd3a77fa · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 680ee922-a081-4184-bcc1-ab836511d940 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56a2430-b014-4f9f-bff4-86196156ea69 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f285e6-ae51-43a5-a5ee-a27270e80710 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dbe9932-8ea6-4758-9406-82d93360e41a · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72053b20-d81c-4e13-869b-b3e84aa7b380 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mapping global dynamics of benchmark creation and saturation in artificial intelligence
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0564aa4-5cb6-4694-8454-4fc8f9af7fc4 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmark
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b56158-c7e9-4246-b35f-dd60e74dd4c9 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e183a41a-e920-47e9-8a48-c2532bcbb09c · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b321f8a4-831b-4f54-a9e0-8ea230480463 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bender, Emily Denton, and Alex Hanna
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a472054a-1414-4ad1-b0a1-652e7056d5b0 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b65cb9b1-d684-4395-a2ad-fc320d5d3399 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Testing - One , Two , Three ... Testing !
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 22c55187-dd1c-467d-8772-e84504bdcd3a · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Roles of English in Evaluating Multilingual Language Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b9c3d3e6-c9e5-4bb0-be05-456267fe790b · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4dc11da-f18e-4b8c-9f60-12df1bb61f7b · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bender, Alex Hanna, and Amandalynne Paullada
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 534bbdae-29cd-496d-8277-abbdc915d5f0 · outbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Gaps in the Safety Evaluation of Generative AI
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0820906a-b442-4c1b-96fa-127cc3db3aa2 · inbound
From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e3eb89-bf89-47cc-b678-bfc63a1c8353 · inbound
VLM@school -- Evaluation of AI image understanding on German middle school knowledge Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd68057e-87f1-4259-aafc-0a5e2ca81e00 · inbound
A Conceptual Framework for AI Capability Evaluations Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb8a087-d194-4022-99b5-ae7bfe110e16 · inbound
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06be70c0-2a6c-4b6d-9947-589ce3f0d3ab · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a3e686-dfa0-4a06-b22f-22a90cce61ab · inbound
Lilith: Developmental Modular LLMs with Chemical Signaling Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107adcf9-e699-4709-903b-29ddec1b7d1d · inbound
Deprecating Benchmarks: Criteria and Framework Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d9021e-42b6-4575-a7eb-8b70337a44f9 · inbound
Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90985508-fb8d-4c1d-8918-c2a5a016e410 · inbound
What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9b9a06-7551-48c2-b207-11eb412f0312 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0bb11615-6862-4f1c-8fdd-7e5f22ddc3e1 · inbound
VERA-MH Concept Paper Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 403e1f22-2bd4-4867-9dbb-ea5ae6e3947c · inbound
AI Consciousness and Existential Risk Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 85ada9e8-d58f-4c86-9f18-21b1d4774eda · inbound
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07cf78a3-31ca-40b9-9cc0-961e5a657f95 · inbound
MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 234dd87e-67d2-4a20-8fa5-f6a6429e8ff6 · inbound
From Human-Level AI Tales to AI Leveling Human Scales Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f67720ca-b7d7-4365-871d-f636c17c07a5 · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eb409c71-7bca-4f56-958b-84895575230e · inbound
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 15b4690f-8238-4d2f-a9e6-4ef43fdc3c52 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5ef79f9c-5b4b-4440-a325-701f8b083403 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e29ccdeb-27b1-4b9b-b1b4-d1073cd12667 · inbound
Simulating the Evolution of Alignment and Values in Machine Intelligence Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 952a38f4-bfb9-4293-b86f-f9f5c3d0001a · inbound
Computational Hermeneutics: Evaluating generative AI as a cultural technology Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e6a615d0-b24d-4a72-ba51-82f817ff86fc · inbound
To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2d70d903-9445-4603-a96e-4fc68942de23 · inbound
Dataset Watermarking for Closed LLMs with Provable Detection Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0226a078-f945-4fc6-8b3e-278471d5bfc0 · inbound
Unsteady Metrics and Benchmarking Cultures of AI Model Builders Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4b7af485-4a67-4466-9444-86cac6b58ff3 · inbound
Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 34ae6e45-61a6-4223-b38c-fccf5d80e4e5 · inbound
MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d4ea5a4e-f8cb-4247-a929-6a3ea6b00793 · inbound
ComplexConstraints and Beyond: Expert Rubrics for RLVR Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 121011bc-ae69-4a14-a5cb-0acd8e8475b0 · inbound
The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6390513f-e9d5-4895-89ed-b04f8dfc184f · inbound
A Technical Typology of AI Systems in Public Administration Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 282
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3ff06c53-3f47-4ec8-9924-1a675112e4dc · inbound
The Foreign Policy AI Evaluation Gap Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68874ed-8aa6-40c8-aee3-a439b3174a5d · inbound
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d59bbba-dbce-43eb-b893-e9613d902022 · inbound
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793a899a-75c2-4ab1-b017-0828169d0495 · inbound
Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.