Pith. sign in

Paper Citation Record · LEDGER

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

As of 23 August 2026, this Paper Citation Record lists 100 of 135 outbound references and 33 inbound Pith citation observations for arXiv:2502.06559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06559 v2

Coverage vector

measured 100 of 135 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:06:55.228037Z

measured 133 of 133 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:34:50.154288Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 135 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved77
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 0c37d026-2980-4063-a91d-9dbb3bb3df9a · outbound

This paper cites Thompson.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Thompson

Reference 1

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.827575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.711009Z digest=sha256:3a2846765eeee3ab368d2612da20e38b34202c22b4b1fd1254b124108cbea266

Observation 484125d1-2ae7-4df3-b436-b1d77ad2717d · outbound

This paper cites Field-building and the epistemic culture of AI safety.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Field-building and the epistemic culture of AI safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.717113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.717113Z digest=sha256:b433639b6fad1f005cbfb1ee10e585692293a4f50a91042cde4b26077f1c294a

Observation 6c1b2e4c-07fe-4278-963b-01609ebcdb51 · outbound

This paper cites When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.722646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.722646Z digest=sha256:8f500d60f5bea1224eb3f7bfdf516136636032ec0d723e16ab5922cebe54db36

Observation 696c2d53-8a8d-462a-9d7b-1ce5de68e947 · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 4

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.799348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.728474Z digest=sha256:cacaad740d2db47214b5fc5ea915cb9ad09c2280aad999cd492999e0ecf3d809

Observation c8b2ce17-cafa-4bcf-bf17-9e10d658e37d · outbound

This paper cites Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Truth Is a Lie : Crowd Truth and the Seven Myths of Human Annotation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.733654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.733654Z digest=sha256:ba31d36d44d15c40d78468998304d6dce0b2c4440139e7ece8b8f1b8475a7b95

Observation a20338bb-4173-4f0b-91bf-111af29a152d · outbound

This paper cites Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.739088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.739088Z digest=sha256:3893f2e3f7cfeeaa821e4ae7e64c922f245f4414736298bcb8f08718d029a6bb

Observation 9809ec0c-2c0c-4853-bdb8-0cdf0367ef5d · outbound

This paper cites Experiences from using snowballing and database searches in systematic literature studies.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Experiences from using snowballing and database searches in systematic literature studies

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.744859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.744859Z digest=sha256:455c60e2758e9fd90fba55b26b2ab1067edc498252bd448ef9f7c8704d7cec58

Observation daa8d07b-660c-4e6a-b008-0f38dba07b34 · outbound

This paper cites It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation It's COMPASlicated: The Messy Relationship between RAI Datasets and Algorithmic Fairness Benchmarks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.749660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.749660Z digest=sha256:e21bed71f20c92f935c59f8615f99e74b57f3aac0d13143c53327bb14e05f704

Observation 6a11a392-f6f5-4e89-8127-98f453d3d61e · outbound

This paper cites Benchmarking in Optimization: Best Practice and Open Issues.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking in Optimization: Best Practice and Open Issues

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.755116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.755116Z digest=sha256:c40912eb1c7b0bee236b353e1a129d55ef41393a5dbd5e6d0866cf30e88bd59b

Observation bb336076-1ff3-44af-9a59-a08018a608af · outbound

This paper cites The Death of the Static AI Benchmark , March 2024.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Death of the Static AI Benchmark , March 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.760557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.760557Z digest=sha256:435d8729c8b73d57caf773b5127d3b2c521c8c72f59434cb0dc751574bb99a65

Observation 80194e49-4368-4eca-9465-d551e6fe5822 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.765506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.765506Z digest=sha256:5393e259cdfb8ea6d64c4b06f609f2309677d1cc48724045c55434325f2fbc88

Observation a0890d5f-d21e-442e-8ac1-12a9e451f0ac · outbound

This paper cites AI auditing: The Broken Bus on the Road to AI Accountability.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI auditing: The Broken Bus on the Road to AI Accountability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.770662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.770662Z digest=sha256:7ea114e14cd1a7aa3223abc76c1b39225764af992adab7cf199b9acd4be0e827

Observation c3e25ed8-e2f8-4f44-b8d4-c27049d49e12 · outbound

This paper cites Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.775599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.775599Z digest=sha256:e13f4d5febcf1e4aa0bf8f7cfc967eaee0413f09f1bed32b42800667b4ba3481

Observation 7f686d32-1905-479c-84be-66c6b6fb57bd · outbound

This paper cites Making Intelligence : Ethical Values in IQ and ML Benchmarks.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Making Intelligence : Ethical Values in IQ and ML Benchmarks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.780076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.780076Z digest=sha256:a8d218780673192c0990f3cf5d2032de2e16db3732d8ca932e8126c8c30ea977

Observation 97b5eb97-07a8-4e52-93ca-cd48646f8309 · outbound

This paper cites Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Stereotyping Norwegian Salmon : An Inventory of Pitfalls in Fairness Benchmark Datasets

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.784923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.784923Z digest=sha256:6202771b48271811926d46545ef465a6efdeeac6e6508e8a4278957751f73efa

Observation 7e708afa-7fdb-4a90-9a59-1a46fa1da1d1 · outbound

This paper cites Bowman and George Dahl.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bowman and George Dahl

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.789694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.789694Z digest=sha256:c153966e6f7f1b9f627aa45a0f14cde8f0eca17514aaea3ed1e825a5d39640ec

Observation 3accfd96-c13e-4391-b0c0-8cbb94169648 · outbound

This paper cites Benchmarking, pages 363--368.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking, pages 363--368

Reference 17

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.750403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.794743Z digest=sha256:4b0f4705dee622cfdad57c01c153924c32bdd79ae9b50ca74105000fc03a765f

Observation 4ac40191-ef53-4541-a50d-b0b9a7814651 · outbound

This paper cites Evaluating AI Evaluation: Perils and Prospects.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating AI Evaluation: Perils and Prospects

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.799539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.799539Z digest=sha256:24543ca763bb93228cea9435e00baab3787bb200df7f596021b87d5ccbfe2b53

Observation 565f89de-82ff-4744-8bc2-0a111fc168ee · outbound

This paper cites Ullman, Fernando Martinez-Plumed, Joshua B.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ullman, Fernando Martinez-Plumed, Joshua B

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.804900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.804900Z digest=sha256:c872d40cdc4292ade7096bc31dccaea2740066ca783c2523105e3b44cdc4027c

Observation 9c3289a9-82a7-4ffc-ba2a-02466e0b23d9 · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.809515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.809515Z digest=sha256:d44297b756a8588a544c047d6731710c0bf566286af383f85e774749da950499

Observation dc8740b4-a76e-4116-ba6a-bb63733cbcbc · outbound

This paper cites A Survey on Evaluation of Large Language Models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation A Survey on Evaluation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.819434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.819434Z digest=sha256:4285900ee1ddc41b0e1da2ee82a2202ff082681ab2a4d692948341a86ef80efb

Observation 25fb02ee-dad9-4db2-856e-9344e71bd96e · outbound

This paper cites Cheng, CS.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Cheng, CS

Reference 23

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.721117Z

Source-reported events for the cited work

correction dated 2022-08-24. Source: crossref record 10.1038/s41928-022-00839-2->10.1038/s41928-022-00798-8:correction, observed 2026-07-11T03:08:33.550039+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-08-08T15:06:54.824738Z digest=sha256:721151e297d229f2fcca8e43db0b556c9f0303fb9d4f3d5cf3027401db6557fa

Observation 1306bf06-d65b-4073-887c-e60cfb9a90b6 · outbound

This paper cites On the Measure of Intelligence.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the Measure of Intelligence

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.829633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.829633Z digest=sha256:659772c2257572ce76240d9ddc0d90f2599989d4b9ee117227460383a2b02615

Observation 0a2d6b0f-b196-4539-aadc-2ef477525f55 · outbound

This paper cites A survey of 25 years of evaluation.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation A survey of 25 years of evaluation

Reference 25

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.704740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.836297Z digest=sha256:b0b58746056c20f21903504bcdebc53708baac1b165d7512d275cf2e56d59437

Observation 06c69184-38f8-4c73-b390-76531cac405f · outbound

This paper cites The Benchmark Lottery.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Benchmark Lottery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.842393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.842393Z digest=sha256:4b9f7523995e763c54a343b3eee693dc6fdd76eb3b4c915d6100794a9fd521de

Observation 7472ef13-715b-4005-84b2-d59542854a89 · outbound

This paper cites On the genealogy of machine learning datasets: A critical history of ImageNet.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the genealogy of machine learning datasets: A critical history of ImageNet

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.847430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.847430Z digest=sha256:80afc56189ed153f5aebe4c79bb58d067947d51f150a5272db2dee074a179b91

Observation bbf73567-e96b-45ca-91cc-b9c5e206b3b1 · outbound

This paper cites Bringing the People Back In: Contesting Benchmark Machine Learning Datasets.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bringing the People Back In: Contesting Benchmark Machine Learning Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.852188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.852188Z digest=sha256:6a7c167125c371e6be9feaefd4e9ac0dacc38085d14bece899238cf17b78b5e7

Observation c76f87d3-f4c7-450c-9d02-f550463139cc · outbound

This paper cites Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Agreements ‘in the wild’: Standards and alignment in machine learning benchmark dataset construction

Reference 30

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.677229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.861689Z digest=sha256:fc1224ae214014978d635b1019fe7576f22e2a22cf3594750888e5ac65f8ce07

Observation 0017ebb3-a924-48e8-80d4-3ef1c890c58b · outbound

This paper cites Utility is in the Eye of the User: A Critique of NLP Leaderboards.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Utility is in the Eye of the User: A Critique of NLP Leaderboards

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.866504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.866504Z digest=sha256:a9aace184b87520cd38c06aecf55b9dfe9e4c847a577d700d30b9ab47273f8c1

Observation e99ca6c3-cee7-472b-9441-7c5e05461bba · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.871752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.871752Z digest=sha256:13d6af7c2b664b894f63ca630d014330527bc793c34e36324e7e3ccf631db900

Observation 8c2ccb33-28c5-487f-b0b4-149dcab46d89 · outbound

This paper cites First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation First Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 a

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.876695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.876695Z digest=sha256:8eda864253563babdb0aa36b7f0e76549040b5282a1ff446b37949320348bb30

Observation 58203012-b7fd-4191-9b28-d2b0b6967556 · outbound

This paper cites Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Second Draft of the General-Purpose AI Code of Practice published, written by independent experts , 2024 b

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.881300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.881300Z digest=sha256:97e50a4ae4dd48300ffd606a0cfb8beb07caeed37b274df7c6570c1074128b11

Observation 16135675-89cf-4cb4-af1a-1b4ab4decb6d · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.885976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.885976Z digest=sha256:d3098d8e5d269fd092c85fea2823c0d2d06794f7eec45865aacbe251e5998b12

Observation 5e8c7efa-28f3-4950-8a14-77141566b3a6 · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.890605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.890605Z digest=sha256:a2a504924b56c04eded1c23ab42c406f46caf5d476eae167054372ca70b223fe

Observation 68abdec4-5a97-4cbb-a2c0-d07552426a14 · outbound

This paper cites Datasheets for Datasets.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Datasheets for Datasets

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.895262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.895262Z digest=sha256:d1c6fa6d36ad7e7cdd8f018ae70dfe994b8a9bd5ca37f3d262d25ca0107ce704

Observation 3de85b09-ff27-4466-a91e-f2fd5e4525cd · outbound

This paper cites Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Repairing the Cracked Foundation : A Survey of Obstacles in Evaluation Practices for Generated Text

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.900368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.900368Z digest=sha256:4ad8d975d7ce7c72a85ac32e8b4e66ec173dc3fb6c51995f7e8871b0645409c0

Observation feab772e-8ce3-47b6-9c30-53d848da43fc · outbound

This paper cites Shortcut Learning in Deep Neural Networks.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Shortcut Learning in Deep Neural Networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.905352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.905352Z digest=sha256:e6a41710b5c9c589042bbf091530a05f4e163ac44816210b56b483471f37519f

Observation bc15d71d-32bd-4810-b3cf-8efe834a4e33 · outbound

This paper cites Are We Done with MMLU?.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Are We Done with MMLU?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.910434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.910434Z digest=sha256:90628f3b71a68ff800a51382d4b3a3a3ea6387620737d3e3602dbe13b124a636

Observation cd55a78b-da1c-4179-9bdc-c9f90c07035b · outbound

This paper cites Diversity in artificial intelligence conferences.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Diversity in artificial intelligence conferences

Reference 41

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.649867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.916120Z digest=sha256:227a7a689ce656980ebc51710373a4cd149cdd405c8e8ace6c3853f392a65b8d

Observation 9ba3440b-6225-41f8-a049-d9647f762fcf · outbound

This paper cites Alignment faking in large language models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Alignment faking in large language models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.921165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.921165Z digest=sha256:c9aeb782ff762c46341e4aec615306eebdb1d4a9b23b41d87c29f5743e8aade4

Observation 0efd86a3-2cc2-4eef-ab17-b0890489515e · outbound

This paper cites COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.930734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.930734Z digest=sha256:ee914c87b6a91698eb05a165995b99311cfb0fbdd2bb8ad3cc1709bb4243b098

Observation 9ae31a2e-06bf-4a6d-8d4a-841d6cc8b96f · outbound

This paper cites Ai hype is built on high test scores.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ai hype is built on high test scores

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.935965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.935965Z digest=sha256:807b2f3fa1bd6261eaef1b009d20f7c6f297666ac9d3c0887d834a16c975c368

Observation 5d51f0fb-7731-40cb-a0dd-ddf5555d5eed · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Measuring Massive Multitask Language Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.941201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.941201Z digest=sha256:50e2c3eb8f238ec4838ca447081f89868721e2682bc7350964ec4c8c773dc247

Observation d81ad73e-70a8-4b69-bf92-e3653a813d0a · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.946419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.946419Z digest=sha256:09fc4751085c9ebafa94d2f59d3130b2dd1acb90d0b92af5c2bdf267423aa1bf

Observation 0219469c-622c-4529-b07d-cefb2c727659 · outbound

This paper cites On the Limitations of Compute Thresholds as a Governance Strategy.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation On the Limitations of Compute Thresholds as a Governance Strategy

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.951473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.951473Z digest=sha256:907804310ca43f8dc1988a6eeabd959c0bfdf3cb4e33eb2379e75368ed96d0cf

Observation 2db1e91e-0e60-4dac-bea1-0c5291fd1b00 · outbound

This paper cites Evaluation Gaps in Machine Learning Practice.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluation Gaps in Machine Learning Practice

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.957293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.957293Z digest=sha256:6e50d8f54518c07b5f627c72fef865ee710420d9e88bc42e7c95e6ba6eaa9197

Observation e4dca48b-bf4e-40c1-bfd1-d618aafcf675 · outbound

This paper cites Systematic literature studies: database searches vs.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Systematic literature studies: database searches vs

Reference 50

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T15:06:57.775230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.962237Z digest=sha256:46f6851ece6014eec6993d0df908b8f703f1b9519ec0c1d6c296ff3afb82b865

Observation dacddb1d-b479-49e0-8b32-7e5bcece080f · outbound

This paper cites Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Escaping the McNamara Fallacy : Toward More Impactful Recommender Systems Research

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.967056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.967056Z digest=sha256:c02236a446cdebf8312c48880ddbdb2cbbe2250e15030a557d12fc08c311cf42

Observation e9c62eaa-2b20-4415-a428-c74b139ceab0 · outbound

This paper cites The Constitution of Algorithms : Ground - Truthing , Programming , Formulating.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Constitution of Algorithms : Ground - Truthing , Programming , Formulating

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.971913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.971913Z digest=sha256:dd512483f705b1f06830a975fc2baf290657b36b8ae076efda26bdeda37d8540

Observation 750eb3c3-0ff1-4a11-bc02-9ee57655392f · outbound

This paper cites Under the radar? examining the evaluation of foundation models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Under the radar? examining the evaluation of foundation models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.976493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.976493Z digest=sha256:efee4e00e51b0cac245f80f56ab611be8b2b2eaa580eb1e56e02088ee2d20740

Observation b0477496-d9a3-4df8-8ce3-56e75c20fccb · outbound

This paper cites Ground truth tracings ( GTT ): On the epistemic limits of machine learning.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Ground truth tracings ( GTT ): On the epistemic limits of machine learning

Reference 54

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.613881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.981043Z digest=sha256:36515ab03ef8173cc8b76c5d3a7f90d68e240e8e9a00627a16fb1cd4e7f909cd

Observation 672fc10c-0bb0-4269-b1d3-0aab8d3a844c · outbound

This paper cites AI Agents That Matter.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI Agents That Matter

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.985509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.985509Z digest=sha256:757288f3a4339236e645dc49f1c88438a322826a8b00bd0ab827e319087338a5

Observation ff0e720a-aff5-4ce8-a171-f858ce4d4746 · outbound

This paper cites Leakage in data mining: Formulation, detection, and avoidance.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Leakage in data mining: Formulation, detection, and avoidance

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.990287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.990287Z digest=sha256:d776612fca26c3d62c9f91fc8256b489e3cea8cf688b03c0a116dc324ae0f786

Observation b1a8f690-3432-4c7c-8e35-71513d7a23af · outbound

This paper cites Everyone Is Judging AI by These Tests.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Everyone Is Judging AI by These Tests

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:54.995109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:54.995109Z digest=sha256:b85c9861cebb6368da034b356f2d93d4e138b78edd63abb413ec49fc028e5798

Observation cf245071-b30c-481b-a6aa-c0c0d7a5245c · outbound

This paper cites Mulvehill, and Deborah L.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mulvehill, and Deborah L

Reference 58

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.598819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:54.999946Z digest=sha256:aa1ab4020dd13a441058832d438db3b33136cbc39613101edd90d246b100e9fa

Observation a6a2a399-860f-4bfa-83e6-33a1cddae108 · outbound

This paper cites Feeling fixes: Mess and emotion in algorithmic audits.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Feeling fixes: Mess and emotion in algorithmic audits

Reference 59

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.582988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.005003Z digest=sha256:c232ce469206c744e535bdce0e2e55c32a213f366c78e3e4ec692dca34fa2a2d

Observation 5e517da3-25ce-4e17-800d-f7f0a2ebb021 · outbound

This paper cites Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.009989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.009989Z digest=sha256:2c0b1425824f57ffa5bf343e40b029a7f56beebc81c5d09eccb475311db832dc

Observation 2e1c6beb-794c-4e4e-add0-2cead3d7f5c8 · outbound

This paper cites From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation From Protoscience to Epistemic Monoculture: How Benchmarking Set the Stage for the Deep Learning Revolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.015239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.015239Z digest=sha256:3128123230ebd94d9d4aae6717bd4192eaa1ea16f148517ae883a5c6965327ab

Observation 93d63b08-f881-4019-832e-4f1cb0049438 · outbound

This paper cites Metaethical Perspectives on 'Benchmarking' AI Ethics.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Metaethical Perspectives on 'Benchmarking' AI Ethics

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:06:57.560951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.020829Z digest=sha256:f1356e3d8f99a21d4a3e21491ce455446915a66341220967443eb80959ba6a14

Observation bc32110d-779b-4610-9e2a-3f26bda585e5 · outbound

This paper cites Questionable practices in machine learning.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Questionable practices in machine learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.026059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.026059Z digest=sha256:d075eb1da7e61ec09259566eb21c391beadef51655fa4ef3640af9d18c9a5b3d

Observation a20a4b3b-91a0-4c83-aa53-351a8cccc922 · outbound

This paper cites Question and answer test-train overlap in open-domain question answering datasets.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Question and answer test-train overlap in open-domain question answering datasets

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.031305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.031305Z digest=sha256:0c3ceabca061312eae6609b70caaf48668d74cf71826c6698d5dfd5512aaa8bc

Observation f7e73f1a-be3b-4574-869d-ebff3a8d7aaa · outbound

This paper cites Holistic Evaluation of Language Models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Holistic Evaluation of Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.036681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.036681Z digest=sha256:cef13d3c0bdd2f09893fdf6d1bcf3c0fc5be7f0356d93975f76c5ef91ae8e582

Observation b8bf805d-b04c-4813-a1ca-07f82e07f1f8 · outbound

This paper cites Rethinking Model Evaluation as Narrowing the Socio-Technical Gap.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.041880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.041880Z digest=sha256:f88623969023f25d2e30107cc4acbf8cf046a4c84d25c4bf8bfd3c15d8b9e8d4

Observation cf0f3579-3aa1-47b0-8995-0582f3e832a5 · outbound

This paper cites Are we learning yet? a meta review of evaluation failures across machine learning.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Are we learning yet? a meta review of evaluation failures across machine learning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.046898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.046898Z digest=sha256:27467e11685a2efdf3913f57c9e8718d57fadf8d5baccda6a6e1bcfe7df57ae9

Observation c509233e-65dc-496a-8ebd-1f0d0023ed99 · outbound

This paper cites ExplainaBoard: An Explainable Leaderboard for NLP.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation ExplainaBoard: An Explainable Leaderboard for NLP

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:06:57.487779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.051780Z digest=sha256:3e0430a5a89d5d46c175ca24712ea8f7c37cd35b22269c05f94ab09676bd2825

Observation d9afcc5e-d8e6-4f7d-bc63-44ebe7af9679 · outbound

This paper cites AI competitions as infrastructures of power in medical imaging.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation AI competitions as infrastructures of power in medical imaging

Reference 69

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T15:06:57.464372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.056702Z digest=sha256:fb5d8bbf1bcb91fc8d0b19c5a66077c3192e9ff257d3f07a6b81da7cbde9479b

Observation d9566273-dcf3-4b57-9649-4fa3175e583f · outbound

This paper cites Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bias in Language Models: Beyond Trick Tests and Toward RUTEd Evaluation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.061220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.061220Z digest=sha256:9ebcf858891f40c6a28eddaa6045963f35373acd862de60950228290bc96120c

Observation 63271f95-0e13-45ce-b248-98bbf8216ed6 · outbound

This paper cites Data contamination: From memorization to exploitation.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Data contamination: From memorization to exploitation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.066279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.066279Z digest=sha256:c52471c2cb1957a2e83c4f3ea6b0a7ed3586e669d0c02202baed8d3f87ec15ca

Observation 2c701896-0398-47f9-8beb-827cb86465bb · outbound

This paper cites Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Practices of Benchmarking : Vulnerability in the Computer Vision Pipeline

Reference 72

Resolution
verified exact
raw_fallback, observed 2026-08-08T15:06:57.371079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.071125Z digest=sha256:6287d135fbe816172b51ba8b30efacb3d6c89724fe84cb08a9d00fea957ea176

Observation b10177e0-a44d-42b7-8306-ec8676ad9faf · outbound

This paper cites Put to the test: For a new sociology of testing.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Put to the test: For a new sociology of testing

Reference 73

Resolution
verified exact
raw_fallback, observed 2026-08-08T15:06:57.246799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.076699Z digest=sha256:767f9d060b893fe779d22c3331fb1495c268a87cefe0866b234deb8db53bd9b1

Observation 46985aa9-a16f-4070-a1c6-89fdbd217fb9 · outbound

This paper cites Mlperf training benchmark.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mlperf training benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.081262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.081262Z digest=sha256:1bdb65df26d42cf9d97fa807d960a7e807c8d12423b25edf000010346fcfa65c

Observation c9f87e9b-4d73-437f-9f10-3e85c932022c · outbound

This paper cites Mlperf: An industry standard benchmark suite for machine learning performance.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mlperf: An industry standard benchmark suite for machine learning performance

Reference 75

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T15:06:57.158956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.085797Z digest=sha256:2b5ad3cbc64717f056f97ea98ccffa75071c06c02085b4a800fe1b54c9175fd3

Observation 58a662b0-38db-4b24-8cf1-35d1af3a5c68 · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.090364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.090364Z digest=sha256:5f562fa78cca339b1d6845ad0260cebffd0a5571f8762187d4cabb649b76c060

Observation 2a4c125f-88e8-42c1-9346-47ca4767b047 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Frontier Models are Capable of In-context Scheming

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.095295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.095295Z digest=sha256:a908851c0cf2896806717d87e5620dd92d24fdf24380ed9f4a91928e59aae9fb

Observation 23568cc2-de9c-48f6-aee7-377fb84f4349 · outbound

This paper cites What Do NLP Researchers Believe? Results of the NLP Community Metasurvey.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation What Do NLP Researchers Believe? Results of the NLP Community Metasurvey

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.100210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.100210Z digest=sha256:96b353d5e94696bfd94e8347ee93137d6fbbc45efe19005e2514775cf8b45b12

Observation f504545f-ffa8-40bd-880e-6f80fcbc1095 · outbound

This paper cites Benchmarking the Benchmarks.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmarking the Benchmarks

Reference 79

Resolution
metadata mismatch
raw_fallback, observed 2026-08-08T15:06:57.028571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.104848Z digest=sha256:0747345016e138daff40b051f744ebaff8e4edb7c4b76e18bcd63760a073a440

Observation ae171c1f-3071-4919-9e8e-09b4bf06203f · outbound

This paper cites How do we know how smart AI systems are? Science, 381 0 (6654), July 2023.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation How do we know how smart AI systems are? Science, 381 0 (6654), July 2023

Reference 80

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.546298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.109305Z digest=sha256:7c4fb7851e11c9ecf632817f79e3157b171314b30c3519d9329eb82e771a8968

Observation fff97b18-4589-4e30-a1df-0d7823c72029 · outbound

This paper cites State of What Art? A Call for Multi-Prompt LLM Evaluation.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation State of What Art? A Call for Multi-Prompt LLM Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.118860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.118860Z digest=sha256:2d5014c7da6ebaaff5c0e5a539e9d8d140ab9e514e4bb881903476effdd9fc28

Observation 4428494f-5418-4eff-80bd-e7bd448a556a · outbound

This paper cites Proxies: The Cultural Work of Standing In.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Proxies: The Cultural Work of Standing In

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.123748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.123748Z digest=sha256:54cf19b6c06700a7edfce1901c005ed3144e8642ce7648a768359e71634cc27a

Observation 468fafed-c739-400b-972d-171bf0681ca3 · outbound

This paper cites GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation GPT -4 and professional benchmarks: the wrong answer to the wrong question, March 2023 a

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.128502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.128502Z digest=sha256:7c38245458b2045a83f3fca22da0b663aeb263c5cb96310a0f635a453394f3f5

Observation a64299a0-a34b-4db1-86d4-fb60e0642331 · outbound

This paper cites Evaluating LLMs is a minefield, 2023 b.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating LLMs is a minefield, 2023 b

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.133342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.133342Z digest=sha256:c81a4da40de63d161da4980d45e07db1d3b2633035e71293dea0e8e742d4003b

Observation de525309-a9a4-45cd-9c6f-56b9e6af60b0 · outbound

This paper cites Feder Cooper, Daphne Ippolito, Christopher A.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Feder Cooper, Daphne Ippolito, Christopher A

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.138308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.138308Z digest=sha256:47b1e8a7a6098a5a57049ad73ffd22ceda0293335dd56afa70a1c79ac1e520bd

Observation cc0240f9-4c51-43b3-be97-98ed3201c4e5 · outbound

This paper cites Scalable Extraction of Training Data from (Production) Language Models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Scalable Extraction of Training Data from (Production) Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.143422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.143422Z digest=sha256:4e064f6f211676cbd1b0a6b85c1a7330176dc84dfff5cfab477d7433f095aef2

Observation 58fab166-1b30-442f-875e-7a55e2d94f62 · outbound

This paper cites Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Evaluating Hosting Provider Security Through Abuse Data and the Creation of Metrics

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.149163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.149163Z digest=sha256:07a3761ab8fd616ea641d24e35ea048f13215d22e0c2172005ad8b0de18f82ce

Observation ac601b34-b7c9-4d6d-835d-4340cd3a77fa · outbound

This paper cites Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Hidden Stratification Causes Clinically Meaningful Failures in Machine Learning for Medical Imaging

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.154615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.154615Z digest=sha256:0fc648e6d9e013bec3da1912013107b6a3797cac9d2cb4bafda656da22a110e1

Observation 680ee922-a081-4184-bcc1-ab836511d940 · outbound

This paper cites Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.159834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.159834Z digest=sha256:77b7ab6bd3365d2a1869e674c97ff31f7d4584fd68c62ebc3120a10fc9b31208

Observation f56a2430-b014-4f9f-bff4-86196156ea69 · outbound

This paper cites The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The social construction of datasets: On the practices, processes, and challenges of dataset creation for machine learning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.164878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.164878Z digest=sha256:7b53e93e94eb7b72e2210de72493fcb60c232ad4feca4e5b056f35db64e80f11

Observation 00f285e6-ae51-43a5-a5ee-a27270e80710 · outbound

This paper cites Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Building Better Datasets: Seven Recommendations for Responsible Design from Dataset Creators

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.169494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.169494Z digest=sha256:f4b6e131c327a6577bd0b84da4d4e2c313f6086277f49d71bac22352fd807a80

Observation 5dbe9932-8ea6-4758-9406-82d93360e41a · outbound

This paper cites an unresolved cited work.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.174479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.174479Z digest=sha256:b15cae2a58b56a0b2dcb8c915d86b6e3d01fc1c8bf2527333d3f184d9c09e237

Observation 72053b20-d81c-4e13-869b-b3e84aa7b380 · outbound

This paper cites Mapping global dynamics of benchmark creation and saturation in artificial intelligence.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.178850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.178850Z digest=sha256:fedee00e658be3a24f8a7243128324d89c502f16263bbe17df2c8c226f6756da

Observation e0564aa4-5cb6-4694-8454-4fc8f9af7fc4 · outbound

This paper cites Benchmark.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Benchmark

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.183830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.183830Z digest=sha256:2f5ecc7210754fa0191b838f3ae0662c68c100a26691fbee7009fff233e14a5c

Observation d7b56158-c7e9-4246-b35f-dd60e74dd4c9 · outbound

This paper cites Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers

Reference 96

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T15:06:56.775567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.188523Z digest=sha256:287d1097c309401359bacd62afbc698169fe486af67c9801fca1b8294bf8a63d

Observation e183a41a-e920-47e9-8a48-c2532bcbb09c · outbound

This paper cites Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Raison d’être of the benchmark dataset: A Survey of Current Practices of Benchmark Dataset Sharing Platforms

Reference 97

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.518509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.193886Z digest=sha256:e2372ca3c8e9828510e01a7f0adfa56eed07938bd715611581e68da1562c8090

Observation b321f8a4-831b-4f54-a9e0-8ea230480463 · outbound

This paper cites Bender, Emily Denton, and Alex Hanna.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bender, Emily Denton, and Alex Hanna

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.198904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.198904Z digest=sha256:4f8924102b505f6eece72982696e673f04e777d200260e610ccd96aea2dc3791

Observation a472054a-1414-4ad1-b0a1-652e7056d5b0 · outbound

This paper cites Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.203597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.203597Z digest=sha256:71810252d3b96fc46c2727fac0495b61139e5eb7ac723b851212aa82e99ca2f4

Observation b65cb9b1-d684-4395-a2ad-fc320d5d3399 · outbound

This paper cites Testing - One , Two , Three ... Testing !.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Testing - One , Two , Three ... Testing !

Reference 100

Resolution
verified exact
doi, observed 2026-08-08T15:06:55.501858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.208380Z digest=sha256:7ee5296e061b1d63d0ed4098b634591f82670ee596cf2e7cd02ce5b5c1ae07b3

Observation 22c55187-dd1c-467d-8772-e84504bdcd3a · outbound

This paper cites The Roles of English in Evaluating Multilingual Language Models.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation The Roles of English in Evaluating Multilingual Language Models

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:06:56.626299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T15:06:55.213500Z digest=sha256:d6771a0e617f5d4587a6cb0986712a09e66f5c45e65284e7b040f5c99a1dcdf0

Observation b9c3d3e6-c9e5-4bb0-be05-456267fe790b · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.218219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.218219Z digest=sha256:7085be73b796f1b90b0cfeeafbff0a8d804f7ec3c6886eb86e6fc84e73336580

Observation c4dc11da-f18e-4b8c-9f60-12df1bb61f7b · outbound

This paper cites Bender, Alex Hanna, and Amandalynne Paullada.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Bender, Alex Hanna, and Amandalynne Paullada

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.223108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.223108Z digest=sha256:bf9c7bd9fde80d92c453c0097644f720f485f35fdc4f5805b36d0bd80204236e

Observation 534bbdae-29cd-496d-8277-abbdc915d5f0 · outbound

This paper cites Gaps in the Safety Evaluation of Generative AI.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Gaps in the Safety Evaluation of Generative AI

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.228037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.228037Z digest=sha256:b581012dca95e8091517e31b90e46d98c8d9dbc2da64fc2c6e368dd333e0812d

Pith citing papers

Observation 0820906a-b442-4c1b-96fa-127cc3db3aa2 · inbound

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation cites this paper.

From Guidelines to Practice: A New Paradigm for Arabic Language Model Evaluation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:34:50.154288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:34:50.154288Z digest=sha256:7e63c89b22dd9b56bc20108d43e096b29fa6a1e59a96c0ac97b9842db1c5aca8

Observation 06e3eb89-bf89-47cc-b678-bfc63a1c8353 · inbound

VLM@school -- Evaluation of AI image understanding on German middle school knowledge cites this paper.

VLM@school -- Evaluation of AI image understanding on German middle school knowledge Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:06:28.735359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:06:28.735359Z digest=sha256:681d11d151e0af393aa712145cb492ae18707e501c277f9a3e9a82a75e6bbfd3

Observation bd68057e-87f1-4259-aafc-0a5e2ca81e00 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:46.923418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:46.923418Z digest=sha256:f0304e7fb0a8dc32417f9f7673569900978a0e122411ffee44622fd3119a5f5c

Observation bfb8a087-d194-4022-99b5-ae7bfe110e16 · inbound

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments cites this paper.

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.187973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.187973Z digest=sha256:e343f9b39147b102fa89f11a91034f7d3ca8bc2cf2c844c29b4d44b8b1e2d4bd

Observation 06be70c0-2a6c-4b6d-9947-589ce3f0d3ab · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:22.007624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:22.007624Z digest=sha256:708557a8f7a267f378518f564c88b2729cb4d9aea8afc35595b5cef34bf79233

Observation b4a3e686-dfa0-4a06-b22f-22a90cce61ab · inbound

Lilith: Developmental Modular LLMs with Chemical Signaling cites this paper.

Lilith: Developmental Modular LLMs with Chemical Signaling Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:48:44.495995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:48:44.495995Z digest=sha256:ce7dd5be467d89acebe34c5d1836d889c19d5708d4947d43bd99fcd7099945aa

Observation 107adcf9-e699-4709-903b-29ddec1b7d1d · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.423706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.423706Z digest=sha256:472287a18fdc7617c6bb1d4799a5dac785ff1a26af0574701860beceba419565

Observation 09d9021e-42b6-4575-a7eb-8b70337a44f9 · inbound

Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation cites this paper.

Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:38.308630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:38.308630Z digest=sha256:1c63dfc8d79feba852d7f7a6a00d6117f8a09e5d201c097b074c00cf1c4ceedd

Observation 90985508-fb8d-4c1d-8918-c2a5a016e410 · inbound

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles cites this paper.

What to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:42.978782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:42.978782Z digest=sha256:4526faee7d95b83b4a920cc1858c1b4a6cc6d5e87ba1e1e45d6b4bf4e3b61dd1

Observation 1f9b9a06-7551-48c2-b207-11eb412f0312 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.457894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:dac1572dacd02d6e2c2886cf2712bb2f99022c2d0e8a418c63380defc42051b2

Observation 0bb11615-6862-4f1c-8fdd-7e5f22ddc3e1 · inbound

VERA-MH Concept Paper cites this paper.

VERA-MH Concept Paper Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:01:01.475860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T07:00:15.411310Z digest=sha256:5a96329e9a2b227c3626c3fac347f2183a3db883b3af499ad6b8a12d9efea337

Observation 403e1f22-2bd4-4867-9dbb-ea5ae6e3947c · inbound

AI Consciousness and Existential Risk cites this paper.

AI Consciousness and Existential Risk Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:30:29.085892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T18:26:03.464658Z digest=sha256:3d3100a2ea3b346f3dcd18f07f6d36a94f532b6fe5c3745fd30b24c5bfa5128f

Observation 85ada9e8-d58f-4c86-9f18-21b1d4774eda · inbound

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation cites this paper.

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:24:17.550034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:24:17.550034Z digest=sha256:08390602256d646b4f3f3ea1d369867bf63cd5bb1fac2679c5b53c99dadcb44e

Observation 07cf78a3-31ca-40b9-9cc0-961e5a657f95 · inbound

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents cites this paper.

MoralityGym: A Benchmark for Evaluating Hierarchical Moral Alignment in Sequential Decision-Making Agents Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:51:25.927108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T10:49:35.846590Z digest=sha256:be8fa6ba9d467949a2d31d8292a303859a53e9b54ffbce8fe9e3d41bef58feba

Observation 234dd87e-67d2-4a20-8fa5-f6a6429e8ff6 · inbound

From Human-Level AI Tales to AI Leveling Human Scales cites this paper.

From Human-Level AI Tales to AI Leveling Human Scales Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:10:18.105298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:08:29.604466Z digest=sha256:5c7b2c053fd48e3abca18b8d1fe1fa2aae7b7843697d71009fe58c43eb7ca8b1

Observation f67720ca-b7d7-4365-871d-f636c17c07a5 · inbound

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI cites this paper.

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:21:23.982835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T10:19:56.003219Z digest=sha256:c33b92d5cf15970bf18f3cf16188381b1911620701fcc4774ccd64f905c4f3dd

Observation eb409c71-7bca-4f56-958b-84895575230e · inbound

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics cites this paper.

RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:28:21.220651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T22:25:35.345087Z digest=sha256:34fb0a033167aa2012b3499af5c98355617137a062c55d72f613bd062806f7e8

Observation 15b4690f-8238-4d2f-a9e6-4ef43fdc3c52 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:20.101226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:c7865bc17771fec33fd08f9826ef6b49ef383a0688f224ac8cdcfb0eaa4953f8

Observation 5ef79f9c-5b4b-4440-a325-701f8b083403 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.798502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:1402d08231e3db1bbb8b1b4779ad6c307f4b632c83e95b597c1e3335a994948c

Observation e29ccdeb-27b1-4b9b-b1b4-d1073cd12667 · inbound

Simulating the Evolution of Alignment and Values in Machine Intelligence cites this paper.

Simulating the Evolution of Alignment and Values in Machine Intelligence Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:48.445792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T20:15:46.311347Z digest=sha256:4e4dfd464b0cc748fa98674b13d8e1e9e17471a34c01cb26a187ca59616cbc0c

Observation 952a38f4-bfb9-4293-b86f-f9f5c3d0001a · inbound

Computational Hermeneutics: Evaluating generative AI as a cultural technology cites this paper.

Computational Hermeneutics: Evaluating generative AI as a cultural technology Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:48:27.153713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T23:48:24.354896Z digest=sha256:e38596bb8cdad05309099d1a0281592453316096d05f344268dbb050d56f8ca3

Observation e6a615d0-b24d-4a72-ba51-82f817ff86fc · inbound

To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems cites this paper.

To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:55:11.735752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T06:44:14.093749Z digest=sha256:385fc13b30955075c5ee6b46cfde0b17c24c747cf6a2b93e2a6c03068a966c7a

Observation 2d70d903-9445-4603-a96e-4fc68942de23 · inbound

Dataset Watermarking for Closed LLMs with Provable Detection cites this paper.

Dataset Watermarking for Closed LLMs with Provable Detection Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:57.214146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T00:53:42.185498Z digest=sha256:91884a3b95fd8927bd7cb6510e6c014fd1cfd958ef1fab5fdeb520b913d96c66

Observation 0226a078-f945-4fc6-8b3e-278471d5bfc0 · inbound

Unsteady Metrics and Benchmarking Cultures of AI Model Builders cites this paper.

Unsteady Metrics and Benchmarking Cultures of AI Model Builders Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:03.336380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T04:54:26.888562Z digest=sha256:a51c1723335e98e198d4035358c52c07f85cf809812be2e4b071fb19d4f2dc7c

Observation 4b7af485-4a67-4466-9444-86cac6b58ff3 · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:52:49.468378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:7c3a693377f9eed0c1027357c95e2956bd2beb058f7199d7f0fa4e1c2e373253

Observation 34ae6e45-61a6-4223-b38c-fccf5d80e4e5 · inbound

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models cites this paper.

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.308800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T20:02:50.169589Z digest=sha256:1d8c173970647e44722707f43714437c57c67d9c39a4da8bfdc2684c2438c9e6

Observation d4ea5a4e-f8cb-4247-a929-6a3ea6b00793 · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:31.360048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:fd2984ede3c97bcf96024c8f363cf5e4cad9a813c4d946c233ef2fa19e38640f

Observation 121011bc-ae69-4a14-a5cb-0acd8e8475b0 · inbound

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act cites this paper.

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.483341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T22:16:25.919667Z digest=sha256:296d6d302904ec76c2306b4c2de04c5dd639cc492249e13ab4bd5fae0694ab2c

Observation 6390513f-e9d5-4895-89ed-b04f8dfc184f · inbound

A Technical Typology of AI Systems in Public Administration cites this paper.

A Technical Typology of AI Systems in Public Administration Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-07-01T02:45:17.178772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-01T02:44:10.129474Z digest=sha256:2b5a1854dac1563e6a43b3aebdc2f75754244c2f09d1fc1ca366b07a6adb9667

Observation 3ff06c53-3f47-4ec8-9924-1a675112e4dc · inbound

The Foreign Policy AI Evaluation Gap cites this paper.

The Foreign Policy AI Evaluation Gap Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T05:51:08.101805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:51:08.101805Z digest=sha256:6fa9166e7a72d18dabc446420f83dcf545b305a126a2792eb6adf86e72807165

Observation d68874ed-8aa6-40c8-aee3-a439b3174a5d · inbound

Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model cites this paper.

Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T20:23:02.816523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T20:23:02.816523Z digest=sha256:5ed9132330c671258474b18e8752b56addc9d5aa66358d0c5627b6f3e1983d78

Observation 3d59bbba-dbce-43eb-b893-e9613d902022 · inbound

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems cites this paper.

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:44:37.517694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:44:37.517694Z digest=sha256:369f1226d69c33133cff2c1536b4be0e4cdea73b6bace9bdfe6c15a9ca2cefc1

Observation 793a899a-75c2-4ab1-b017-0828169d0495 · inbound

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems cites this paper.

Multi-Dimensional Assessment for AI Cognition (MAAC): A Theoretical Framework for Process-Oriented Cognitive Evaluation of Text-Based AI Systems Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T00:44:45.631232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:44:45.631232Z digest=sha256:00de87ab33d2c9781da6141cc8c576c149ac085bd4c139ca5b0fee42b5152c7b