Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2608.06202.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:36.904062Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
65 of 65 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26221a13-0976-4b7c-8ff5-babd1a283942 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2026 , note=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622a2c7a-90e9-427b-bd74-9178d043a17c · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 462ca983-a498-4f75-b75c-16294ff53ff8 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dee58e1d-f0a9-4e4e-8c85-9406f8f71f35 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 10th SIGHAN Workshop on Chinese Language Processing (SIGHAN-10) , pages=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d6aa5bc-2a07-47f9-8374-28c6e9e277b0 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71b7280e-9268-4942-ae0b-d9bd554abc3b · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60b402ce-04b2-4459-b5f6-e38590f0a660 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f4de7f-479c-4700-a2e9-6569fef31de5 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) European Semantic Web Conference , pages=
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c2e8861-ca7c-4252-972b-efc4ed7283d9 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions of the Association for Computational Linguistics , volume=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a50989-6931-4d0e-a9fa-d8b75f58de0a · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68e4a836-098f-4d8f-b437-4c400f68a726 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Evaluating Large Language Models Trained on Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70309c71-8603-499e-99da-07f30a0a32c8 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 conference of the North American chapter of the association for computational linguistics: human language technologies , pages=
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfb01546-9fc0-4fbd-97f4-7857da88a24e · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the association for computational linguistics: EMNLP 2020 , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7228e6f-6477-4bf1-98c1-836015154b6f · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d72a67a-258c-4fa6-8534-0578564f6b3c · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8d60a6d-d40b-46f5-8eb6-d5ffcf91416b · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) C row S -Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47d6d3e-4761-4dc4-a5f8-a6b66cf1012b · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on Machine Learning Research , doi =
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ef6a630-248e-4bea-aecf-078d979fbb64 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) ACM transactions on intelligent systems and technology , volume=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7776f9c7-d798-446f-a516-ba187f2bbd87 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Transactions on machine learning research , year=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83be2d93-8c2c-4244-b90c-003d12d088db · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Measuring Massive Multitask Language Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe0054e4-ac38-4184-9bbe-44a120d787cb · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: NAACL 2024 , pages=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086b3cad-130d-45e1-8776-16f61cd85221 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) P ro SA : Assessing and Understanding the Prompt Sensitivity of LLM s
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f92216a6-268c-4b31-bd92-8fc6e0a4b4bb · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e24bf35-9c64-4a7b-aa78-09c43b5ead4d · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis , pages=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe0a5809-6866-4f94-a79c-1da3448024a0 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Understanding User Experience in Large Language Model Interactions
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d46bb22a-b864-4e17-9d3f-89312334bcac · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8b092d3-0c98-4656-96af-edda0a1640e5 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c553b29-49b6-4375-aa24-bdf47c7ce632 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da414d8c-eecd-4044-b747-d975df546e44 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) arXiv preprint arXiv:2509.19364 , year=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3f3f2b-b54d-4663-a092-ccc87aea3b80 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a9935d0-f7b6-46d0-acd9-2b2a9276d937 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) 2023 , isbn =
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6161cf9b-f2c6-4a51-bb84-19e368a3ece4 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dccb30e-fa97-4f3c-9bda-fba0e3b7e762 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b012f5-87a5-48b6-bcab-0b05fd694512 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) S iren ' s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b346a3ab-f126-4e6a-bc5c-5e514e1e25f6 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f074f6-1ad5-452b-af55-9103cc4219d2 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1313c07f-0d7c-4d4e-9d24-d81a7fb31217 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a458f07a-0ffa-4f31-865f-231be2c48040 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , number=
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96133c27-3520-4a5a-9715-6591749ebe40 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Practices for
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 856bbd95-784c-4347-abaf-daea68706d34 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f43044d6-d349-4364-bbd8-fb8fb3bedbbb · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Aligning AI With Shared Human Values
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d834dcae-40d0-4cbc-95ab-648cf68eb808 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Toward an Evaluation Science for Generative AI Systems
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa23e92-20a0-4295-a933-fcdb476df63d · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Science , volume=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca37549-c4c5-45e5-970c-15f40d1808ec · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bcf96d0-68e2-4d3c-b056-ed6c588043bb · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b8b964-fc11-4886-a522-ddbc7f20ed3f · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09ebe1db-820d-48ac-9d25-7ef4f3d17a2b · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AI and the Everything in the Whole Wide World Benchmark
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3f7131e-5e38-46df-9f2a-7ea4c19805f0 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2025
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2808b874-9e30-4988-997d-b30941a981eb · outbound
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90fdb2a7-d6a1-4bdf-a364-b147057f1c14 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff3cc6b-80e4-4297-a894-ab5196b95f4d · outbound
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a18bd0c-da49-4965-981a-f5c1d621b604 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Sociotechnical Safety Evaluation of Generative AI Systems
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3825143f-fd23-4b60-8b37-122987b4325c · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Lessons from the Trenches on Reproducible Evaluation of Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 095f31e1-3024-4640-83ae-201405d6b4e2 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Advances in Neural Information Processing Systems , volume=
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b54a70-5013-4aab-b741-c30f055a0581 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 29th international conference on computational linguistics , pages=
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7ea2215-28e0-437e-82e8-bb84696ebfb7 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) International Conference on Learning Representations , volume=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e9c0f34-3323-44d4-89bd-3ce922ffa349 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Findings of the Association for Computational Linguistics: ACL 2024 , pages=
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7b4a3c1-3323-44c2-8021-4b673020cabd · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91789132-3150-4b20-8918-30ab698702b7 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Darrell, Trevor and Norouzi, Narges and Gonzalez, Joseph E
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 198004bf-5cd5-4244-8dbd-4afd50b28c74 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26b99a95-b552-483f-91de-7b8bd9d9cc69 · outbound
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a6ad17b-8612-40c6-a6b1-a94cab07c62f · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e89b97-aa78-4cdf-87ea-310d3df49387 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) and Metaxa, Dana
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af29f3f7-4e3a-4bc3-9c8a-b23f72efdb87 · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) Proceedings of the 2020 conference on fairness, accountability, and transparency , pages=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b748991-47bb-4794-885b-a35805ba49ce · outbound
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations) In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.