Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:49:34.499531Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2412.01020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:49:34.499531Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:56.287352Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:28:44.552688Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 959064f1-1cef-4252-906d-e03d12c2d360 · outbound
AI Benchmarks and Datasets for LLM Evaluation https://aisafetybulgaria.c om/
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 768576e5-db86-4bd5-b740-1828ab195762 · outbound
AI Benchmarks and Datasets for LLM Evaluation Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 595bf9b7-49cb-428f-8c43-23511f8d17b0 · outbound
AI Benchmarks and Datasets for LLM Evaluation https://compl-ai.org/
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b852ae81-e737-4721-831a-9bfe82e099b8 · outbound
AI Benchmarks and Datasets for LLM Evaluation Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5a99daa-ed8a-4c81-b982-79c205a95d25 · outbound
AI Benchmarks and Datasets for LLM Evaluation Robustbench: a standardized adversarial ro bustness bench- mark
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 75850ed8-80ce-4434-8daf-38b73a9c410e · outbound
AI Benchmarks and Datasets for LLM Evaluation https://artificialintelligenceact.eu /the-act/
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b267cce2-1bd9-4cfc-9b5d-b9ab5095e816 · outbound
AI Benchmarks and Datasets for LLM Evaluation https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1232386a-9562-4f1a-a896-34461e041741 · outbound
AI Benchmarks and Datasets for LLM Evaluation Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ba014203-b1b9-423d-a334-fd5ee4d2adb7 · outbound
AI Benchmarks and Datasets for LLM Evaluation Measuring massive multita sk language understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5d3b5543-7ead-435d-8400-7e2da9421296 · outbound
AI Benchmarks and Datasets for LLM Evaluation Measuri ng mathe- matical problem solving with the MATH dataset
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8dcf9c9a-e335-440e-92ff-49c4d3a6b235 · outbound
AI Benchmarks and Datasets for LLM Evaluation Weld, and Luke Zett lemoyer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 31d31134-2521-42bc-890c-fcb31d6d2140 · outbound
AI Benchmarks and Datasets for LLM Evaluation OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e6adef-20f7-49ad-8eda-32d82c7a0935 · outbound
AI Benchmarks and Datasets for LLM Evaluation Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a3ab6eb7-bfbb-4464-ae23-cf7c389c373f · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b024e696-d46b-48e5-baa3-d74efc343611 · outbound
AI Benchmarks and Datasets for LLM Evaluation GLoRE: Evaluating Logical Reasoning of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9509a7b-ff82-44b9-8da8-450226793d40 · outbound
AI Benchmarks and Datasets for LLM Evaluation Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ab148b2-0aa1-4d3d-b87b-0a3b5c1f0b6e · outbound
AI Benchmarks and Datasets for LLM Evaluation Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0fd44be2-0d82-4afa-b5ea-4905c02a7fc3 · outbound
AI Benchmarks and Datasets for LLM Evaluation Abstractive text summarization u sing sequence- to-sequence rnns and beyond
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 73b65f13-8f29-4de2-b2bf-a0fc493bcc68 · outbound
AI Benchmarks and Datasets for LLM Evaluation Adversarial NLI: A new benchmark for natura l language understanding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation caa6d431-23da-4cb6-a193-00024761feeb · outbound
AI Benchmarks and Datasets for LLM Evaluation https://oecd.ai/en/ai-pr inciples
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a9363cd5-ffaf-4dcd-91b5-f9b2e87961e3 · outbound
AI Benchmarks and Datasets for LLM Evaluation The LAMBADA dataset: Word prediction requir- ing a broad discourse context
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation de3aa830-3499-4d37-b28c-7f1490d7eda4 · outbound
AI Benchmarks and Datasets for LLM Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32cc2676-71cc-441d-a613-b2dc3c66fdad · outbound
AI Benchmarks and Datasets for LLM Evaluation Winogrande: An adversarial winograd schema challenge at sc ale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6348edc0-763c-4d98-981e-065481892f0a · outbound
AI Benchmarks and Datasets for LLM Evaluation Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 141e0544-db56-4537-8ad7-db555dc02946 · outbound
AI Benchmarks and Datasets for LLM Evaluation Sur- vey of different large language model architectures: Trends , benchmarks, and challenges
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5c0eef8c-bb5a-4fe3-8961-57df44b03318 · outbound
AI Benchmarks and Datasets for LLM Evaluation Concep tnet 5.5: An open multilingual graph of general knowledge
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8e2cb9b3-8344-4777-9005-28c997947f30 · outbound
AI Benchmarks and Datasets for LLM Evaluation Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a20935e0-0fbc-4411-88fd-e8cd5c050af2 · outbound
AI Benchmarks and Datasets for LLM Evaluation Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cf6d6eb4-7de6-4ec0-b3e1-a900a43a10d8 · outbound
AI Benchmarks and Datasets for LLM Evaluation A corpus for reasoning about natural language gr ounded in photographs, 2019
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a87cdc31-c0d4-4d97-9ffc-8a2595dc2b75 · outbound
AI Benchmarks and Datasets for LLM Evaluation Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 436f95fd-5862-4c93-964f-41ceb378d71f · outbound
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5870f8e5-7fd8-4410-af03-e0a4673eb9d5 · outbound
AI Benchmarks and Datasets for LLM Evaluation Commonsenseqa: A question answering challenge targeting c ommonsense knowledge
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation da8f0378-5315-42b6-87ca-298895d91458 · outbound
AI Benchmarks and Datasets for LLM Evaluation Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954456a9-cbf4-4331-b187-94545c406c76 · outbound
AI Benchmarks and Datasets for LLM Evaluation Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 25f27a24-2548-4309-b237-0cff8158c062 · outbound
AI Benchmarks and Datasets for LLM Evaluation Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 64a38680-7320-44cf-b4b3-208c97dd72fa · outbound
AI Benchmarks and Datasets for LLM Evaluation Introducing v0.5 of the AI Safety Benchmark from MLCommons
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef0f9f7-90d2-4a76-8294-9d88f86a62e0 · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dd9703b6-3513-4e6c-b93a-19a841bcb088 · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9694d149-2b58-40a5-b696-3d6f82c39ce7 · outbound
AI Benchmarks and Datasets for LLM Evaluation CORD-19: The COVID-19 Open Research Dataset
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12ca922d-2dcb-43e9-8888-3c4b06c828be · outbound
AI Benchmarks and Datasets for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 758a8961-3d6a-4944-b487-63ddebaff14b · outbound
AI Benchmarks and Datasets for LLM Evaluation CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c69c960-5975-41e9-98fb-689d2e8a499f · outbound
AI Benchmarks and Datasets for LLM Evaluation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a57a991-f3da-4a06-8e1a-a1c0b5aecbbe · outbound
AI Benchmarks and Datasets for LLM Evaluation Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5e7dfec3-743e-430c-9517-9d9a8d40c46b · outbound
AI Benchmarks and Datasets for LLM Evaluation Eval- uating the quality of hallucination benchmarks for large vi sion-language models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f81d44-da0d-41b9-ab53-27ec1ca13751 · outbound
AI Benchmarks and Datasets for LLM Evaluation https://z-inspection.org/
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 42407999-af0d-4843-9c5c-43c813337483 · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8cc604dd-0a84-4bfb-ac5e-869ec535b2d4 · outbound
AI Benchmarks and Datasets for LLM Evaluation Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 79e80c6a-d471-4375-b5de-162f8947408f · outbound
AI Benchmarks and Datasets for LLM Evaluation MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f429688-6ded-4117-9a01-d6e88f1f15cd · outbound
AI Benchmarks and Datasets for LLM Evaluation Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb8602bc-0507-49ad-b393-0c6d3a0fdaff · outbound
AI Benchmarks and Datasets for LLM Evaluation Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2207480c-5c0f-4a4c-9051-22aba14151ab · outbound
AI Benchmarks and Datasets for LLM Evaluation Instruction-Following Evaluation for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56dd0cbf-1cce-40a4-965c-225238405a70 · outbound
AI Benchmarks and Datasets for LLM Evaluation CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf60c4e1-4baf-46d8-b430-995295a8caff · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c1835ad-51b0-4747-97f1-e0064a538d51 · outbound
AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ddab16a9-f941-4960-9d2b-7ecee0ac701b · inbound
Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning AI Benchmarks and Datasets for LLM Evaluation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.