Pith. sign in

Paper Citation Record · LEDGER

AI Benchmarks and Datasets for LLM Evaluation

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2412.01020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01020 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:49:34.499531Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:56.287352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.552688Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959064f1-1cef-4252-906d-e03d12c2d360 · outbound

This paper cites https://aisafetybulgaria.c om/.

AI Benchmarks and Datasets for LLM Evaluation https://aisafetybulgaria.c om/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.494189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.275900Z digest=sha256:21600075f4cd5562ebe7b432d9faaa7662575dabbbd543180f94b9ca7a2a5f1a

Observation 768576e5-db86-4bd5-b740-1828ab195762 · outbound

This paper cites Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems.

AI Benchmarks and Datasets for LLM Evaluation Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.481294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.280630Z digest=sha256:2a59f9d3b5b628e5bba6e47bc753cae963aa232c78643dadb9ff479a26e1291e

Observation 595bf9b7-49cb-428f-8c43-23511f8d17b0 · outbound

This paper cites https://compl-ai.org/.

AI Benchmarks and Datasets for LLM Evaluation https://compl-ai.org/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.468167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.285007Z digest=sha256:8f3e9584861e2f7906b643624a37369cc57e20955a2b871d7f5ffffbd245bc7a

Observation b852ae81-e737-4721-831a-9bfe82e099b8 · outbound

This paper cites Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance.

AI Benchmarks and Datasets for LLM Evaluation Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.289369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.289369Z digest=sha256:46288b9b1ed13acdea5ab4551ce9b861165f117b714595211b60960ae18bdafa

Observation c5a99daa-ed8a-4c81-b982-79c205a95d25 · outbound

This paper cites Robustbench: a standardized adversarial ro bustness bench- mark.

AI Benchmarks and Datasets for LLM Evaluation Robustbench: a standardized adversarial ro bustness bench- mark

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.455802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.293794Z digest=sha256:95ca6dead336d7f2bcd57a937c17135c4dcbc78e11bdcd0cc8ded97ddb0abaab

Observation 75850ed8-80ce-4434-8daf-38b73a9c410e · outbound

This paper cites https://artificialintelligenceact.eu /the-act/.

AI Benchmarks and Datasets for LLM Evaluation https://artificialintelligenceact.eu /the-act/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.442610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.298264Z digest=sha256:8b62f10a9abb8072b44902c2c6992c987875c80cc7110bf752eaff883d7de386

Observation b267cce2-1bd9-4cfc-9b5d-b9ab5095e816 · outbound

This paper cites https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai.

AI Benchmarks and Datasets for LLM Evaluation https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.429455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.302992Z digest=sha256:3533a44994f7a8458574692d43de3f9a035548c03af78c041aaa642f472e6548

Observation 1232386a-9562-4f1a-a896-34461e041741 · outbound

This paper cites Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024.

AI Benchmarks and Datasets for LLM Evaluation Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.416377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.307310Z digest=sha256:e09fead9b6677317ede962594883c424215fd8b5ead8d3b60556f693602a2818

Observation ba014203-b1b9-423d-a334-fd5ee4d2adb7 · outbound

This paper cites Measuring massive multita sk language understanding.

AI Benchmarks and Datasets for LLM Evaluation Measuring massive multita sk language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.401262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.311537Z digest=sha256:ac4abe564c6b230ba4cc150d9fabab2eeb6242071b295edc6936c76b037a08d1

Observation 5d3b5543-7ead-435d-8400-7e2da9421296 · outbound

This paper cites Measuri ng mathe- matical problem solving with the MATH dataset.

AI Benchmarks and Datasets for LLM Evaluation Measuri ng mathe- matical problem solving with the MATH dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.386821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.315694Z digest=sha256:61ecbab1970050805280bb6ef0a2df65e4da70aa5b8c6cea2fdaf7d96e313a39

Observation 8dcf9c9a-e335-440e-92ff-49c4d3a6b235 · outbound

This paper cites Weld, and Luke Zett lemoyer.

AI Benchmarks and Datasets for LLM Evaluation Weld, and Luke Zett lemoyer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.373656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.319762Z digest=sha256:ad42f3cb85825557f19b5b602545e24c6198709b8903f379c7fda0380b5da7df

Observation 31d31134-2521-42bc-890c-fcb31d6d2140 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

AI Benchmarks and Datasets for LLM Evaluation OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.323908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.323908Z digest=sha256:c9a3f9c375409e37f9c85c9fc8dec81fbb9e957f11e0e473eed57fd8c086d295

Observation 10e6adef-20f7-49ad-8eda-32d82c7a0935 · outbound

This paper cites Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024.

AI Benchmarks and Datasets for LLM Evaluation Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.359667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.328632Z digest=sha256:4c6b9224a2cadf48654ff03b989109c3ea8937a32af1e5c11dcc77433b21f747

Observation a3ab6eb7-bfbb-4464-ae23-cf7c389c373f · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.346655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.332779Z digest=sha256:510ec782beaaadd9df1e3eeb1722b0e46b25885e4ab07375d12cb149c21b129f

Observation b024e696-d46b-48e5-baa3-d74efc343611 · outbound

This paper cites GLoRE: Evaluating Logical Reasoning of Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation GLoRE: Evaluating Logical Reasoning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.337026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.337026Z digest=sha256:4c3ce241cf7aed2c1f84dd873019d8b91a382dc3834212590dc52b7577b94205

Observation f9509a7b-ff82-44b9-8da8-450226793d40 · outbound

This paper cites Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning.

AI Benchmarks and Datasets for LLM Evaluation Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.333475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.341499Z digest=sha256:b4672ca335051a0798a86978c23cf7ca03f955a15f9152f746206eaaff1dbb11

Observation 4ab148b2-0aa1-4d3d-b87b-0a3b5c1f0b6e · outbound

This paper cites Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019.

AI Benchmarks and Datasets for LLM Evaluation Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.321026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.346322Z digest=sha256:7f7c9fb43dcde572639da32ebdc97c0f642ac58fc742051db92b83b38511e05b

Observation 0fd44be2-0d82-4afa-b5ea-4905c02a7fc3 · outbound

This paper cites Abstractive text summarization u sing sequence- to-sequence rnns and beyond.

AI Benchmarks and Datasets for LLM Evaluation Abstractive text summarization u sing sequence- to-sequence rnns and beyond

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.308111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.351220Z digest=sha256:93047a5fc5e17cf0803dff064d52cfa2c5bdb641f192e35b0803b64e19951a68

Observation 73b65f13-8f29-4de2-b2bf-a0fc493bcc68 · outbound

This paper cites Adversarial NLI: A new benchmark for natura l language understanding.

AI Benchmarks and Datasets for LLM Evaluation Adversarial NLI: A new benchmark for natura l language understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.293849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.355412Z digest=sha256:a4f0ecba1c644ebe8f4fdcfee3ab2237845ebf4414bdd3b29a1a1cb024e2b83c

Observation caa6d431-23da-4cb6-a193-00024761feeb · outbound

This paper cites https://oecd.ai/en/ai-pr inciples.

AI Benchmarks and Datasets for LLM Evaluation https://oecd.ai/en/ai-pr inciples

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.280799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.359484Z digest=sha256:f738da7ef71cc227e4b1a72cab7a164bd80ef902b5920f9c8089a91c728eba4b

Observation a9363cd5-ffaf-4dcd-91b5-f9b2e87961e3 · outbound

This paper cites The LAMBADA dataset: Word prediction requir- ing a broad discourse context.

AI Benchmarks and Datasets for LLM Evaluation The LAMBADA dataset: Word prediction requir- ing a broad discourse context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.266378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.363678Z digest=sha256:48c2b50f438a27063dc557e435afd83cc63b82da19f5f11ae96505d8a7432c6f

Observation de3aa830-3499-4d37-b28c-7f1490d7eda4 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

AI Benchmarks and Datasets for LLM Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.367867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.367867Z digest=sha256:001b251d3a8df86f1f07e31a5e7952d69d75e3169e852588332c309af355602d

Observation 32cc2676-71cc-441d-a613-b2dc3c66fdad · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at sc ale.

AI Benchmarks and Datasets for LLM Evaluation Winogrande: An adversarial winograd schema challenge at sc ale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.251863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.372425Z digest=sha256:247712af0e7e6946af0ca76cb6dab0ef086e546cccb2166471d9ba5e05e19629

Observation 6348edc0-763c-4d98-981e-065481892f0a · outbound

This paper cites Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F.

AI Benchmarks and Datasets for LLM Evaluation Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.239019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.376471Z digest=sha256:050a74b8eecb651bb3bf28633322ec07a305ec3ca1e2e126704f16fc88fb2008

Observation 141e0544-db56-4537-8ad7-db555dc02946 · outbound

This paper cites Sur- vey of different large language model architectures: Trends , benchmarks, and challenges.

AI Benchmarks and Datasets for LLM Evaluation Sur- vey of different large language model architectures: Trends , benchmarks, and challenges

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.224479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.380505Z digest=sha256:a6f1eebede42e6978e29c5669ecd77773a53b97c8af5f3b0381a09aa946ad994

Observation 5c0eef8c-bb5a-4fe3-8961-57df44b03318 · outbound

This paper cites Concep tnet 5.5: An open multilingual graph of general knowledge.

AI Benchmarks and Datasets for LLM Evaluation Concep tnet 5.5: An open multilingual graph of general knowledge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.211305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.384428Z digest=sha256:2848edace5fab1f35da5ee98d7eabcfe2e77a8f9ecd8f7891303d92a82c607f1

Observation 8e2cb9b3-8344-4777-9005-28c997947f30 · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing.

AI Benchmarks and Datasets for LLM Evaluation Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.198454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.388365Z digest=sha256:613e1e1090b1765ef7124f6f51e4e43a200da57f3ad50fd8d21333bea6f74af7

Observation a20935e0-0fbc-4411-88fd-e8cd5c050af2 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

AI Benchmarks and Datasets for LLM Evaluation Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.185345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.392518Z digest=sha256:1ca7d9ddbbf1057221a4562e3728b776141fccf31ffc5fef0183c4382c0aeee5

Observation cf6d6eb4-7de6-4ec0-b3e1-a900a43a10d8 · outbound

This paper cites A corpus for reasoning about natural language gr ounded in photographs, 2019.

AI Benchmarks and Datasets for LLM Evaluation A corpus for reasoning about natural language gr ounded in photographs, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.172027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.396498Z digest=sha256:5feea6db7b4c53a1d8dc032a2bd0771d5f923facf01cfe3f220acb18c801c781

Observation a87cdc31-c0d4-4d97-9ffc-8a2595dc2b75 · outbound

This paper cites Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024.

AI Benchmarks and Datasets for LLM Evaluation Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.157059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.400470Z digest=sha256:db3b40f5c1a893cb62748703b5fd8fa14b95ec6c9932aefda111802133129c6f

Observation 436f95fd-5862-4c93-964f-41ceb378d71f · outbound

This paper cites Le, Ed H.

AI Benchmarks and Datasets for LLM Evaluation Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.142715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.404516Z digest=sha256:0669af5d4286182a282fc18c032651c6fd9f3ac06e3ce8558678fee8590fd441

Observation 5870f8e5-7fd8-4410-af03-e0a4673eb9d5 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting c ommonsense knowledge.

AI Benchmarks and Datasets for LLM Evaluation Commonsenseqa: A question answering challenge targeting c ommonsense knowledge

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.128940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.408653Z digest=sha256:defab2ca230740c475a686dd09cf59d1071d6bf12b621fc98e2a5cc82f564fbb

Observation da8f0378-5315-42b6-87ca-298895d91458 · outbound

This paper cites Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls.

AI Benchmarks and Datasets for LLM Evaluation Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.412604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.412604Z digest=sha256:dd3c7ff30d072932774472524642084b2ea9ee47e136f874a4ec76c3d78b9ba9

Observation 954456a9-cbf4-4331-b187-94545c406c76 · outbound

This paper cites Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C.

AI Benchmarks and Datasets for LLM Evaluation Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.115471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.416422Z digest=sha256:6f4c9a42b05ecf331a3f9272b73128d385539a6b54364c5fbdf9800cfc7cc202

Observation 25f27a24-2548-4309-b237-0cff8158c062 · outbound

This paper cites Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V.

AI Benchmarks and Datasets for LLM Evaluation Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.102325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.420387Z digest=sha256:17233ac634014d79466b655617d9ec8a31b83ff911afa5f70756c8541c512f90

Observation 64a38680-7320-44cf-b4b3-208c97dd72fa · outbound

This paper cites Introducing v0.5 of the AI Safety Benchmark from MLCommons.

AI Benchmarks and Datasets for LLM Evaluation Introducing v0.5 of the AI Safety Benchmark from MLCommons

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.424309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.424309Z digest=sha256:d51f8830e18ca17a3b0227a638bef38e83399a182a17919b8d8a9fda1e6dd280

Observation 8ef0f9f7-90d2-4a76-8294-9d88f86a62e0 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.089272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.428747Z digest=sha256:288ed8d259983f561b72d915c1ec547012ba8832606bc816f99f3f98832af561

Observation dd9703b6-3513-4e6c-b93a-19a841bcb088 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.074927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.432590Z digest=sha256:40bdb91b3d41694aa5073917623d8f66900341325fb33fb6fbe502a784917b65

Observation 9694d149-2b58-40a5-b696-3d6f82c39ce7 · outbound

This paper cites CORD-19: The COVID-19 Open Research Dataset.

AI Benchmarks and Datasets for LLM Evaluation CORD-19: The COVID-19 Open Research Dataset

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.436441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.436441Z digest=sha256:584cbc630524a802edf729f76fba3ab5cb850fd904f9531bd70bb992241c8736

Observation 12ca922d-2dcb-43e9-8888-3c4b06c828be · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

AI Benchmarks and Datasets for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.440600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.440600Z digest=sha256:f78b4546d1b5949f174f25bff685e453119665585b85ef4d7bd73c9a3e539b6c

Observation 758a8961-3d6a-4944-b487-63ddebaff14b · outbound

This paper cites CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.061640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.444965Z digest=sha256:1873e7d543d8410111180146314c0ee28999a95ad6aef5d7aeec5e43d360eb60

Observation 8c69c960-5975-41e9-98fb-689d2e8a499f · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

AI Benchmarks and Datasets for LLM Evaluation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.448945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.448945Z digest=sha256:13d0d62aefdd7cfb4b465f6c77928907da1a70074befec4fe93213c91a3653dc

Observation 8a57a991-f3da-4a06-8e1a-a1c0b5aecbbe · outbound

This paper cites Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering.

AI Benchmarks and Datasets for LLM Evaluation Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.453224Z digest=sha256:4f725a560fd1081929badedc07d3953bd0f38c777d13955816f1263fab0162f2

Observation 5e7dfec3-743e-430c-9517-9d9a8d40c46b · outbound

This paper cites Eval- uating the quality of hallucination benchmarks for large vi sion-language models.

AI Benchmarks and Datasets for LLM Evaluation Eval- uating the quality of hallucination benchmarks for large vi sion-language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.457155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.457155Z digest=sha256:6ec6645d198b493af1b3d30dbbbfb133905bcdd3991b97ed6fa85c54f826038f

Observation 20f81d44-da0d-41b9-ab53-27ec1ca13751 · outbound

This paper cites https://z-inspection.org/.

AI Benchmarks and Datasets for LLM Evaluation https://z-inspection.org/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.032335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.461684Z digest=sha256:386c64f976231dee115bfdbc8e3a2ebf722bc8b7fd7171b79aee9dfbb956d853

Observation 42407999-af0d-4843-9c5c-43c813337483 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.018580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.465655Z digest=sha256:2ddeda1e7f48cfcc4ca09ce3e37e7fdf7b716564268855f38fd290db5a5c87eb

Observation 8cc604dd-0a84-4bfb-ac5e-869ec535b2d4 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R.

AI Benchmarks and Datasets for LLM Evaluation Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.990279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.474079Z digest=sha256:5081e39956dd910fa4dea8836bef8ebabf567e1457c71a32354496bc33ea238c

Observation 79e80c6a-d471-4375-b5de-162f8947408f · outbound

This paper cites MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.478030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.478030Z digest=sha256:6c76a915a9096e0abebd9f8776c448d8d57b097bceb153d9b5c66d2415fc3695

Observation 3f429688-6ded-4117-9a01-d6e88f1f15cd · outbound

This paper cites Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024.

AI Benchmarks and Datasets for LLM Evaluation Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.976392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.482834Z digest=sha256:a5449221db01c4189948a6e8e1eae3392ff04ae08f82c6283f1ef32455de8f8b

Observation bb8602bc-0507-49ad-b393-0c6d3a0fdaff · outbound

This paper cites Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024.

AI Benchmarks and Datasets for LLM Evaluation Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.961429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.486810Z digest=sha256:9b263279893260342e5197d774f48c1636f7d835c1e14b29f24024d108fc1c73

Observation 2207480c-5c0f-4a4c-9051-22aba14151ab · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.490762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.490762Z digest=sha256:2da4115fa90f18e463e2e0ec52f894819a41bd4b5bf2e9053f83aa08026c748e

Observation 56dd0cbf-1cce-40a4-965c-225238405a70 · outbound

This paper cites CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.495253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.495253Z digest=sha256:0f9facec9614ef03c768f6e664b3e9a210c5e00c1ea6fb8eccf1bc821b682051

Observation cf60c4e1-4baf-46d8-b430-995295a8caff · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:34.946685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.499531Z digest=sha256:73f3463c0ed1e357023af5eb68992e8450bfacbe3f998c8f35ca648f3cc2cdd9

Observation 3c1835ad-51b0-4747-97f1-e0064a538d51 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.004190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T04:49:34.469869Z digest=sha256:b61e59043440b3ef6ced2e73b87bacaf4d5389bcd7d69a1eb794873cee9bff7e

Pith citing papers

Observation ddab16a9-f941-4960-9d2b-7ecee0ac701b · inbound

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning cites this paper.

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning AI Benchmarks and Datasets for LLM Evaluation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.554094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T04:07:56.287352Z digest=sha256:3dd88f8f7cda916177b1b5a28f33b88b8970e0aca35af5ca18531ff53139d788