Pith. sign in

Paper Citation Record · LEDGER

AI Benchmarks and Datasets for LLM Evaluation

As of 22 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2412.01020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01020 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:49:34.499531Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T04:07:56.287352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:28:44.552688Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 959064f1-1cef-4252-906d-e03d12c2d360 · outbound

This paper cites https://aisafetybulgaria.c om/.

AI Benchmarks and Datasets for LLM Evaluation https://aisafetybulgaria.c om/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.494189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.275900Z digest=sha256:917c6ef138ccf2a86796b615da62bcac5a213e7e28f95c6b10d562d6435e17ac

Observation 768576e5-db86-4bd5-b740-1828ab195762 · outbound

This paper cites Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems.

AI Benchmarks and Datasets for LLM Evaluation Tpcx-ai - an industry standard benc hmark for artificial intelligence and machine learning systems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.481294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.280630Z digest=sha256:d4923a133ed69119829ca24c3bdd8771acea989c64e21d28337575a0a241d449

Observation 595bf9b7-49cb-428f-8c43-23511f8d17b0 · outbound

This paper cites https://compl-ai.org/.

AI Benchmarks and Datasets for LLM Evaluation https://compl-ai.org/

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.468167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.285007Z digest=sha256:75653f80f9341b100aa22ebee516eb86b5ff4fb97b275abefc7d79b429b99663

Observation b852ae81-e737-4721-831a-9bfe82e099b8 · outbound

This paper cites Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance.

AI Benchmarks and Datasets for LLM Evaluation Relevai-reviewer: A benchmark on AI reviewers for survey paper relevance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.289369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.289369Z digest=sha256:73602ffa15397ba5206722b760dc6ea6756ab5b0486c4ecd48d5ca221934b653

Observation c5a99daa-ed8a-4c81-b982-79c205a95d25 · outbound

This paper cites Robustbench: a standardized adversarial ro bustness bench- mark.

AI Benchmarks and Datasets for LLM Evaluation Robustbench: a standardized adversarial ro bustness bench- mark

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.455802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.293794Z digest=sha256:9a2343d3a7ad2d3fa14a3d464e102e2df067112357cdd8a2d440cc78fa6db7a2

Observation 75850ed8-80ce-4434-8daf-38b73a9c410e · outbound

This paper cites https://artificialintelligenceact.eu /the-act/.

AI Benchmarks and Datasets for LLM Evaluation https://artificialintelligenceact.eu /the-act/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.442610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.298264Z digest=sha256:6ecd5583fa95a8436ae3ed1e9e4ab8a4319b0201108b63c7103f5138fd81d8cc

Observation b267cce2-1bd9-4cfc-9b5d-b9ab5095e816 · outbound

This paper cites https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai.

AI Benchmarks and Datasets for LLM Evaluation https://di gital- strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.429455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.302992Z digest=sha256:86b9b25973be6c1ffe0e9a9e74d1b74b46662c4e23f5209b8f54ce19a3ff479f

Observation 1232386a-9562-4f1a-a896-34461e041741 · outbound

This paper cites Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024.

AI Benchmarks and Datasets for LLM Evaluation Compl-ai framework: A technical interpretation and llm benchmarkin g suite for the eu artificial intelligence act, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.416377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.307310Z digest=sha256:1742c99ebd33d69961dfb28bc9683577988b13cc05a57d1d5137a738df95b425

Observation ba014203-b1b9-423d-a334-fd5ee4d2adb7 · outbound

This paper cites Measuring massive multita sk language understanding.

AI Benchmarks and Datasets for LLM Evaluation Measuring massive multita sk language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.401262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.311537Z digest=sha256:d5143755635fa21c1d7bffeb089b6231365fac764c2b1d029a6049962513ad1d

Observation 5d3b5543-7ead-435d-8400-7e2da9421296 · outbound

This paper cites Measuri ng mathe- matical problem solving with the MATH dataset.

AI Benchmarks and Datasets for LLM Evaluation Measuri ng mathe- matical problem solving with the MATH dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.386821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.315694Z digest=sha256:657ca85ac3f0cda410595c34c9180806d5f75d48917877eb024419ce240d99a1

Observation 8dcf9c9a-e335-440e-92ff-49c4d3a6b235 · outbound

This paper cites Weld, and Luke Zett lemoyer.

AI Benchmarks and Datasets for LLM Evaluation Weld, and Luke Zett lemoyer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.373656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.319762Z digest=sha256:57cdbb508508c0220c855ef13d392c7480143ea51494c427b446af33dc9d4696

Observation 31d31134-2521-42bc-890c-fcb31d6d2140 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

AI Benchmarks and Datasets for LLM Evaluation OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.323908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.323908Z digest=sha256:569be9f8620f3c2a25bed48c2ae1a6703f52360ed62977b5704c84ede83d2d9e

Observation 10e6adef-20f7-49ad-8eda-32d82c7a0935 · outbound

This paper cites Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024.

AI Benchmarks and Datasets for LLM Evaluation Back- doorllm: A comprehensive benchmark for backdoor attacks on large lan- guage models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.359667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.328632Z digest=sha256:d944a0f488b31fbb349284159c3b248d87349e0b6789154ecc27f78bf5599558

Observation a3ab6eb7-bfbb-4464-ae23-cf7c389c373f · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.346655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.332779Z digest=sha256:483963d0adc1b751c2964bd595524b6225fbd973b69ed32ec22b805d94d880e5

Observation b024e696-d46b-48e5-baa3-d74efc343611 · outbound

This paper cites GLoRE: Evaluating Logical Reasoning of Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation GLoRE: Evaluating Logical Reasoning of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.337026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.337026Z digest=sha256:3db6876c74323bd7379d620f2bfdc3b4193f4a6b4218a0bf2f0c4e0c7778c364

Observation f9509a7b-ff82-44b9-8da8-450226793d40 · outbound

This paper cites Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning.

AI Benchmarks and Datasets for LLM Evaluation Metabox: A benchmark plat- form for meta-black-box optimization with reinforcement l earning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.333475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.341499Z digest=sha256:02f352f8b2f79b8e76b4e8900a5bdde8f0992b21accffc6d88a74a03d7cde10a

Observation 4ab148b2-0aa1-4d3d-b87b-0a3b5c1f0b6e · outbound

This paper cites Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019.

AI Benchmarks and Datasets for LLM Evaluation Ok-vqa: A visual question answering benchmark requi ring external knowledge, 2019

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.321026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.346322Z digest=sha256:8782ab3f3a25f3c08fdf4020bd651e496a2b95d137d1541e79f34df333bd5171

Observation 0fd44be2-0d82-4afa-b5ea-4905c02a7fc3 · outbound

This paper cites Abstractive text summarization u sing sequence- to-sequence rnns and beyond.

AI Benchmarks and Datasets for LLM Evaluation Abstractive text summarization u sing sequence- to-sequence rnns and beyond

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.308111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.351220Z digest=sha256:4876135383a221de1b13067a21b2e04f86a29b8cce2684e49f43d72b6cab0f95

Observation 73b65f13-8f29-4de2-b2bf-a0fc493bcc68 · outbound

This paper cites Adversarial NLI: A new benchmark for natura l language understanding.

AI Benchmarks and Datasets for LLM Evaluation Adversarial NLI: A new benchmark for natura l language understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.293849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.355412Z digest=sha256:89a1f35a47f0fb006810635bdff49da6d43e604a6c88d3300f4e2dec88653ecb

Observation caa6d431-23da-4cb6-a193-00024761feeb · outbound

This paper cites https://oecd.ai/en/ai-pr inciples.

AI Benchmarks and Datasets for LLM Evaluation https://oecd.ai/en/ai-pr inciples

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.280799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.359484Z digest=sha256:14a2f41c95d1f86d80d147995df7121dd40031c16301c8e02f719e282482610b

Observation a9363cd5-ffaf-4dcd-91b5-f9b2e87961e3 · outbound

This paper cites The LAMBADA dataset: Word prediction requir- ing a broad discourse context.

AI Benchmarks and Datasets for LLM Evaluation The LAMBADA dataset: Word prediction requir- ing a broad discourse context

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.266378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.363678Z digest=sha256:98b2d984526d5ad252fe81882396c70de0ac73958e592ab61a053ef93b808144

Observation de3aa830-3499-4d37-b28c-7f1490d7eda4 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

AI Benchmarks and Datasets for LLM Evaluation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.367867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.367867Z digest=sha256:a6225d609dfe89c9e2b6b3d6f40809eeb54930adb771fa84adfc761c2690f44f

Observation 32cc2676-71cc-441d-a613-b2dc3c66fdad · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at sc ale.

AI Benchmarks and Datasets for LLM Evaluation Winogrande: An adversarial winograd schema challenge at sc ale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.251863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.372425Z digest=sha256:1fc64520ed3802f1c8860fe508b65eb442f007b84833fada319b242001d51372

Observation 6348edc0-763c-4d98-981e-065481892f0a · outbound

This paper cites Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F.

AI Benchmarks and Datasets for LLM Evaluation Hadfiel d, Richard Ngo, Konstantin Pilz, George Gor, Emma Bluemke, Sarah Shoker, Ja net Egan, Robert F

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.239019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.376471Z digest=sha256:119f8778e3ee22779fac944984f5599c432683ce89e04ad06b5d44563f4ea6e7

Observation 141e0544-db56-4537-8ad7-db555dc02946 · outbound

This paper cites Sur- vey of different large language model architectures: Trends , benchmarks, and challenges.

AI Benchmarks and Datasets for LLM Evaluation Sur- vey of different large language model architectures: Trends , benchmarks, and challenges

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.224479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.380505Z digest=sha256:4a04f4bef49791e9f336cfcea5c9e540f517504ed6df920583a1cdabecc96953

Observation 5c0eef8c-bb5a-4fe3-8961-57df44b03318 · outbound

This paper cites Concep tnet 5.5: An open multilingual graph of general knowledge.

AI Benchmarks and Datasets for LLM Evaluation Concep tnet 5.5: An open multilingual graph of general knowledge

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.211305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.384428Z digest=sha256:77eb67069644fc646fb0dbbff635e65aadd7e2088ae84ba20ed7ff148978060b

Observation 8e2cb9b3-8344-4777-9005-28c997947f30 · outbound

This paper cites Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing.

AI Benchmarks and Datasets for LLM Evaluation Musr: Testing the limits of chain-of-thought with multiste p soft reason- ing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.198454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.388365Z digest=sha256:bcea1859ae0088719da4422a583e3ca862989802618ce75fae3fbd977eebd3c1

Observation a20935e0-0fbc-4411-88fd-e8cd5c050af2 · outbound

This paper cites Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024.

AI Benchmarks and Datasets for LLM Evaluation Conflictbank: A benchmark for evaluating the influence of knowledge conflicts in llm, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.185345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.392518Z digest=sha256:761a09b20cec323a69375159e0e81345e7fdb72d7dc8e80ba9a48b0ec942a00a

Observation cf6d6eb4-7de6-4ec0-b3e1-a900a43a10d8 · outbound

This paper cites A corpus for reasoning about natural language gr ounded in photographs, 2019.

AI Benchmarks and Datasets for LLM Evaluation A corpus for reasoning about natural language gr ounded in photographs, 2019

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.172027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.396498Z digest=sha256:0ea9aab3dc3b8f4a28f9adff2d978d5268dd9b52cac1848dbf5691fdc06746a4

Observation a87cdc31-c0d4-4d97-9ffc-8a2595dc2b75 · outbound

This paper cites Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024.

AI Benchmarks and Datasets for LLM Evaluation Table meets llm: Can large language models understand struc tured table data? a benchmark and empirical study, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.157059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.400470Z digest=sha256:922f448963ccc857a35b360ec67a315591d671ad7c3b3fc1cbafff7e6119ae06

Observation 436f95fd-5862-4c93-964f-41ceb378d71f · outbound

This paper cites Le, Ed H.

AI Benchmarks and Datasets for LLM Evaluation Le, Ed H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.142715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.404516Z digest=sha256:c7eb23fb3d6a831700ad44b5a291722dce4e3a907056f3661b53f63b59a80e91

Observation 5870f8e5-7fd8-4410-af03-e0a4673eb9d5 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting c ommonsense knowledge.

AI Benchmarks and Datasets for LLM Evaluation Commonsenseqa: A question answering challenge targeting c ommonsense knowledge

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.128940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.408653Z digest=sha256:06ededa2ac21539e2a918cf0773b7b1e1f9d2cf094739a7fb089330520785458

Observation da8f0378-5315-42b6-87ca-298895d91458 · outbound

This paper cites Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls.

AI Benchmarks and Datasets for LLM Evaluation Ltlbench: Towards bench marks for evaluating temporal logic reasoning in large language mode ls

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.412604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.412604Z digest=sha256:a2c414183f3b3c5027f143ce3e8a6385eb62d6b7d994c7ea30d815422fdc0bf8

Observation 954456a9-cbf4-4331-b187-94545c406c76 · outbound

This paper cites Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C.

AI Benchmarks and Datasets for LLM Evaluation Kirkpatrick, Feiyi Wang, Tom Gibbs, Venkatram Vishwanath, Mallikarjun Shankar, Geoffrey C

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.115471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.416422Z digest=sha256:4aefbd10b7ff869f47639d1d4b30d343be4ab2e72776e53bf12faa81aa972dc8

Observation 25f27a24-2548-4309-b237-0cff8158c062 · outbound

This paper cites Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V.

AI Benchmarks and Datasets for LLM Evaluation Madai, Emilie Wiinb lad Mathez, 25 Jesmin Jahan Tithi, Magnus Westerlund, Renee Wurth, and Rob erto V

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.102325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.420387Z digest=sha256:7f25fb0935c8923d6413744c09dd97a566f465e839fb7c189fc508c0c3319318

Observation 64a38680-7320-44cf-b4b3-208c97dd72fa · outbound

This paper cites Introducing v0.5 of the AI Safety Benchmark from MLCommons.

AI Benchmarks and Datasets for LLM Evaluation Introducing v0.5 of the AI Safety Benchmark from MLCommons

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.424309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.424309Z digest=sha256:04febbf0372a22ed408096d6c538c4436a743df9e52dac25ec1955e787607292

Observation 8ef0f9f7-90d2-4a76-8294-9d88f86a62e0 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.089272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.428747Z digest=sha256:9f336b1ee28a0c8600235ff0e58c63fe7a89d6b27d333716ca5ae1f4c8358d10

Observation dd9703b6-3513-4e6c-b93a-19a841bcb088 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.074927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.432590Z digest=sha256:06da6d695a200e09bbfdd1500dd73a739dedab84c9f65c4b719fb04b9687bda2

Observation 9694d149-2b58-40a5-b696-3d6f82c39ce7 · outbound

This paper cites CORD-19: The COVID-19 Open Research Dataset.

AI Benchmarks and Datasets for LLM Evaluation CORD-19: The COVID-19 Open Research Dataset

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.436441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.436441Z digest=sha256:e4c6617b39bc3e19ced43c0aed0a1acbdf341b9442355d58cac6044075276a08

Observation 12ca922d-2dcb-43e9-8888-3c4b06c828be · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

AI Benchmarks and Datasets for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.440600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.440600Z digest=sha256:3a8554c00bf6518e0a9a3ed6a2f5759b941b804332b10abb40f0d728a3243a7a

Observation 758a8961-3d6a-4944-b487-63ddebaff14b · outbound

This paper cites CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A comprehensive benchmark for evaluating causal reasoning capabilities of large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.061640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.444965Z digest=sha256:38203139084e5b72307ef2078bf4efca1bcd137eff56371d6f5d902d14099360

Observation 8c69c960-5975-41e9-98fb-689d2e8a499f · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

AI Benchmarks and Datasets for LLM Evaluation LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.448945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.448945Z digest=sha256:74e6e9c06304e601d9ed6c2386e3c8f2914c107828c496ce1236bc36c4fe3d3d

Observation 8a57a991-f3da-4a06-8e1a-a1c0b5aecbbe · outbound

This paper cites Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering.

AI Benchmarks and Datasets for LLM Evaluation Quick and (not so) dirty: Unsupervised selection of justification sentences f or multi-hop ques- tion answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.453224Z digest=sha256:b04f3d46022134918f7b01da7c7a8c80c72227857c8e514d6965a2d3502a7272

Observation 5e7dfec3-743e-430c-9517-9d9a8d40c46b · outbound

This paper cites Eval- uating the quality of hallucination benchmarks for large vi sion-language models.

AI Benchmarks and Datasets for LLM Evaluation Eval- uating the quality of hallucination benchmarks for large vi sion-language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.457155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.457155Z digest=sha256:414e04c5c6e0537adc85897d11bfa5408e7f452eff3ca703b71e5be6075ce63d

Observation 20f81d44-da0d-41b9-ab53-27ec1ca13751 · outbound

This paper cites https://z-inspection.org/.

AI Benchmarks and Datasets for LLM Evaluation https://z-inspection.org/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:35.032335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.461684Z digest=sha256:0cffb6eef2952c64fd1c4305dceb0189bc2bd9435b8c7eb02c148dc50359bf9a

Observation 42407999-af0d-4843-9c5c-43c813337483 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.018580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.465655Z digest=sha256:dbcce1be86569d0f70bd7e134de6fe853debdc3fd1f1f49001e2f4c0f8fc5240

Observation 8cc604dd-0a84-4bfb-ac5e-869ec535b2d4 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R.

AI Benchmarks and Datasets for LLM Evaluation Hellaswag: Can a machine really finish your sentence? In Anna Korho- nen, David R

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.990279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.474079Z digest=sha256:0c76b8032e03f83cc3ff97cc389b6e53f08ba1228f0f7443f68f77a2f1f5a6db

Observation 79e80c6a-d471-4375-b5de-162f8947408f · outbound

This paper cites MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation MultiTrust: A Comprehensive Benchmark Towards Trustworthy Multimodal Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.478030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.478030Z digest=sha256:0395331798439a61a359ccd94f27d7bce53148f0e51142090c0db6d2dd5c396a

Observation 3f429688-6ded-4117-9a01-d6e88f1f15cd · outbound

This paper cites Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024.

AI Benchmarks and Datasets for LLM Evaluation Reef- knot: A comprehensive benchmark for relation hallucinatio n evaluation, analysis and mitigation in multimodal large language model s, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.976392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.482834Z digest=sha256:a66611328626180f107839398d953138abdefa069639393048d83800ebb2a936

Observation bb8602bc-0507-49ad-b393-0c6d3a0fdaff · outbound

This paper cites Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024.

AI Benchmarks and Datasets for LLM Evaluation Revolutionizing data base q&a with large language models: Comprehensive benchmark and evalua tion, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:49:34.961429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.486810Z digest=sha256:8888250e9054239c7a485a235c4cc44b42253f18d21a242bb8993ee05ca71b6e

Observation 2207480c-5c0f-4a4c-9051-22aba14151ab · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AI Benchmarks and Datasets for LLM Evaluation Instruction-Following Evaluation for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.490762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.490762Z digest=sha256:727560cf6e787ca237d060529df6c283b0be2642dd98180a4613518e607c8251

Observation 56dd0cbf-1cce-40a4-965c-225238405a70 · outbound

This paper cites CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs.

AI Benchmarks and Datasets for LLM Evaluation CausalBench: A Comprehensive Benchmark for Causal Learning Capability of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:49:34.495253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:49:34.495253Z digest=sha256:2a880e7626e345c0f412ae775a1c10a15fc2995d267440a5a47dd1296ebcadfe

Observation cf60c4e1-4baf-46d8-b430-995295a8caff · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:34.946685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.499531Z digest=sha256:4b664194f533d4b619e8d5a9c1f48454dec5bac4ccdfa23a9b60c8c8345b0110

Observation 3c1835ad-51b0-4747-97f1-e0064a538d51 · outbound

This paper cites an unresolved cited work.

AI Benchmarks and Datasets for LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:49:35.004190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T04:49:34.469869Z digest=sha256:3f7c879eed3a1fc87986cda8fcc3fc93b6355856068afd80b0438e644d2d8b49

Pith citing papers

Observation ddab16a9-f941-4960-9d2b-7ecee0ac701b · inbound

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning cites this paper.

Formalizing and Mitigating Structural Distortion in LLM Attention for Graph Reasoning AI Benchmarks and Datasets for LLM Evaluation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:28:44.554094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T04:07:56.287352Z digest=sha256:a63ee82409468d972258995dba998385b34a545ad34609f325d0e1aa586aa712