Pith. sign in

Paper Citation Record · LEDGER

Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.09880.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.09880 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:46.197830Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:41.039569Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a670cd0f-1687-4bc1-96ff-f2dccbbd6bbf · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.042123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:1af245b1b36a8b10b67afcd306602ede0489722eaace644b05cc663b9dc2a5b0

Observation 2f38d407-5bf1-4fe6-ad94-b89fce08349a · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.396135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:718e7b46c1147f0ebc27a0032228ba2ebb208420e8fb5e58f3abbfebdd25530c

Observation d00f127b-f3c5-4b99-b092-d7e1cc5deca9 · inbound

Evaluating the Sensitivity of LLMs to Prior Context cites this paper.

Evaluating the Sensitivity of LLMs to Prior Context Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.197830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.197830Z digest=sha256:dedaf27e6664389cb6a3610658bce48e0b82f1a42e3f05df3caa4702f9b1ef7a

Observation ac7fea2d-36bc-4c4a-96dd-4c3d6f2e3949 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:49.091271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:49.091271Z digest=sha256:f88c76a04a6a83db43e8af988c70e1fe3259b849151d308cc85b25cbd09a5437

Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · inbound

Benchmarking the Pedagogical Knowledge of Large Language Models cites this paper.

Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.697366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.697366Z digest=sha256:515a928514787953d1e728e042f16865be744fd26a3375099060a60079da4f32

Observation 730348d7-1f43-4ae1-8aea-8f0731132ca0 · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:42.082274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:42.082274Z digest=sha256:ecdd483f7d6df4b2dd0309851565925356f92dcdbfc7a329125159bbe9b3d6e2

Observation 9c272ff5-c8e7-4443-82c4-911beb380184 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:23.835725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:23.835725Z digest=sha256:e97746bf59d3144bddb74cb493ab0c710bedd90da9a8c597709440eb4054b89f

Observation 6e79f6e4-785c-4eab-9a33-5d27681e7eac · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.466197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.466197Z digest=sha256:b7ee5ad28475f3b865d874f6ca64b3b1e8e504a29113a377a9ff29747a240233

Observation 8326260e-4bf2-4873-b1c8-2522a225cb5b · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:20:34.404981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:62c5b57bf87b6c8d1313b4f1271b4e5c4bf79da31ea2799faff283d921f363be

Observation e2b61dff-efae-45ec-847d-e10efba11dfd · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:51.487379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:51.487379Z digest=sha256:fdc8a48bd13c12761ecf3e1b71f9ad9ddfacec3961ae5e90142227d85adfba4b

Observation 0c9c93e9-8fc0-4bda-a6c1-b5b050d53f65 · inbound

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation cites this paper.

Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T14:42:44.631644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:42:44.631644Z digest=sha256:2b52aacd93989afd5f0cbc7e3a3b4793c5264c70f7b67ce8616da541989295bf