Pith. sign in

Paper Citation Record · LEDGER

Training on the Benchmark Is Not All You Need

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2409.01790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01790 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:30:34.695025Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:19:14.868572Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 17a8cc32-8cd0-499f-9d89-3ab18a11bf20 · inbound

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? cites this paper.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Training on the Benchmark Is Not All You Need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.695025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.695025Z digest=sha256:fe4dcbb2baf30515ad71e3e857a521dbecfb78c6ff571c42913955bb93b38ffa

Observation e82a9929-e61d-434f-9c9b-887540602385 · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Training on the Benchmark Is Not All You Need

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.134851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.134851Z digest=sha256:c3a27878a1e72e460b8a1c08cb20f7f67be4a10940a2c0c1bb55d48ca086875d

Observation 11a417e7-a81b-403a-b639-4071f017c38e · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Training on the Benchmark Is Not All You Need

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.442357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.442357Z digest=sha256:72e6d29a94915f90d425512dbf415ce1d69ae0c7fa088e602c9e9ee4282a670d

Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · inbound

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation cites this paper.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.501910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.501910Z digest=sha256:6912593c007c54ed9d35c12e59d098068db2fbc677c6ba64a858fb241a617cf1

Observation 5ee39aaa-cdd3-4af3-9b7b-a25b7cd89e8d · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Training on the Benchmark Is Not All You Need

Reference 168

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:15.045306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T10:19:07.723477Z digest=sha256:dfead8e43e2eaa0be9c5c372027007429c6f8e25d006624f5a5d63a33768899e