Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:01:55.836203Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2507.19219.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:01:55.836203Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:41:33.026345Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T05:46:24.255634Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de701f72-c710-4c19-aa79-1cdb149ba1a5 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework , " * write output.state after.block = add.period write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef661ef1-c1da-4461-bfe7-e42144758803 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23bedcb-7cda-496c-9e55-d50cec278a01 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework online" 'onlinestring :=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25c9e4e2-3691-4a75-bddf-066fb4661db1 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework write newline
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a171cd-56dd-4186-95aa-4a526f425728 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Qwen2 Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d1b9016e-2025-468d-aa54-c3a51a77b0aa · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework G.; and Chapelle, C
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bf7bd0dc-f3de-4581-8a44-4ccb29f39f76 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20545eb8-2c45-488e-88b6-b1014d396d38 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1a6acc29-5978-4d64-a9fa-a3e88cf3e8f3 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 909831b6-c8c3-4b24-bc45-54b8228231d0 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a6f395b1-a8f4-467d-9e14-e60b4c8dceab · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 975ef086-0d2f-4479-be7e-f50b7669ff4f · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Training Verifiers to Solve Math Word Problems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2580364-1796-4722-9084-5472b51f73ac · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 38f91d24-a655-4e12-bc39-52964c09dd74 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c996d1c-ca6c-4adf-b9bd-3b7c14346336 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 197b8bb9-3113-42e5-838b-060263718a05 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Textbooks Are All You Need
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adee98e0-fa9c-4f64-952b-435a2bd005f3 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64af435e-867f-448a-94c5-e241839bdff4 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6005ea8-7433-46bc-bc2d-d750f52730c1 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8adae2e-d0ec-4b1d-9a8d-fc7ace04c78c · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Investigating Data Contamination for Pre-training Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9793d50a-88dd-449c-9dcc-7cda7679c799 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 75c7a7a7-ee09-46b9-993b-f3fd2f476ff1 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32f46696-e39d-45a0-8542-9ef7e0fbd0fb · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework E.; and Stoica, I
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1bab9562-35e8-488f-ab1a-41570e4bb1c3 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Textbooks Are All You Need II: phi-1.5 technical report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50b439a7-defb-4651-8ee1-477d7e163daa · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ba9996c-c450-4e83-8b35-e85b6cf4e6e7 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9a1eb895-66d7-4f7d-b13a-e5a0135f118f · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8de1b18e-c067-44e0-9b9e-8c0f82b02c4b · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 10941fed-2f6d-4968-be0a-6c7fedbc8ae0 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6a548f3-fcfc-425b-a430-c3c840c0c252 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GPT-4 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f01e147-ff93-4e73-9c78-df41f77cf51c · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GPT-4o System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb003761-1184-4789-b462-f56a05221eea · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4204b1cb-f08f-4bb8-933d-288413222b3d · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 886fe9bf-27e8-4f60-b7e1-4ce63b4081be · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation df5bf444-4711-4382-9608-1659f3dfa045 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1461311b-bddd-401c-a678-87a7554ef371 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b0bb6115-a0dd-48cb-b99d-ab1f8521afe6 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework R.; Zhang, S.; Sun, Y.; and Wang, W
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6ab5017e-a309-42cc-8af4-4e8d3216c147 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db0d18a-3759-437d-8980-54c6b335f7b5 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework HelpSteer2-Preference: Complementing Ratings with Preferences
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b274764-4af1-479b-8f8c-d41e653f5620 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28de5af8-b5f2-4ac0-9dc9-7265f8dac8ad · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d6b58b7-c008-4b48-8b5f-6e16d748d806 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6f6084a4-9800-4ef4-aab1-8a4103bece7a · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Benchmark Data Contamination of Large Language Models: A Survey
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4ec8ab-bfca-45cd-acfa-51f14490acd4 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be476101-66eb-474d-bf41-3bee196f84c9 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8fd61c-cbd6-49f4-921d-6eccb7cdbad9 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1709b14-4be9-4403-b2fa-bcbd841b2cd5 · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 91f5e9fc-20ee-4e8d-ac92-aef7568cd88e · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Z.; Yang, D.; and Xie, X
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ac12e6b6-1d26-47f5-9171-0072111cbdab · outbound
How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ebece4ca-6d1a-44e5-a104-5cf02be18e75 · inbound
Measuring AI Reasoning: A Guide for Researchers How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6a1e66ce-ddef-4ac5-a4d7-52e11b478d8e · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.