Pith. sign in

Paper Citation Record · LEDGER

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2507.19219.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19219 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:01:55.836203Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T05:41:33.026345Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T05:46:24.255634Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de701f72-c710-4c19-aa79-1cdb149ba1a5 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.508876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.508876Z digest=sha256:a73608eff42d2ee42bbcd1c69730495c63f9a7089fb7796f5412998db3ea2a9e

Observation ef661ef1-c1da-4461-bfe7-e42144758803 · outbound

This paper cites write newline.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.518400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.518400Z digest=sha256:f5f92649188fc20e22ca4db27b24033af4d9b01771e128b06149bf6ca46c6644

Observation a23bedcb-7cda-496c-9e55-d50cec278a01 · outbound

This paper cites online" 'onlinestring :=.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework online" 'onlinestring :=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.526142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.526142Z digest=sha256:1c746d1541ab2f4630e19217179d98b95746e73ad790947d6f1a6c6278792f75

Observation 25c9e4e2-3691-4a75-bddf-066fb4661db1 · outbound

This paper cites write newline.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework write newline

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.533594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.533594Z digest=sha256:3d7a9d65140086baf9f55c11db23bf7d8c1c0f22714cec92fb9ffb193919be81

Observation d3a171cd-56dd-4186-95aa-4a526f425728 · outbound

This paper cites Qwen2 Technical Report.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Qwen2 Technical Report

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.851023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.541887Z digest=sha256:aaafbde17895c24d9262f9c178716eb30d5a869ef74ef8f964b542a3ba933ea2

Observation d1b9016e-2025-468d-aa54-c3a51a77b0aa · outbound

This paper cites G.; and Chapelle, C.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework G.; and Chapelle, C

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.828269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.552138Z digest=sha256:8e81bb1865845b5d987e70e906c2ae93032fc8bd03bb721f7e01f3e618b062fc

Observation bf7bd0dc-f3de-4581-8a44-4ccb29f39f76 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.559260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.559260Z digest=sha256:4a1b09a28afd28f9fb292307224f82cacc0fb5ba26b736f77ef7c644105a36bd

Observation 20545eb8-2c45-488e-88b6-b1014d396d38 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.793642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.566010Z digest=sha256:1b99afcbd8accc5c79e914ce5c3241f0c18998669988c62308ac6929d7bbbd81

Observation 1a6acc29-5978-4d64-a9fa-a3e88cf3e8f3 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.772998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.573762Z digest=sha256:08eea0b45c46a4836c8cb91ce5f446bf6feafb2246855bd0988d26f85b3b4bb8

Observation 909831b6-c8c3-4b24-bc45-54b8228231d0 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.754109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.582581Z digest=sha256:211d73782c2c8a033843f1571a792f8e579437268d09541c8427edd7103314e1

Observation a6f395b1-a8f4-467d-9e14-e60b4c8dceab · outbound

This paper cites N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.734151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.592757Z digest=sha256:b85c0ec1f4f87243b88804e0302676c1beb13286d774be2a3b63f900bde2478e

Observation 975ef086-0d2f-4479-be7e-f50b7669ff4f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.602383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.602383Z digest=sha256:7807729671922f43bb123ae36c5ecbf600761a5a806d743934d541a02cdd45be

Observation c2580364-1796-4722-9084-5472b51f73ac · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.714746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.608638Z digest=sha256:bd96a3b292c6024964d49aa4e55bf84290fdbf9b53556193d1a50c4a34c284fb

Observation 38f91d24-a655-4e12-bc39-52964c09dd74 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.694654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.614316Z digest=sha256:cb2ea31d755b4676afe047cdc3efb59bd5d4f75bc72fb942ecb45997cf98bdec

Observation 5c996d1c-ca6c-4adf-b9bd-3b7c14346336 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.619561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.619561Z digest=sha256:69c08f7963b7f4063db822d4499603b5b2c47b565d3bb6631443ca0aa875fbf0

Observation 197b8bb9-3113-42e5-838b-060263718a05 · outbound

This paper cites Textbooks Are All You Need.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Textbooks Are All You Need

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.624898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.624898Z digest=sha256:249d9d9ad9a5215fdc50f4681aae3da35aa78382dfd28d88c64ccb5f20fe7f7e

Observation adee98e0-fa9c-4f64-952b-435a2bd005f3 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.630645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.630645Z digest=sha256:5154cb3733a6806715da0046d916a4f655ed6d4021be46949560e690a79c9fd2

Observation 64af435e-867f-448a-94c5-e241839bdff4 · outbound

This paper cites OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.636266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.636266Z digest=sha256:dd18769fadfb61e57df33c9618c8d9162cd0c31a8ed1dd7583c09fd507569d1a

Observation d6005ea8-7433-46bc-bc2d-d750f52730c1 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.642060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.642060Z digest=sha256:edbad30f645f3404600a7a505854133e0260b47edfe87c606893e9c8ddeefb4d

Observation f8adae2e-d0ec-4b1d-9a8d-fc7ace04c78c · outbound

This paper cites Investigating Data Contamination for Pre-training Language Models.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Investigating Data Contamination for Pre-training Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.648414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.648414Z digest=sha256:73e8498d4983f7d490c737878d2d92ea65930149bd42c35e7de5c73b7bc60c41

Observation 9793d50a-88dd-449c-9dcc-7cda7679c799 · outbound

This paper cites E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; and Narasimhan, K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.631862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.654260Z digest=sha256:e3fa4a9506a0524e0e6aa04d2c18d5f7a7ab48ea6e5bb33fd2096690ef20840f

Observation 75c7a7a7-ee09-46b9-993b-f3fd2f476ff1 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.662468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.662468Z digest=sha256:f03d2c91a9c15cd9c8eba00939ace556a90c09f47cf9c895ce5eb535eae38a38

Observation 32f46696-e39d-45a0-8542-9ef7e0fbd0fb · outbound

This paper cites E.; and Stoica, I.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework E.; and Stoica, I

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.610206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.668498Z digest=sha256:109a6ddca8dd92a439a7ec8de92ccaac5e3cc0b961c1ae2e0f56c8d5cd25afa9

Observation 1bab9562-35e8-488f-ab1a-41570e4bb1c3 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Textbooks Are All You Need II: phi-1.5 technical report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.674430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.674430Z digest=sha256:1a26955af2d94e7c05e0f2230b9c12fd258d853f729ba1b35b49663aa31e379f

Observation 50b439a7-defb-4651-8ee1-477d7e163daa · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.591048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.681134Z digest=sha256:13cd432929f0fe717327faf72efb450d7958d8b70902c91ea3282fd1e42f3778

Observation 0ba9996c-c450-4e83-8b35-e85b6cf4e6e7 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.573262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.686898Z digest=sha256:d8f1edbaf75c6bc46d4466e1dc2f9e91cd7efc01afa38c17bde05b9b0f457c63

Observation 9a1eb895-66d7-4f7d-b13a-e5a0135f118f · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.693029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.693029Z digest=sha256:a3feb4706992bcef075a9078587066aed6676a912cd68eaa1a1dc4cfb91e3a7e

Observation 8de1b18e-c067-44e0-9b9e-8c0f82b02c4b · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.551689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.699070Z digest=sha256:639f4712c2439bf5cd9fdf31caced80b1f21466cfc0e39241857d676177f2279

Observation 10941fed-2f6d-4968-be0a-6c7fedbc8ae0 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.704411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.704411Z digest=sha256:60e3d030ccb6835306f81ff0f8b07efd2d9a5b466bfdb7887f0f909970dc138e

Observation a6a548f3-fcfc-425b-a430-c3c840c0c252 · outbound

This paper cites GPT-4 Technical Report.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.709827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.709827Z digest=sha256:b3bd8bd955dc84f2d5486a145132f2695a4a204bf343b2d4162e21bb17ac5206

Observation 9f01e147-ff93-4e73-9c78-df41f77cf51c · outbound

This paper cites GPT-4o System Card.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.715343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.715343Z digest=sha256:31a51282307090db60d7ef905f8afd892661267ad9e53013365b7ca80be6efa0

Observation cb003761-1184-4789-b462-f56a05221eea · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.532668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.721411Z digest=sha256:6b7062fc461a2aa00857b7cfdd4640fd45afc8f873a8303e005f194b2822fea1

Observation 4204b1cb-f08f-4bb8-933d-288413222b3d · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.513939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.726781Z digest=sha256:2230ca894139503f20afc694a8c4ecbb9940b3ead388895d1e24910bc0d964e6

Observation 886fe9bf-27e8-4f60-b7e1-4ce63b4081be · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.495555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.732456Z digest=sha256:9c3459ff9f0902f011dcf0d29d0abf3bf95f340ffa35bc692158a219106ad45f

Observation df5bf444-4711-4382-9608-1659f3dfa045 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.738519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.738519Z digest=sha256:7b14027071413f66eff097afb7d16ebb4a9dab1d9b4d32e7b1d95326d75340d1

Observation 1461311b-bddd-401c-a678-87a7554ef371 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.473936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.744656Z digest=sha256:aaa2923de8831889666de578f3db30f183cb6ef6e9e5b8f451668df996e3c862

Observation b0bb6115-a0dd-48cb-b99d-ab1f8521afe6 · outbound

This paper cites R.; Zhang, S.; Sun, Y.; and Wang, W.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework R.; Zhang, S.; Sun, Y.; and Wang, W

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.451035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.750766Z digest=sha256:2741962abceed142659e70cba80dbeec9a4b6249dd858754acd047360fabcc5f

Observation 6ab5017e-a309-42cc-8af4-4e8d3216c147 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.758217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.758217Z digest=sha256:433502fb7c948522e4a08248627f15c268012b9b877d72427ce478b0bdb977eb

Observation 3db0d18a-3759-437d-8980-54c6b335f7b5 · outbound

This paper cites HelpSteer2-Preference: Complementing Ratings with Preferences.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework HelpSteer2-Preference: Complementing Ratings with Preferences

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.764894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.764894Z digest=sha256:9d0f8b3dcff66eb88a877ed66dea49e509fde389863d29520f6f0a41adcefcf9

Observation 5b274764-4af1-479b-8f8c-d41e653f5620 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.772078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.772078Z digest=sha256:027d9dfbc11309189878e450a2f9f118d89ade80742806ee4a29b0bf3e1e45f0

Observation 28de5af8-b5f2-4ac0-9dc9-7265f8dac8ad · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.779492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.779492Z digest=sha256:ee83f525ff123457f32414838ac73ea51f713da13c8f7a126a4f0ac1ae363526

Observation 1d6b58b7-c008-4b48-8b5f-6e16d748d806 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.428925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.786651Z digest=sha256:23a0a536531d19914bd0a029e816df9d68f27165c42a55f321d9e0024ce54380

Observation 6f6084a4-9800-4ef4-aab1-8a4103bece7a · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Benchmark Data Contamination of Large Language Models: A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.792774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.792774Z digest=sha256:3d00c4d318b7d48141ab3de410dcc085a61b26c3f981a0b6d1a18b268d0a1c11

Observation ea4ec8ab-bfca-45cd-acfa-51f14490acd4 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.799239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.799239Z digest=sha256:c49eb2cb526c97ae6d8d0feed853c7889a31218f8e7f04c9c9c2c4f636e1ce85

Observation be476101-66eb-474d-bf41-3bee196f84c9 · outbound

This paper cites Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:01:55.808143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:01:55.808143Z digest=sha256:5d2ec2437e7b971cded4b109dc4663f36954aae42f489aeb134a5eaad57f5d41

Observation 2c8fd61c-cbd6-49f4-921d-6eccb7cdbad9 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.409971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.813918Z digest=sha256:cf0df606d3d1874a22acbb7762a7db8e2dddf264b11905a9e913bbd5e1ddb113

Observation f1709b14-4be9-4403-b2fa-bcbd841b2cd5 · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.391357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.820500Z digest=sha256:8e7136c4b59b995f1c02c3dd92a3eae3c77a1ad80a8dfc7a9caa7b0e1aa21ff9

Observation 91f5e9fc-20ee-4e8d-ac92-aef7568cd88e · outbound

This paper cites Z.; Yang, D.; and Xie, X.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Z.; Yang, D.; and Xie, X

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:01:56.371048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.827990Z digest=sha256:536c0e6f52db5a39318b475e693535056ce46c6cf187522a7ca08f81b9c36749

Observation ac12e6b6-1d26-47f5-9171-0072111cbdab · outbound

This paper cites an unresolved cited work.

How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:01:56.351817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-15T18:01:55.836203Z digest=sha256:f5fb744fe78712c9af3617fa14243beef1fd91ff643321e3f4cdb0d1dee5adb0

Pith citing papers

Observation ebece4ca-6d1a-44e5-a104-5cf02be18e75 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-26T02:03:01.207018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:ee1985eb499d953eb1c1b8d236a444c93a7ebfb2de1c6026bc361de5f4b1d58b

Observation 6a1e66ce-ddef-4ac5-a4d7-52e11b478d8e · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-26T02:03:01.207018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:41:33.026345Z digest=sha256:c2e201fad87fefbe81e07b019fc383a065f06ecfe69c039d49e7844719e948cf