Pith. sign in

Paper Citation Record · LEDGER

Towards Contamination Resistant Benchmarks

As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2505.08389.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08389 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:10.470063Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:19:50.354875Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:23:07.041987Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69940e1e-cddd-4802-a9de-6a158853e6e8 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Towards Contamination Resistant Benchmarks What learning algorithm is in-context learning? investigations with linear models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.035486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.035486Z digest=sha256:24e96b4b89e09f98fd7f3a43b40608beeec299c4db696ee8f6a1e87809eb39e6

Observation fb7d25df-a325-4e73-a1b0-4198f9551556 · outbound

This paper cites The Surprising Effectiveness of Test-Time Training for Few-Shot Learning.

Towards Contamination Resistant Benchmarks The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.060656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.060656Z digest=sha256:da22074d6f46e6810180c35e2707d9aaf18466b30cc9422d4340395832206d35

Observation 679693a7-2f00-4adf-bb6c-f512b27a3bbb · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms.

Towards Contamination Resistant Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.363940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.096677Z digest=sha256:203466d04d9f00b143064f20fbb1d750e3c6d9159bdda905b82ae1c20e7c2365

Observation 61240b48-40e3-4d10-b95e-e38734ebd56b · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Towards Contamination Resistant Benchmarks Do, Yan Xu, and Pascale Fung

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.106629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.106629Z digest=sha256:46f72fae58e6671ffdf34a0a95c7f10fd4d585cb51f280f8a54bb479e4a325ad

Observation da1627ff-46a5-47f4-b407-c5ffef23943e · outbound

This paper cites Managing extreme ai risks amid rapid progress.

Towards Contamination Resistant Benchmarks Managing extreme ai risks amid rapid progress

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.114432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.114432Z digest=sha256:6ea65763495b04fac0864f0f285ddb16eda7a7347726b59c15ea25ced11d7a0b

Observation 96c32ab6-d82b-4475-b5a9-ce0a2804a633 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al.

Towards Contamination Resistant Benchmarks Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.332096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.126612Z digest=sha256:54e15dac09055b7bacb6dcffbc35c4dab5a72be2935534b1ed6b6d6a4346fe1a

Observation a222112e-aace-4c1a-bd86-27d64ddfa00c · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Towards Contamination Resistant Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.135591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.135591Z digest=sha256:0f014059192d6fb7da5103568882a87454624c167bea967198609f8cbaf1f864

Observation 9afb45fe-9725-4a00-953f-ba4af404d928 · outbound

This paper cites TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs.

Towards Contamination Resistant Benchmarks TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.141863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.141863Z digest=sha256:bbdfa740acb4f2588abe2f991966f71e616ab46f01c034bb52f0e6dfac1185fd

Observation 87c02f98-6247-419a-8ac1-f330f1090157 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Towards Contamination Resistant Benchmarks Palm: Scaling language modeling with pathways

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.315065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.156182Z digest=sha256:402c7f43085d062b8c836f343b1fec8420df03b721ca0c09407498e572a29b20

Observation 9b19019a-fbde-4bf2-9906-fbb9ef6bf8e6 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Towards Contamination Resistant Benchmarks Scaling Instruction-Finetuned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.164195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.164195Z digest=sha256:7cf1058089b2066daac75bbe7ac30ea7c843247e3d2cf7e610b77ef9b1305b7b

Observation 5a102519-49cd-4d08-8674-c0cfcf141338 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Towards Contamination Resistant Benchmarks Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.171454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.171454Z digest=sha256:2cda99eaf6337feea9507de4c3167dfb6e1e0e0acfb73b117503652c0440f55a

Observation 67c3b58c-0549-4070-9353-333d8b900a68 · outbound

This paper cites Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers.

Towards Contamination Resistant Benchmarks Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.300322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.179869Z digest=sha256:8e434367afca51434a30b1f2a9a46464125cfaadd8b5b03ad80a698d87395986

Observation 0db4f8e0-25b6-4f1f-ae6c-6bb1978294c1 · outbound

This paper cites Generalization or memorization: Data contamination and trustworthy evaluation for large language models.

Towards Contamination Resistant Benchmarks Generalization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.284886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.185895Z digest=sha256:2e175dafaa7d4a1d10d8ee40ba5b70cbbbbeda467b6d4fb31972d1c1c4d39472

Observation 9682585c-cd52-4507-93c8-ea1ec6a6b3c6 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Contamination Resistant Benchmarks The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.190458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.190458Z digest=sha256:7a5d299072d4f913e80e027421d6702d9caaaaf88f3c6f6c0d2bc5ca22ffc7ca

Observation 57bfe173-e544-4f93-879f-1cadb603ae5a · outbound

This paper cites Cole, Fangyu Liu, and William W.

Towards Contamination Resistant Benchmarks Cole, Fangyu Liu, and William W

Reference 15

Resolution
verified exact
doi, observed 2026-08-15T22:00:10.787781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.196731Z digest=sha256:cf280d8e30b60958c69619b1c4d21a320801c8883f6cf6a00f6eaf69ca5e42b8

Observation 06e56a22-f3b8-4d18-943a-df202aaec5ab · outbound

This paper cites What can transformers learn in-context? A case study of simple function classes.

Towards Contamination Resistant Benchmarks What can transformers learn in-context? A case study of simple function classes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.269783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.201762Z digest=sha256:f8cd79b80d764c8ff994ab2fbed3fbcf13f9e708f1550625aa2024f071779fd4

Observation 78704c4c-d05c-4b92-b5ac-329530667238 · outbound

This paper cites Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt.

Towards Contamination Resistant Benchmarks Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.254168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.207651Z digest=sha256:39df86738c39384a61d6ad522bc6816601d409bf262a82e5605cde4a24c58542

Observation 9c551f1c-83c8-45a8-9a2d-e3c0714bbc48 · outbound

This paper cites Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias.

Towards Contamination Resistant Benchmarks Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.213556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.213556Z digest=sha256:7fef00f73372aae78c6831b9a5ec7e4eee1eb91f9f42d2ab3469770b8f20c300

Observation 0e3b71f0-7d93-4473-815c-50eda4a52d82 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Towards Contamination Resistant Benchmarks LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.217836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.217836Z digest=sha256:f68ae02ffcb65de6fa59e5b6a6a5467dc2f6b927d7574d8f1346e23a23be555d

Observation ff868702-fe1a-423e-99fa-885499a739df · outbound

This paper cites Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models.

Towards Contamination Resistant Benchmarks Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.240666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.224057Z digest=sha256:c68e009c25fc3a64c0a43155541dff1f06b28caf1d0a6ebbc8153b4a20c23331

Observation faa7495c-5684-4147-86ab-518b2fae3cc3 · outbound

This paper cites Large language models are zero-shot reasoners.

Towards Contamination Resistant Benchmarks Large language models are zero-shot reasoners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.224617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.256194Z digest=sha256:a1179e9e37d7ef01d8cb67e60a6c7a2838fb73b6f38837a7ad16f116965df22e

Observation 120ad59c-f67d-498d-af83-09969a6ee94a · outbound

This paper cites Task contamination: Language models may not be few-shot anymore.

Towards Contamination Resistant Benchmarks Task contamination: Language models may not be few-shot anymore

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.263224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.263224Z digest=sha256:537f2419731948c154ebc7c3fc2c392aa83c972ff4c4f320ad2bd48c48f3a881

Observation c769f84a-a61c-44a5-a3c1-355cca4e40c3 · outbound

This paper cites Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages.

Towards Contamination Resistant Benchmarks Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.269455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.269455Z digest=sha256:3f837dc7ac7c81b2f7017b45a13e1d3c0861b7fb9d432b25c3c9035e421228d1

Observation 8a2aba1a-0999-4c81-b000-1a3fe0e4a1df · outbound

This paper cites an unresolved cited work.

Towards Contamination Resistant Benchmarks Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:00:11.202371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.274302Z digest=sha256:5ad257e664bd583e71ef4745506ca971b81dc616951f6ae5865d47704e260bdc

Observation 91be474c-debb-4978-af9e-e49d7067a69e · outbound

This paper cites an unresolved cited work.

Towards Contamination Resistant Benchmarks Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:00:11.188688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.280041Z digest=sha256:5bd6074623cadb8bb10bad2b7e98c36d3a4d4617d447a4fe63ccb361c7d0cb30

Observation f81a0d5a-71e7-493d-802d-61f26c9b7f61 · outbound

This paper cites Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation.

Towards Contamination Resistant Benchmarks Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.289031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.289031Z digest=sha256:86c9992dd5f86c33369a0f505447dcff475c0a49cd1bc08a6b73c8f13a772edf

Observation 3a828076-cf65-4803-bece-354477b49413 · outbound

This paper cites Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D.

Towards Contamination Resistant Benchmarks Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.293932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.293932Z digest=sha256:29c9ef8a0e9f9217e61d1c11bcb4a4b3239658b7a5fdd7fa7c48caf73dc6e66c

Observation 3fd22f3c-3528-4a6a-8fc5-7c039a60179a · outbound

This paper cites When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1.

Towards Contamination Resistant Benchmarks When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.298851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.298851Z digest=sha256:4d3572b758afd74c04796fffd27afb3c580f487453f35f3f32b32d02a79864f8

Observation d3fbf80f-7f20-4aa9-949e-88a247feca54 · outbound

This paper cites Sources of hallucination by large language models on inference tasks.

Towards Contamination Resistant Benchmarks Sources of hallucination by large language models on inference tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.303348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.303348Z digest=sha256:8139aebdff1fcf4fe3946239bce5ac5d4567d279db64bd90b3548fbfd2791a36

Observation 987d4032-2c32-437a-b2bd-6c8a72860bb0 · outbound

This paper cites Language Models Implement Simple Word2Vec-style Vector Arithmetic.

Towards Contamination Resistant Benchmarks Language Models Implement Simple Word2Vec-style Vector Arithmetic

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.309873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.309873Z digest=sha256:0d8be935e009762bba4284d83cede41fef56489e962d0946ba18e093655def1c

Observation 2681c611-dad7-4fec-a5f4-722a252c47b6 · outbound

This paper cites In-context learning generalizes, but not always robustly: The case of syntax.

Towards Contamination Resistant Benchmarks In-context learning generalizes, but not always robustly: The case of syntax

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.322626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.322626Z digest=sha256:8013d3ee53e59639eab798333b33cd78f671111b8da66bc84ea0e69e6c9acf8f

Observation cff4232b-5ded-4a76-84dd-241b01e606b0 · outbound

This paper cites Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation.

Towards Contamination Resistant Benchmarks Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.328149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.328149Z digest=sha256:f80e0a7566a7b83f75164c73922cef1d73cc234b4f8bc2b4ccd0123a06f66cff

Observation 67ed016a-5efd-4da6-bf20-2aa0df251489 · outbound

This paper cites Know what you don't know: Unanswerable questions for squad.

Towards Contamination Resistant Benchmarks Know what you don't know: Unanswerable questions for squad

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.333342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.333342Z digest=sha256:a82965998a765a0b363943d6f2b62814d3bdf11db0a7f7920d0262da7329d29a

Observation a9af2165-f0c5-4b2d-b9ec-99c587bbab03 · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

Towards Contamination Resistant Benchmarks A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.339578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.339578Z digest=sha256:72d369e836e28a6ad43157836ecb5e82d42e38bfe1f221d1056be91c9733731c

Observation 29ced63f-40b0-47cb-b9c7-00ca0ccff7b5 · outbound

This paper cites Prompt programming for large language models: Beyond the few-shot paradigm.

Towards Contamination Resistant Benchmarks Prompt programming for large language models: Beyond the few-shot paradigm

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.345705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.345705Z digest=sha256:46224600bb5a1f939f99ba583d3552e6256d30922ec309858e95dbbbbaa97a3f

Observation 42649257-1eac-41fb-8027-bb5b2e74d59e · outbound

This paper cites A natural experiment on LLM data contamination in code generation.

Towards Contamination Resistant Benchmarks A natural experiment on LLM data contamination in code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.174912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.353768Z digest=sha256:570e3137e5b27c584f6d73082f10564f07da37938f21f50e29e0ac36b88a838b

Observation efe9794c-e9ff-4b6f-9d06-2e5f1528e393 · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Towards Contamination Resistant Benchmarks NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.358528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.358528Z digest=sha256:ed70d44e5c826c74788a3f2ed749216e4bb22324c9a18686c554a0d5504a377f

Observation 1905e593-47d4-498b-b7b1-f4089f73099e · outbound

This paper cites Language models are greedy reasoners: A systematic formal analysis of chain-of-thought.

Towards Contamination Resistant Benchmarks Language models are greedy reasoners: A systematic formal analysis of chain-of-thought

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.157774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.364032Z digest=sha256:26614cc2aee4b403849679a46d4c7b36227578cf141852421c7abfc7ce9791bb

Observation a9aff1ad-87af-4268-86ff-04027412bfba · outbound

This paper cites an unresolved cited work.

Towards Contamination Resistant Benchmarks Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:00:11.145225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.369363Z digest=sha256:f92747dede2631c6d5102b13d9ed0fc802287a1535326507442224c5e85565b7

Observation d874dcdd-8913-499f-9040-60253a9f8e25 · outbound

This paper cites LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content.

Towards Contamination Resistant Benchmarks LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.374324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.374324Z digest=sha256:0c0831f69f26cdca46b6b4dbbda239d4746c900b6aff2b663d5a8b5e4e130ca8

Observation 206f4690-61c5-4bef-b395-d58e57f13858 · outbound

This paper cites Language models are multilingual chain-of-thought reasoners.

Towards Contamination Resistant Benchmarks Language models are multilingual chain-of-thought reasoners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.380705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.380705Z digest=sha256:8f1dea576681b285fb2af0f1fdbe26bb591f97444b00549a1233290a4d164720

Observation 8b42a115-268b-457f-8442-b96fa8ed7686 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Towards Contamination Resistant Benchmarks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.385552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.385552Z digest=sha256:f7b6e947884e3b6ba58185a57434f36e11df4b5f2dbf440ab5a02359e9e48614

Observation 4cd750b4-6f2e-423d-afa2-42bc77824238 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Towards Contamination Resistant Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.117221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.390704Z digest=sha256:962f5b628851209201c44b30da1230e05eb69eaf93f23dd822dc6cc7efccf35d

Observation 5405adbb-fdac-4e6b-98d9-d107edeadaeb · outbound

This paper cites an unresolved cited work.

Towards Contamination Resistant Benchmarks Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-15T22:00:10.618573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.395927Z digest=sha256:891e84522328d67d93d85ef8037445b2a1ee6b02b8fccf5b528cb1717ae6e2ee

Observation 8b42b87c-fe06-4142-b93a-bcc0cc850ca2 · outbound

This paper cites an unresolved cited work.

Towards Contamination Resistant Benchmarks Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:00:11.100511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.401979Z digest=sha256:f9e12c41fa6299b746209eafbc7e4d0966e418eb834cd5885217e45cae06ffd1

Observation 3d52a43f-9462-4e91-838d-1191469dd8d1 · outbound

This paper cites Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning.

Towards Contamination Resistant Benchmarks Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.410510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.410510Z digest=sha256:6db8992e501d406526697264d7ce37a05dbeceabd53bf0c4f559856f7cb10881

Observation 62cfa28b-8cf7-406a-ae08-9ff25a4369f9 · outbound

This paper cites Emergent analogical reasoning in large language models.

Towards Contamination Resistant Benchmarks Emergent analogical reasoning in large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.419376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.419376Z digest=sha256:4cf34e7603b37d06e41710d70b35be65f51219b40ec8ca79ea0e7851bf7bac52

Observation 1b51b3e6-bd18-48c8-809c-7c2f0eb35081 · outbound

This paper cites Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus.

Towards Contamination Resistant Benchmarks Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.071915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.423734Z digest=sha256:96b09e6844bc9284021d06e1d618380101f24650788977a8fe7433587a00a239

Observation b6c125aa-dd19-4f26-bcaa-b743fbf0b004 · outbound

This paper cites Chi, Quoc V.

Towards Contamination Resistant Benchmarks Chi, Quoc V

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.056355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.428533Z digest=sha256:d69792dffe25820fcd3c13588c753116c558dbf5bc9e7cf05c58397d837e8aab

Observation 1e998387-12eb-473e-b624-66cee9e76a4d · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Towards Contamination Resistant Benchmarks LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.433948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.433948Z digest=sha256:9fadfef15bd4f42a24a24b90605985f4bf0253b2a8e32489d9c36d5a3878839c

Observation b5a4af3f-3c19-4c74-80d1-433faf371b24 · outbound

This paper cites Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts.

Towards Contamination Resistant Benchmarks Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:00:11.038296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T22:00:10.443014Z digest=sha256:33e6295a57be24271bc68cb0a23d8aa3634e892497abde2150182975e65c8c0b

Observation 7c7ec1f2-e068-407e-8440-e7518e97d44d · outbound

This paper cites Qwen2.5 Technical Report.

Towards Contamination Resistant Benchmarks Qwen2.5 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.450042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.450042Z digest=sha256:588ec169d5054d5b02e04cb1ebaf3ff788f997fb4caf69db59f398c61cfa9102

Observation 95ca64ea-383a-474a-a83e-d399c0dea817 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Towards Contamination Resistant Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.454387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.454387Z digest=sha256:d8bc02e9d2aea3da08bbdcd724dced8a348e15de5fc47c513ded113d40be39b0

Observation d833ab05-622d-4d38-8aed-abd04f39f60f · outbound

This paper cites How Language Model Hallucinations Can Snowball.

Towards Contamination Resistant Benchmarks How Language Model Hallucinations Can Snowball

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.460187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.460187Z digest=sha256:56694c12e0fbbc62c4eb6d3ed230bd55f70dc8db3e96c425beed842d2e756e7a

Observation 0c9024f9-24d6-4bf0-b6c3-8fde42d56ac7 · outbound

This paper cites Trained Transformers Learn Linear Models In-Context.

Towards Contamination Resistant Benchmarks Trained Transformers Learn Linear Models In-Context

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.464998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.464998Z digest=sha256:d18028166c95f7875559ecb691b85903dccd86730aaa6c6dc235aeb201fa7d8a

Observation ff9803c0-7e34-4f03-955c-2ea6bd322445 · outbound

This paper cites Why Does ChatGPT Fall Short in Providing Truthful Answers?.

Towards Contamination Resistant Benchmarks Why Does ChatGPT Fall Short in Providing Truthful Answers?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:00:10.470063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:00:10.470063Z digest=sha256:e1b4632f31de697f5064c305216c22ccff32598d4770479290b59eed2cd92c13

Pith citing papers

Observation eeba85aa-00bb-43c1-af8b-210a0fb75436 · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Towards Contamination Resistant Benchmarks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:23:07.043308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:85c893588cb89f1cce8bfc9cd27cfdb7dafa20eac34d42e9beb41ca21626ae04