Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:10.470063Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:2505.08389.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:00:10.470063Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T07:19:50.354875Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T07:23:07.041987Z
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69940e1e-cddd-4802-a9de-6a158853e6e8 · outbound
Towards Contamination Resistant Benchmarks What learning algorithm is in-context learning? investigations with linear models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb7d25df-a325-4e73-a1b0-4198f9551556 · outbound
Towards Contamination Resistant Benchmarks The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 679693a7-2f00-4adf-bb6c-f512b27a3bbb · outbound
Towards Contamination Resistant Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61240b48-40e3-4d10-b95e-e38734ebd56b · outbound
Towards Contamination Resistant Benchmarks Do, Yan Xu, and Pascale Fung
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1627ff-46a5-47f4-b407-c5ffef23943e · outbound
Towards Contamination Resistant Benchmarks Managing extreme ai risks amid rapid progress
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96c32ab6-d82b-4475-b5a9-ce0a2804a633 · outbound
Towards Contamination Resistant Benchmarks Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, et al
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a222112e-aace-4c1a-bd86-27d64ddfa00c · outbound
Towards Contamination Resistant Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9afb45fe-9725-4a00-953f-ba4af404d928 · outbound
Towards Contamination Resistant Benchmarks TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87c02f98-6247-419a-8ac1-f330f1090157 · outbound
Towards Contamination Resistant Benchmarks Palm: Scaling language modeling with pathways
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9b19019a-fbde-4bf2-9906-fbb9ef6bf8e6 · outbound
Towards Contamination Resistant Benchmarks Scaling Instruction-Finetuned Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a102519-49cd-4d08-8674-c0cfcf141338 · outbound
Towards Contamination Resistant Benchmarks Training Verifiers to Solve Math Word Problems
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67c3b58c-0549-4070-9353-333d8b900a68 · outbound
Towards Contamination Resistant Benchmarks Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0db4f8e0-25b6-4f1f-ae6c-6bb1978294c1 · outbound
Towards Contamination Resistant Benchmarks Generalization or memorization: Data contamination and trustworthy evaluation for large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9682585c-cd52-4507-93c8-ea1ec6a6b3c6 · outbound
Towards Contamination Resistant Benchmarks The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57bfe173-e544-4f93-879f-1cadb603ae5a · outbound
Towards Contamination Resistant Benchmarks Cole, Fangyu Liu, and William W
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 06e56a22-f3b8-4d18-943a-df202aaec5ab · outbound
Towards Contamination Resistant Benchmarks What can transformers learn in-context? A case study of simple function classes
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 78704c4c-d05c-4b92-b5ac-329530667238 · outbound
Towards Contamination Resistant Benchmarks Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in chatgpt
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9c551f1c-83c8-45a8-9a2d-e3c0714bbc48 · outbound
Towards Contamination Resistant Benchmarks Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3b71f0-7d93-4473-815c-50eda4a52d82 · outbound
Towards Contamination Resistant Benchmarks LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff868702-fe1a-423e-99fa-885499a739df · outbound
Towards Contamination Resistant Benchmarks Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation faa7495c-5684-4147-86ab-518b2fae3cc3 · outbound
Towards Contamination Resistant Benchmarks Large language models are zero-shot reasoners
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 120ad59c-f67d-498d-af83-09969a6ee94a · outbound
Towards Contamination Resistant Benchmarks Task contamination: Language models may not be few-shot anymore
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c769f84a-a61c-44a5-a3c1-355cca4e40c3 · outbound
Towards Contamination Resistant Benchmarks Language Ranker: A Metric for Quantifying LLM Performance Across High and Low-Resource Languages
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a2aba1a-0999-4c81-b000-1a3fe0e4a1df · outbound
Towards Contamination Resistant Benchmarks Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 91be474c-debb-4978-af9e-e49d7067a69e · outbound
Towards Contamination Resistant Benchmarks Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f81a0d5a-71e7-493d-802d-61f26c9b7f61 · outbound
Towards Contamination Resistant Benchmarks Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a828076-cf65-4803-bece-354477b49413 · outbound
Towards Contamination Resistant Benchmarks Thomas McCoy, Shunyu Yao, Dan Friedman, Mathew D
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd22f3c-3528-4a6a-8fc5-7c039a60179a · outbound
Towards Contamination Resistant Benchmarks When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3fbf80f-7f20-4aa9-949e-88a247feca54 · outbound
Towards Contamination Resistant Benchmarks Sources of hallucination by large language models on inference tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987d4032-2c32-437a-b2bd-6c8a72860bb0 · outbound
Towards Contamination Resistant Benchmarks Language Models Implement Simple Word2Vec-style Vector Arithmetic
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2681c611-dad7-4fec-a5f4-722a252c47b6 · outbound
Towards Contamination Resistant Benchmarks In-context learning generalizes, but not always robustly: The case of syntax
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff4232b-5ded-4a76-84dd-241b01e606b0 · outbound
Towards Contamination Resistant Benchmarks Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ed016a-5efd-4da6-bf20-2aa0df251489 · outbound
Towards Contamination Resistant Benchmarks Know what you don't know: Unanswerable questions for squad
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9af2165-f0c5-4b2d-b9ec-99c587bbab03 · outbound
Towards Contamination Resistant Benchmarks A Comprehensive Survey of Contamination Detection Methods in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29ced63f-40b0-47cb-b9c7-00ca0ccff7b5 · outbound
Towards Contamination Resistant Benchmarks Prompt programming for large language models: Beyond the few-shot paradigm
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42649257-1eac-41fb-8027-bb5b2e74d59e · outbound
Towards Contamination Resistant Benchmarks A natural experiment on LLM data contamination in code generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation efe9794c-e9ff-4b6f-9d06-2e5f1528e393 · outbound
Towards Contamination Resistant Benchmarks NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1905e593-47d4-498b-b7b1-f4089f73099e · outbound
Towards Contamination Resistant Benchmarks Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9aff1ad-87af-4268-86ff-04027412bfba · outbound
Towards Contamination Resistant Benchmarks Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d874dcdd-8913-499f-9040-60253a9f8e25 · outbound
Towards Contamination Resistant Benchmarks LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206f4690-61c5-4bef-b395-d58e57f13858 · outbound
Towards Contamination Resistant Benchmarks Language models are multilingual chain-of-thought reasoners
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b42a115-268b-457f-8442-b96fa8ed7686 · outbound
Towards Contamination Resistant Benchmarks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd750b4-6f2e-423d-afa2-42bc77824238 · outbound
Towards Contamination Resistant Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5405adbb-fdac-4e6b-98d9-d107edeadaeb · outbound
Towards Contamination Resistant Benchmarks Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8b42b87c-fe06-4142-b93a-bcc0cc850ca2 · outbound
Towards Contamination Resistant Benchmarks Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d52a43f-9462-4e91-838d-1191469dd8d1 · outbound
Towards Contamination Resistant Benchmarks Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62cfa28b-8cf7-406a-ae08-9ff25a4369f9 · outbound
Towards Contamination Resistant Benchmarks Emergent analogical reasoning in large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b51b3e6-bd18-48c8-809c-7c2f0eb35081 · outbound
Towards Contamination Resistant Benchmarks Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6c125aa-dd19-4f26-bcaa-b743fbf0b004 · outbound
Towards Contamination Resistant Benchmarks Chi, Quoc V
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1e998387-12eb-473e-b624-66cee9e76a4d · outbound
Towards Contamination Resistant Benchmarks LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a4af3f-3c19-4c74-80d1-433faf371b24 · outbound
Towards Contamination Resistant Benchmarks Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7c7ec1f2-e068-407e-8440-e7518e97d44d · outbound
Towards Contamination Resistant Benchmarks Qwen2.5 Technical Report
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ca64ea-383a-474a-a83e-d399c0dea817 · outbound
Towards Contamination Resistant Benchmarks A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d833ab05-622d-4d38-8aed-abd04f39f60f · outbound
Towards Contamination Resistant Benchmarks How Language Model Hallucinations Can Snowball
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9024f9-24d6-4bf0-b6c3-8fde42d56ac7 · outbound
Towards Contamination Resistant Benchmarks Trained Transformers Learn Linear Models In-Context
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9803c0-7e34-4f03-955c-2ea6bd322445 · outbound
Towards Contamination Resistant Benchmarks Why Does ChatGPT Fall Short in Providing Truthful Answers?
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeba85aa-00bb-43c1-af8b-210a0fb75436 · inbound
LLM Benchmark Datasets Should Be Contamination-Resistant Towards Contamination Resistant Benchmarks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.