Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:29.070019Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2505.24263.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:29.070019Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 48403ae3-c436-4c4d-8b1b-b0882126a52f · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552d78f0-3bb1-4696-8fe9-f0b5a18c4c1a · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381ca8e0-ebc8-45b7-beb2-c466d9774410 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d579d1d-a646-4908-bd15-bd401705277f · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac6a04c4-0f96-44ff-a4d0-5d758d64bf20 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Language Models are Few-Shot Learners
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d2be1ac-92ae-4e22-9908-3792e513e538 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Quantifying Memorization Across Neural Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1faed743-79ba-435e-a9d9-a5ba3cb31840 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb46aeb-d547-4ee8-8969-fc01656f5fca · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d447cf9b-c72d-4059-a0f0-3344f8a8dc83 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Does Data Contamination Detection Work (Well) for LLMs? A Survey and Evaluation on Detection Assumptions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b4366e9-05f8-49cd-a94a-145685687d6d · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Are We Done with MMLU?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d8c614-fee6-4973-98cc-704b3e9632b4 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Time Travel in LLMs: Tracing Data Contamination in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f148745-1121-4a19-8c0b-296d37a97674 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f015e85b-9744-4278-af9d-5a08ebfe5715 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f872dafe-c61b-4cd0-b962-c71c5ba91f46 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Measuring Massive Multitask Language Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa753c5-b81e-466b-9385-c969a3c49ecf · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f23ba1cc-33d7-4198-8548-4725c998a439 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LoRA: Low-Rank Adaptation of Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 088ddac3-4d2c-4796-97d3-8db2db9abe7b · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Membership Inference Attacks on Machine Learning: A Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a55477a-f0cc-4719-8e26-4ed2979c2896 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0af4880-be2f-4f7f-b237-dd4aed103c85 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aca29502-3b06-4411-898b-37b2eb2c6e40 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Lin, Jacob Hilton, and Owain Evans
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e29cbdeb-17ed-44ca-be0f-87b2bd22759a · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation DeepSeek-V3 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537db04c-fc95-4991-b246-5bc604ff91bc · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 022bc4e4-6bca-4107-a61d-e4d2afd41e0d · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation GPT-4 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85948a63-eefd-4c53-8708-a7d5933f2fbb · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457c0c24-c0e0-4a64-bfc4-7e22fc3024c4 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Qwen2.5 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d45bf4-cb03-4b7e-b8af-fae0dddd9161 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Leveraging Large Language Models for Multiple Choice Question Answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1645894-ad5c-4459-87c3-38c1f24c1685 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a2cc5c-414e-4f72-93f1-84b947cfecfd · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation SocialIQA: Commonsense Reasoning about Social Interactions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877f8036-0eb7-428d-b4b5-b2904ec2975f · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemini: A Family of Highly Capable Multimodal Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17509d7d-145e-4f20-b4c0-cafb062a0914 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma: Open Models Based on Gemini Research and Technology
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c169a4fc-480e-431e-9945-25576704a8e1 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Gemma 2: Improving Open Language Models at a Practical Size
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d931bec7-4dd3-40f8-b778-8aef1d04f872 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation LLaMA: Open and Efficient Foundation Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7265672e-7d0f-4ee7-aab8-3d9c9f00974d · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc895bca-1fb5-4f50-872f-432387d4c5bd · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd67ca9-a106-43a1-a3b5-d6a351c50b4f · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Benchmarking Benchmark Leakage in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe21523-ee1b-44fc-b0db-954c381c2c42 · outbound
Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.