Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:08.508522Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.00612.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:08.508522Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:34.737273Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:44:37.557080Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c2aeb9ec-465f-4ccf-9404-aee96c36798b · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8c22902-2b58-4f99-91ee-1e07904347a6 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation State of M achine L earning C ompetitions in 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 82f00423-27f5-41d9-800c-9eb480426413 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation O'Reilly Media, Inc
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ebf4d11-2256-49a4-b587-9a7808d78ef6 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f7d2c28-8a9d-4356-b339-7d1903077d25 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e68b148a-181b-48f2-b4c9-82895ace4329 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3b0a38d-5db8-4d84-9fa4-dda913daa44c · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation On the Measure of Intelligence
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 415b1356-0d77-421b-bbde-11119672a83b · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061cb8f5-012f-4a8c-b48c-d6412bce6b4f · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Ghemawat, S
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed25cbb-504d-4591-8fe5-ed9d98e14a30 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation ImageNet: A large-scale hierarchical image database
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b9bd959-6eaa-4d55-959c-b870e52b30fb · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d6542c68-c1fe-46f2-a63a-b1ef12074db1 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A Large Scale Benchmark for Uplift Modeling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3039d003-c610-4879-88f8-ff0b0efa5590 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Open LLM Leaderboard v2
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4a35a751-5357-46ef-b0b1-dd12084dc449 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation D., Piovesan, D., Joshi, P., Reade, W., and Howard, A
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7fa1e65d-f378-43e1-a8ee-0c87c72f442d · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation C., Buzzard, K., Gowers, T., Liu, P
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ecb7b5e6-4fda-4e80-806c-fbdf6f736bf5 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a1988d-e785-4bf5-a6c2-7a2ccb1401d3 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Measuring Massive Multitask Language Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7fb3a0c-6d0d-452f-93d8-23d5fd6eb86c · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7906ca23-2851-497f-816b-5a5468c61c4e · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efc8e7ee-bd8f-4701-94a0-4f1a66af230b · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1369f21-5136-4284-b4cf-e02846435563 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Leakage in data mining: Formulation, detection, and avoidance
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d33c3f-3cd8-443f-9735-a8f000cbda54 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation UCI Machine Learning Repository , 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a204bfe4-6f92-4eb7-916a-b074161a39ec · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 82492e54-9fa5-4fa5-b004-27e834579321 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Cortes, C
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 34574055-789f-4d08-9b61-199df12ddf14 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Schwartz, R
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca12d1b-cdca-4582-a023-cb2c9624a50b · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation P., Santorini, B., and Marcinkiewicz, M
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16091625-c137-4813-94aa-7a29072cd3bf · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a8a7229e-80ca-477f-8557-cc29adfd875c · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation MTEB: Massive Text Embedding Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41a3d596-f9a1-4e58-a666-da80db14dfe8 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Handbook of Statistical Analysis & Data Mining Applications
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 33eb3a42-c2ef-4607-8528-c3b79516b80f · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Proving Test Set Contamination in Black Box Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f342500f-f49f-4cf5-bc1c-f57a0279e4ec · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Humanity's Last Exam
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e285ab-e312-4e59-9494-0b9075291cc1 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Google - Fast or Slow? Predict AI Model Runtime
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b2a0c61-722e-4c8a-9551-a574d5c54fdc · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1ae38a-d777-4a88-a27e-1eb2fc478ce7 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Do I mage N et classifiers generalize to I mage N et? In Chaudhuri, K
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 07a3dcd8-7ba4-453f-ae53-d662b2700175 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Bozsolik, T
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation db495b5d-8d47-4a52-a407-88b6b60da1f9 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation LANL Earthquake Prediction
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0cf042a5-352d-4e5f-8b3e-3d232ed99d5c · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A meta-analysis of overfitting in machine learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f409989d-4ae7-4024-9435-224faaa3b9b0 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A meta-analysis of overfitting in machine learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b4242be1-6b05-4837-ad09-0d76456679fc · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation L., and Agirre, E
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39736dd-54cb-4f6b-b51b-e53c17d526e1 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SEAL leaderboards
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation aac54d00-ccb9-42b8-bfdc-b4a79d74e741 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation D., Reade, W., Wang, S., Croft, S., and Chen, Y
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5aed9029-79b1-487a-8801-687ffb7185be · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f09942ce-a8d0-485c-9ec6-1353a71bc167 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation OpenML: networked science in machine learning
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a810f462-966e-4a59-9110-7ceadd7ac6d3 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The Nature of Statistical Learning Theory
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0adbc3c5-f0ea-4086-85ac-2af304c952d4 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9a7f3ba9-417e-4b42-8156-ca648e361734 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation S., Naidu, S
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 303ebe59-835d-4a75-bf91-db8cea792572 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation AI mathematical olympiad - progress prize 1
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 40debefe-96ae-4249-b31c-8418f9165a83 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation TalkingData AdTracking Fraud Detection Challenge
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8414c02b-b2c3-471f-bc89-cc7e97aa1fea · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61e55a1-867c-46cd-82d9-3907d7d1a70b · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc7db32-4344-48f5-8769-50b9212333d2 · outbound
Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd3fd3f4-6ca8-4869-8edb-d911de07a0e1 · inbound
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d47b433a-ecdd-4c60-bf5f-c86b2e713985 · inbound
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.