Pith. sign in

Paper Citation Record · LEDGER

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation

As of 20 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.00612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00612 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:42:08.508522Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:44:34.737273Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:44:37.557080Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2aeb9ec-465f-4ccf-9404-aee96c36798b · outbound

This paper cites Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.081999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.081999Z digest=sha256:35117020bc7498915d57319f7fe5d7ec241fe727ffa0a38f307cfb2671c6aacc

Observation b8c22902-2b58-4f99-91ee-1e07904347a6 · outbound

This paper cites State of M achine L earning C ompetitions in 2024.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation State of M achine L earning C ompetitions in 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.797985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.088128Z digest=sha256:ef2c6a90b97b3c975e9cacc7964774e83744a19e9f6c14dbf31097e2db8ca1cb

Observation 82f00423-27f5-41d9-800c-9eb480426413 · outbound

This paper cites O'Reilly Media, Inc.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation O'Reilly Media, Inc

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.729022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.095670Z digest=sha256:e3b99f4332e4d7eb80959e4186108d8762bd3eff7fa9e7bc93bb994f91541caf

Observation 7ebf4d11-2256-49a4-b587-9a7808d78ef6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.102876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.102876Z digest=sha256:fcdae444b80c7fbd937ed711e50069b800de730f743ef85543d34d487b1cec55

Observation 0f7d2c28-8a9d-4356-b339-7d1903077d25 · outbound

This paper cites E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation E., Stoica, I., Mooney, P., Dane, S., Howard, A., and Keating, N

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.683405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.109984Z digest=sha256:8bc722762ef6fae3001317a7e3ddfa302cefa6acf710781a4593f0e93d427217

Observation e68b148a-181b-48f2-b4c9-82895ace4329 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.121864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.121864Z digest=sha256:272dfef4a4358f6d3c1ce07500703a448bea7bc87562c4f044ae0391be2fcaaa

Observation b3b0a38d-5db8-4d84-9fa4-dda913daa44c · outbound

This paper cites On the Measure of Intelligence.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation On the Measure of Intelligence

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.132335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.132335Z digest=sha256:db1246eb2afd4252b7a304a8607e136c83de8e897fbdfa638087f892b6e17e23

Observation 415b1356-0d77-421b-bbde-11119672a83b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.141550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.141550Z digest=sha256:511adb22bf258d934904854bf1e3ee2a96b1d9c97a8504ebdbd72eba003f58ef

Observation 061cb8f5-012f-4a8c-b48c-d6412bce6b4f · outbound

This paper cites and Ghemawat, S.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Ghemawat, S

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.149789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.149789Z digest=sha256:36a7a09c2bdb05554986e4565feb852c0d0e6d816815c77a55e54975de8813c6

Observation 2ed25cbb-504d-4591-8fe5-ed9d98e14a30 · outbound

This paper cites ImageNet: A large-scale hierarchical image database.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation ImageNet: A large-scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.156677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.156677Z digest=sha256:53da5384cb9fcab5aaae3af71683ce4e23e283cdd4a38f91d6bd11e12ae617f9

Observation 6b9bd959-6eaa-4d55-959c-b870e52b30fb · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:42:10.644886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.166238Z digest=sha256:45544545b214c6fc1eae7f04af2d4e88856bed33e7cd81f591d9c954e078a564

Observation d6542c68-c1fe-46f2-a63a-b1ef12074db1 · outbound

This paper cites A Large Scale Benchmark for Uplift Modeling.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A Large Scale Benchmark for Uplift Modeling

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.611087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.185543Z digest=sha256:f3cbd6fd87f2fb9e31b74efb14f35646148669324799ede331f9e17840589e66

Observation 3039d003-c610-4879-88f8-ff0b0efa5590 · outbound

This paper cites Open LLM Leaderboard v2.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Open LLM Leaderboard v2

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.573477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.192499Z digest=sha256:ba7c7b105eb081cb74b7510d564436c603c93d2a72af3826e962a6dd1353bea9

Observation 4a35a751-5357-46ef-b0b1-dd12084dc449 · outbound

This paper cites D., Piovesan, D., Joshi, P., Reade, W., and Howard, A.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation D., Piovesan, D., Joshi, P., Reade, W., and Howard, A

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.540691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.201437Z digest=sha256:1eb25ab3093b08148734ec0ec0030f64c5f053028bf23463d489276a7a6f94a4

Observation 7fa1e65d-f378-43e1-a8ee-0c87c72f442d · outbound

This paper cites C., Buzzard, K., Gowers, T., Liu, P.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation C., Buzzard, K., Gowers, T., Liu, P

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.477635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.208659Z digest=sha256:a2b3e50c84b3f6b8cb4f5d1635560c8106e02e6e5bbbf93170899553ad21f6d2

Observation ecb7b5e6-4fda-4e80-806c-fbdf6f736bf5 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.218905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.218905Z digest=sha256:a3e517aa8fe0769ced1bcaefb7684dc0af38315cd9da2db79e3aa66f64d0bbd9

Observation 12a1988d-e785-4bf5-a6c2-7a2ccb1401d3 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.230671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.230671Z digest=sha256:96cd53a305387c75cfad37bc7adc64971c685b01a41fa86b2e9b1ca46cdb2869

Observation f7fb3a0c-6d0d-452f-93d8-23d5fd6eb86c · outbound

This paper cites The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.238393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.238393Z digest=sha256:0b26f56137c7620b2ff7868bb4570a9cbf08851ea9784b6b2e00760bc0dd1d8d

Observation 7906ca23-2851-497f-816b-5a5468c61c4e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.252331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.252331Z digest=sha256:f07abeccc93044da0e6e9ce36908d1db7db78970d8247441b2b929ba34d752a7

Observation efc8e7ee-bd8f-4701-94a0-4f1a66af230b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.262560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.262560Z digest=sha256:40c47c7973624c2d875d369df92058dfdf4e15951cdeafb807fa78e7881a15e8

Observation f1369f21-5136-4284-b4cf-e02846435563 · outbound

This paper cites Leakage in data mining: Formulation, detection, and avoidance.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Leakage in data mining: Formulation, detection, and avoidance

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.270368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.270368Z digest=sha256:8055d53a6e115909c7ca648b8bb316f984c773ff91e38d4d17d40d2c8d823331

Observation b4d33c3f-3cd8-443f-9735-a8f000cbda54 · outbound

This paper cites UCI Machine Learning Repository , 2025.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation UCI Machine Learning Repository , 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.442617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.278562Z digest=sha256:4051200708ded0cef03d0709f5f03a89481933b5d0458d8678d044f687b90b18

Observation a204bfe4-6f92-4eb7-916a-b074161a39ec · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:42:10.415860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.285617Z digest=sha256:3809dfafbc1650f3138ba1c4578bc9dce766927b48aa818d0f5cc7fcfef03c8e

Observation 82492e54-9fa5-4fa5-b004-27e834579321 · outbound

This paper cites and Cortes, C.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Cortes, C

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.392751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.293799Z digest=sha256:ae066792f746e29eeecaf461c90fc4dc46cb1e92b23b0e3fe1cdda2f35b1d596

Observation 34574055-789f-4d08-9b61-199df12ddf14 · outbound

This paper cites and Schwartz, R.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Schwartz, R

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.303084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.303084Z digest=sha256:fe0cd42777b8c61a1daa331d75a75c0d9c48fb76d934099d58a2fdbc680c8004

Observation 1ca12d1b-cdca-4582-a023-cb2c9624a50b · outbound

This paper cites P., Santorini, B., and Marcinkiewicz, M.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation P., Santorini, B., and Marcinkiewicz, M

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.311535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.311535Z digest=sha256:38af6b627adbc57a102ede0ef6c714fb7f344ef4013749772269d8d7568379ef

Observation 16091625-c137-4813-94aa-7a29072cd3bf · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:42:10.331147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.319394Z digest=sha256:49ca0005ead6713f938c37ae5643e1b9bde79d24163f2e3b097bc9b5f890b2cd

Observation a8a7229e-80ca-477f-8557-cc29adfd875c · outbound

This paper cites MTEB: Massive Text Embedding Benchmark.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation MTEB: Massive Text Embedding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.326197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.326197Z digest=sha256:9950ebf15eae02c072bfe534d5d80109e8c91da3d4475479f19b6aa47b7d5667

Observation 41a3d596-f9a1-4e58-a666-da80db14dfe8 · outbound

This paper cites Handbook of Statistical Analysis & Data Mining Applications.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Handbook of Statistical Analysis & Data Mining Applications

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.306032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.333661Z digest=sha256:85bc29312d5c83a41ae6d0835dbfd3cfad69f1f7fe1f80762e35259371fe0521

Observation 33eb3a42-c2ef-4607-8528-c3b79516b80f · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Proving Test Set Contamination in Black Box Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.343263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.343263Z digest=sha256:a5b2d37c15818103bfaa9b3bab45dc082d057429ffb97b68062a4128454ca28f

Observation f342500f-f49f-4cf5-bc1c-f57a0279e4ec · outbound

This paper cites Humanity's Last Exam.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Humanity's Last Exam

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.351149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.351149Z digest=sha256:37901c256c37e7e8d900d49813b22d1bca5521602e2b718404adb31c3f18de6c

Observation 59e285ab-e312-4e59-9494-0b9075291cc1 · outbound

This paper cites Google - Fast or Slow? Predict AI Model Runtime.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Google - Fast or Slow? Predict AI Model Runtime

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.278179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.365145Z digest=sha256:fb91188902c8e440a6ff5d43730c59082f75a8b35a7444728705b89b21837520

Observation 3b2a0c61-722e-4c8a-9551-a574d5c54fdc · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.373337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.373337Z digest=sha256:78e67bcd52f3db64401753cd1be3851f50425b34f922983d57c83829893cf9ec

Observation cb1ae38a-d777-4a88-a27e-1eb2fc478ce7 · outbound

This paper cites Do I mage N et classifiers generalize to I mage N et? In Chaudhuri, K.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Do I mage N et classifiers generalize to I mage N et? In Chaudhuri, K

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.249400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.379693Z digest=sha256:5fae757aa6792e69628e26eeeb419e0c80d0dbf86189a2d22d7d0b6944ba7d8c

Observation 07a3dcd8-7ba4-453f-ae53-d662b2700175 · outbound

This paper cites and Bozsolik, T.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation and Bozsolik, T

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.194279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.390323Z digest=sha256:b8ec43fba716fe59a8a2585453feca4aeba60f6e37e25e6e89fc6cceb4acbcbc

Observation db495b5d-8d47-4a52-a407-88b6b60da1f9 · outbound

This paper cites LANL Earthquake Prediction.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation LANL Earthquake Prediction

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.157638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.395723Z digest=sha256:c8227632185ed61765afd44d1feb1f5fbc7f98f58446b91cf56d6a268e21f950

Observation 0cf042a5-352d-4e5f-8b3e-3d232ed99d5c · outbound

This paper cites A meta-analysis of overfitting in machine learning.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A meta-analysis of overfitting in machine learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.124189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.402748Z digest=sha256:b8d8a7926e737cd75c09691db2cf149debc351d3f41f5eacdf092d2be144b371

Observation f409989d-4ae7-4024-9435-224faaa3b9b0 · outbound

This paper cites A meta-analysis of overfitting in machine learning.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation A meta-analysis of overfitting in machine learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.084979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.408640Z digest=sha256:160894dfaaa801537fc65097d0026b3a22d34934b11c2e5b443e4040cb89ead7

Observation b4242be1-6b05-4837-ad09-0d76456679fc · outbound

This paper cites L., and Agirre, E.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation L., and Agirre, E

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.413784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.413784Z digest=sha256:73505596c19a137357e050d0ab30eb39daf105a7e5e4c54c08679aa0bfe87680

Observation b39736dd-54cb-4f6b-b51b-e53c17d526e1 · outbound

This paper cites SEAL leaderboards.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation SEAL leaderboards

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.061612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.422269Z digest=sha256:35e6a267ad9f0fd1a3d4d2588a1fa2e0fdf004e8ea949ed11778a3dd49579b54

Observation aac54d00-ccb9-42b8-bfdc-b4a79d74e741 · outbound

This paper cites D., Reade, W., Wang, S., Croft, S., and Chen, Y.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation D., Reade, W., Wang, S., Croft, S., and Chen, Y

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:10.034426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.429994Z digest=sha256:9bb885ef87320f76df5838de69297c27756c5f392cc2b985c89077e4e884eb74

Observation 5aed9029-79b1-487a-8801-687ffb7185be · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.434933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.434933Z digest=sha256:3a2533d0d91a9f92051c931fac066ac7faf13adfa3614f27a2bb4b963b82789b

Observation f09942ce-a8d0-485c-9ec6-1353a71bc167 · outbound

This paper cites OpenML: networked science in machine learning.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation OpenML: networked science in machine learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.445356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.445356Z digest=sha256:3a6a59eb07e366a94fd4a791eafc622856385b6296a8840ad65b212b3d64ce65

Observation a810f462-966e-4a59-9110-7ceadd7ac6d3 · outbound

This paper cites The Nature of Statistical Learning Theory.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation The Nature of Statistical Learning Theory

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:09.987542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.450973Z digest=sha256:111a3cdb50809a95a039f344644fe7fc2af9276e70ef7dfab58e408e51f7742c

Observation 0adbc3c5-f0ea-4086-85ac-2af304c952d4 · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 45

Resolution
verified exact
doi, observed 2026-08-16T04:42:08.589818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.456200Z digest=sha256:4f414bdab572c2beb9b4a7c4b3a4034aff56ba952af600deab7ad3cba49d4a20

Observation 9a7f3ba9-417e-4b42-8156-ca648e361734 · outbound

This paper cites S., Naidu, S.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation S., Naidu, S

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:09.962603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.464070Z digest=sha256:9996c720d89ae8a9f10ca0bdda33d5f3bee83e8032fa1bd87210462c61b27c2f

Observation 303ebe59-835d-4a75-bf91-db8cea792572 · outbound

This paper cites AI mathematical olympiad - progress prize 1.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation AI mathematical olympiad - progress prize 1

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:09.924177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.474508Z digest=sha256:de5b724b74b647577c8d64f43343ca1618c689270a151afc96cf60c1abebbb59

Observation 40debefe-96ae-4249-b31c-8418f9165a83 · outbound

This paper cites TalkingData AdTracking Fraud Detection Challenge.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation TalkingData AdTracking Fraud Detection Challenge

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:42:09.873976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.479708Z digest=sha256:024344af18dcf7fd662add60b386cdbc360f6fa9a0b4497719a6a37005f5e6ea

Observation 8414c02b-b2c3-471f-bc89-cc7e97aa1fea · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.486427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.486427Z digest=sha256:f1c94b84646a64a2eed086afba993949c42c25d80b346100f5ec78574ba3e941

Observation c61e55a1-867c-46cd-82d9-3907d7d1a70b · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:42:08.497577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:42:08.497577Z digest=sha256:8c066f528dbe8ea3f02dddae1d697d93d2a2568410041bcd95e67b1643bc71f3

Observation 5fc7db32-4344-48f5-8769-50b9212333d2 · outbound

This paper cites an unresolved cited work.

Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:42:09.836041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-16T04:42:08.508522Z digest=sha256:aaa629d48fc0179dd75e79dae0fecdb2555df3e14ce3143b2336a4e47dc8c3a3

Pith citing papers

Observation fd3fd3f4-6ca8-4869-8edb-d911de07a0e1 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:44:37.671156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:44:34.737273Z digest=sha256:6c9b68bea6a6e7526a1ba08a350ce351cd4390377a16776bcd419dbcea60ae70

Observation d47b433a-ecdd-4c60-bf5f-c86b2e713985 · inbound

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security cites this paper.

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security Position: AI Competitions Provide the Gold Standard for Empirical Rigor in GenAI Evaluation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:08.637388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:19:08.637388Z digest=sha256:ddb2761d32ee4679857d9a0c2b52802a651628c2c4e6d9083d76a98436370d25