{"id":"9697c1a0-e436-4e25-8123-eda99f3e4b18","arxiv_id":"2505.00612","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors argue that time-bound AI competitions with hidden test data are the gold standard for generative AI evaluation, and that public static benchmarks should be considered contaminated once published.","lead":"This paper argues that AI competitions, like those hosted on Kaggle, should be treated as the most reliable way to test generative AI models because test questions are kept secret until the deadline. It claims public static benchmarks are already leaked once their questions appear online, so their results should be trusted less.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.1's 'leaked the moment it is shared' rule is asserted, not established; the paper's own §2.2 evidence on public-benchmark stability cuts against it, so the impossibility claim is under-supported.","rationale":"The reader's weakest_assumption correctly identifies the Section 4.1 rule of thumb as load-bearing. My analysis sharpens this: the paper's own §2.2 citations are a direct empirical counterexample to the impossibility claim for public benchmarks, and the paper never reconciles that tension. The competition case studies (CAFA, AIMO, Konwinski) are real and support the value of time-bound, secret-test evaluation, and the paper gives useful structural arguments about parallelized evaluation and isolated code execution that deserve credit. However, because the normative conclusion is built on an unproven empirical premise about universal leakage, CONDITIONAL remains the appropriate verdict: the position is worth engaging with as a testable proposal, not as an established fact. No change to the reader's verdict is needed.","tokens_in":18215,"tokens_out":6378,"duration_ms":71125,"concrete_test":"Controlled leakage experiment: create ~200 fresh expert-authored questions (e.g., math/coding) split into matched private set P and public set Q. Evaluate k current LLMs on both at t0, then publish Q on a public repository for three months. At t1, after at least one new LLM checkpoint has been released, evaluate k models on both P and Q. Compute (a) the Q−P score gap over time and (b) Spearman rank correlation between P and Q rankings at t1; additionally apply n-gram/membership contamination detectors to Q and re-rank models on the flagged-clean subset. If the clean-subset ranking on Q closely matches the private ranking on P, the Section 4.1 impossibility claim is falsified and static benchmarks can be decontaminated; if not, the rule of thumb is supported. For the API part, audit a no-logging endpoint by probing whether any private prompts are later recoverable from the model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion—static public GenAI benchmarks cannot be robust, so competitions are the gold standard—depends on Section 4.1's rule that an evaluation is 'leaked the moment it has been shared online or sent over the wire.' The paper treats this as a practical certainty, but the support is a small set of selected Kaggle leakage anecdotes (SETI, TalkingData, LANL) plus the observation that LLM training data is hard to audit. Those anecdotes establish that leakage happens and is hard to avoid; they do not establish that public exposure generally invalidates GenAI model rankings, nor that contamination cannot be detected and filtered. The unresolved tension is with the paper's own §2.2, which cites Recht et al. (2019) and Roelofs et al. (2019) showing that heavily reused public benchmarks and public leaderboards preserved rank ordering on fresh data. If public exposure does not systematically corrupt rankings, or if decontamination is feasible, then static benchmarks can remain valid, and time-bound secret-test evaluation is one valuable option rather than the unique gold standard.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper, authored by Kaggle employees, argues that traditional IID-based static benchmarking is fundamentally inadequate for generative AI, and that the dominant problem is leakage and contamination. The authors assert that an evaluation is effectively leaked the moment its test data has been shared online or sent over the wire, and that consequently no published static benchmark can be robust to leakage. They argue that AI competitions, which combine time-bound secret tests, parallel independent attempts, and anti-cheating mechanisms, provide the gold standard for empirical rigor in GenAI evaluation, and they recommend that the field shift from reproducible static benchmarks to repeatable processes and value meta-analyses of competition results. The paper supports these claims with Kaggle leakage case studies, a review of existing leak-avoidance benchmark designs, and descriptions of competition structures that use prospective ground truth, novel task generation, and post-deadline data collection.","tokens_in":18380,"tokens_out":5099,"duration_ms":49632,"significance":"If the paper's central impossibility claim were established, it would be a significant and timely intervention in how the field evaluates generative models. The paper's real strengths are its concrete catalog of leakage mechanisms from a decade of Kaggle competitions (SETI, TalkingData, LANL, Predict AI Model Runtime), and its explicit, actionable descriptions of leak-resilient competition designs such as CAFA 5, AIMO, the WSDM Cup, and the Konwinski Prize. These examples provide a useful practical blueprint, and the paper is honest enough to discuss alternative viewpoints. The fundamental weakness is that the load-bearing empirical premise—that any shared evaluation is irrevocably leaked and that a static published benchmark cannot be trustworthy—is asserted rather than demonstrated, and it sits in tension with the paper's own citation of Recht et al. (2019) and Roelofs et al. (2019a,b), which show rank-order stability under heavy public reuse. As a position statement it is valuable, but the central generalization currently outruns the evidence.","major_comments":[{"comment":"The rule of thumb that an evaluation is 'leaked the moment it has been shared online or sent over the wire' and the ensuing claim that 'we simply cannot have a published static benchmark that is robust to leakage' are load-bearing for the paper's conclusion, but they are supported only by a small set of Kaggle leakage anecdotes. The case studies in Section 4 (SETI, TalkingData, LANL, and the AI Model Runtime competition) involve file metadata, row ordering, randomization seeds, and synthetic-signal artifacts; they do not establish that LLM results on static benchmarks are generally invalidated by contamination, nor do they rule out detection and filtering of contaminated examples. This is an overgeneralization from a selected sample, and since the 'gold standard' recommendation depends on the impossibility claim, it needs substantially stronger support or a clearly qualified formulation.","section":"Section 4.1"},{"comment":"The paper's own Section 2.2 cites Recht et al. (2019), which showed that ImageNet rank ordering was preserved on brand new data despite massive reuse, and Roelofs et al. (2019a,b), which showed that public leaderboard performance on Kaggle competitions was a strong indicator of private holdout rank ordering. That evidence directly cuts against the universal 'leaked the moment shared' rule, which treats any public exposure as catastrophically invalidating. The authors should either explain why GenAI evaluation is qualitatively different in ways that make contamination destroy the validity of comparisons (e.g., unbounded output spaces, memorization), or soften the impossibility claim to a claim about individual model scores rather than rank ordering. Without such engagement, the central thesis is internally inconsistent with evidence the paper itself presents.","section":"Section 2.2 vs 4.1"},{"comment":"The paper claims that AI competitions provide 'the gold standard for empirical rigor' but never defines the criteria for that designation, nor does it systematically compare competitions against the alternatives reviewed in Section 5 (unreleased holdout sets, dynamic benchmarks, and community benchmarks). The possible weaknesses of competitions—such as task-selection bias, incentive effects of prizes, deadline-driven submission strategies, and the fact that leakage also occurs in competitions, as Section 4 recounts—are not weighed transparently against the claimed benefits. For a prescriptive position paper, the comparison should be made explicit and the term 'gold standard' operationally defined (e.g., contention-resistance, statistical power, ecological validity, and reproducibility).","section":"Section 6.2 and 8"},{"comment":"The asserted 'fundamental tension' between reproducibility and robustness, likened to the Heisenberg uncertainty principle, is presented rhetorically rather than argued. The text offers no proof or evidence that no reproducible static benchmark can also be robust to leakage. The hybrid FACTS Grounding Leaderboard discussed in Section 5.1, which publishes half the test set while keeping the other half private precisely to allow reproducibility checks, suggests a tradeoff with design options rather than an impossibility. The claim should be demoted from a universal impossibility to a practical difficulty or supported with an explicit argument.","section":"Section 4.1"}],"minor_comments":[{"comment":"The paper should include an explicit conflict-of-interest statement: the authors are Kaggle employees and the paper argues for increased use of AI competition results, so readers would benefit from a transparent disclosure beyond the affiliation line.","section":"Author affiliations"},{"comment":"The claim that ensembling submissions 'obtains little-to-no improvement to top ranked solutions' is presented without data or citation; either provide the underlying analysis or qualify it as an informal observation.","section":"Section 8"},{"comment":"The phrase 'gold standard' is used as a conclusion rather than a defined criterion; a brief operational definition would make the position more testable and reduce the risk of circular reasoning.","section":"Abstract and Section 1"},{"comment":"The quoted Nectar dataset description contains 'Antropic/hh-rlhf', which appears to be a typo; please mark it as '[sic]' or correct the quotation.","section":"Section 3.2"},{"comment":"The FACTS Grounding Leaderboard is mentioned as a hybrid approach but only briefly; one or two sentences on how the public/private split is used would help the reader understand the reproducibility-robustness tradeoff the paper discusses.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is authored by Kaggle employees and makes a recommendation that benefits the platform's perceived importance, which is a potential conflict of interest that the public text does not explicitly flag beyond the affiliation line. This is not disqualifying, but an explicit statement would help. The reader's stress-test concern about Section 4.1 lands: the 'leaked the moment shared' rule is under-supported and inconsistent with the paper's own Section 2.2 evidence. If the authors are willing to soften the impossibility claim into a practical-risk claim and add a systematic comparison of alternatives, the paper could be publishable as a position statement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a position piece from current Kaggle employees arguing that AI competitions should be treated as the gold standard for GenAI evaluation. It is not a research result, and it does not pretend to be. The core argument is that GenAI evaluation needs novelty-centric generalization rather than IID-style splits, that leakage and contamination are the dominant threats, and that the time-bound, parallelized, secret-test structure of competitions is the most practical way to meet those threats. That argument is coherent, and the paper does some things well: it names the reproducibility-versus-robustness tradeoff explicitly, it proposes a clear rule of thumb (an evaluation is leaked the moment it is shared or sent over the wire), and it gives concrete, instructive examples of leakage from Kaggle competitions. It also discloses the authors' affiliation in the abstract and acknowledgments, which is more than many papers do. I found the discussion of prospective ground truth, post-deadline data collection, and novel task generation genuinely useful, and the call for meta-analyses across evaluations is a good one.\n\nThe soft spots are real but not fatal. The central claim that static public benchmarks are fundamentally invalid rests on Section 4.1's rule of thumb, which is asserted rather than established. The supporting anecdotes show leakage happens and is hard to avoid, but they do not show that public exposure always corrupts model rankings. That matters because the paper itself cites Recht et al. and Roelofs et al. in Section 2.2 showing that heavily reused benchmarks and public leaderboards preserved rank ordering on fresh data. The paper never resolves that tension, and the impossibility claim is stronger than the evidence warrants. The evidence for competition effectiveness is also selected from the authors' own platform, which makes the 'gold standard' phrasing feel rhetorical even if the underlying practices are sound. I would like to see the claim softened to something like 'competitions are one of the most robust available mechanisms' and the failure modes of competitions themselves discussed more frankly. None of this sinks the paper; it is a position paper, and it does its job of laying out a defensible position with concrete mechanisms.\n\nWho should read it: anyone working on LLM evaluation methodology, benchmark design, or contamination. It deserves a serious referee and would benefit from a revision that either weakens the impossibility claim or engages head-on with the Recht/Roelofs evidence. I would not cite it as an empirical result, but I would cite it as a clear statement of a design philosophy.","headline":"A clearly argued, well-disclosed position paper that makes a real case for competitions, but overreaches when it claims static benchmarks can never be trusted.","tokens_in":780,"tokens_out":912,"would_cite":true,"duration_ms":21290,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI Competitions, not static benchmarks, should set the standard for GenAI evaluation.","keywords":["generative AI evaluation","data leakage","benchmark contamination","LLM evaluation","AI competitions","novelty-centric generalization","static benchmarks","evaluation robustness"],"falsifier":"Construct a matched pair of evaluations from the same source—one a long-public static benchmark, the other a fresh set written after model training cutoffs—and run the same models on both; the claim is falsified if rank ordering and score gaps on the public benchmark closely reproduce on the fresh set across many tasks and models.","tokens_in":17963,"feed_emoji":"🏆","tokens_out":7244,"duration_ms":70206,"temperature":0.7,"pith_summary":"Machine learning's traditional paradigm of a fixed training set and a hidden test set, reused as a public benchmark, assumes test examples are drawn from the same distribution as training data and stay unseen. The paper argues that this paradigm breaks for generative models, whose input and output spaces are effectively unbounded and whose outputs feed back into context. It therefore treats leakage and contamination as the central dangers, adopting the rule of thumb that an evaluation is leaked the moment it is shared online or sent to a model over the wire. From that premise it concludes that public static benchmarks cannot be trusted for GenAI and that time-bound AI competitions, which test many models in parallel on data that does not exist at training time, provide the strongest available empirical rigor.","feed_headline":"Static GenAI benchmarks are invalid once shared","feed_subtitle":"New paper: only time-bound competitions with hidden test data can yield trustworthy GenAI comparisons.","key_machinery":"The central object is the AI Competition, defined by the paper as a task with an objective evaluation function in which independent teams make parallel, time-bound attempts. Its load-bearing feature is that test data is kept secret and, in the strongest cases, does not exist during the training phase—via prospective ground truth, novel task generation, or post-deadline data collection. The paper also places the leakage rule of thumb—'leaked the moment it has been shared online or sent over the wire'—as the mechanism that invalidates static benchmarks.","core_discovery":"Generative AI cannot be reliably evaluated with the traditional train/test-split benchmark, because GenAI models have nearly unbounded input and output spaces, no well-defined ground truth, and feedback loops that break the IID assumption; and because any evaluation data that is shared online or sent to a model is, by the paper's rule of thumb, already leaked. The paper's central claim is that the field should therefore regard static public benchmarks as invalidated once published and treat AI Competitions—time-bound, parallel, independently attempted tasks with hidden test data—as the gold standard for empirical rigor. It argues that competition structures such as prospective ground truth, novel task generation, and post-deadline data collection can create leak-proof evaluations, and that reproducibility should be sacrificed for robustness by replacing immutable benchmarks with repeatable processes.","pith_inferences":["The paper leaves implicit that high-stakes model choices, such as deploying a model in medicine or law, should not be made from public leaderboard scores once the leakage rule of thumb is accepted.","A testable extension of the paper's position: competition platforms could publish periodic 'contamination audits' by measuring model probabilities on held-out competition data before and after public release, giving an empirical handle on the leakage rule.","The same time-bound, secret-test structure could be ported to agentic and tool-use evaluations, where leakage through environment data is even harder to detect than in text benchmarks.","If the rule of thumb holds, benchmark designers should shift effort from bigger static question sets to renewable pipelines and agreements with API providers not to train on evaluation traffic."],"forward_implications":["Once an evaluation is published online or sent to a model, treat its results as potentially contaminated; static benchmark scores should carry much less weight in GenAI comparisons.","Evaluation should be organized as simultaneous, time-bound attempts on novel tasks, so that every model sees the test material for the first time at the same moment.","In cases where reproducibility and robustness conflict, robustness wins: better a one-time trustworthy result than a repeatable result that may be contaminated.","The field should invest in meta-analyses of competition results, synthesizing methods and outcomes across many tasks rather than trusting any single static benchmark.","Competition-style anti-cheating structures—hidden test data, trusted offline execution, and post-deadline data collection—should become the model for general GenAI evaluation."],"supporting_citations":[{"why":"Defines the novelty-based measure of intelligence that motivates the paper's centrality of novel-task evaluation.","marker":"Chollet (2019)"},{"why":"Meta-analysis showing that benchmark reuse did not produce overfitting in classical ML, used to argue leakage rather than overfitting is the critical risk.","marker":"Roelofs et al. (2019b)"},{"why":"Re-evaluation of ImageNet models on freshly collected data shows rank-order stability, the evidence that static benchmarks worked under IID assumptions.","marker":"Recht et al. (2019)"},{"why":"Supplies the formal treatment of leakage in data mining that frames the paper's leakage and contamination discussion.","marker":"Kaufman et al. (2012)"},{"why":"Documents contamination and evaluation malpractice in closed-source LLMs, supporting the 'leaked once shared' rule.","marker":"Balloccu et al. (2024)"},{"why":"Proves test-set contamination can be detected in black-box language models, evidence that contamination is real and measurable.","marker":"Oren et al. (2023)"},{"why":"CAFA 5 uses prospective ground truth, serving as the paper's demonstration that leak-proof competitions are feasible.","marker":"Friedberg et al. (2023)"},{"why":"LiveBench is the dynamic-benchmark approach the paper contrasts with competitions, showing current leakage mitigation and its limits.","marker":"White et al. (2025)"},{"why":"Konwinski Prize uses post-deadline data collection, a second feasibility proof for leak-proof competition structure.","marker":"Konwinski et al. (2024)"},{"why":"OpenVaccine results improved state of the art and generalized to unseen data, evidence that competition outcomes carry real-world validity.","marker":"Wayment-Steele et al. (2022)"}],"fun_headline_variants":["AI competitions are the gold standard for GenAI evaluation","Static GenAI benchmarks leak; hidden-test competitions don't","For trusted GenAI results, use time-bound competitions","Why AI competitions outshine static GenAI benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument stands or falls on the empirical premise that any evaluation data shared online or sent over the wire to a model is effectively leaked and cannot be kept out of training data; if some shared test sets remain effectively unknown, or contamination can be detected and filtered, static benchmarks could stay valid and competitions lose their uniquely protected status.","fun_headline_variants_meta":{"raw":{"variants":["AI competitions are the gold standard for GenAI evaluation","Static GenAI benchmarks leak; hidden-test competitions don't","For trusted GenAI results, use time-bound competitions","Why AI competitions outshine static GenAI benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1412,"prompt_tokens":876,"completion_tokens":536,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":473}},"tokens_in":492,"tokens_out":536,"duration_ms":5384,"temperature":1.0,"reasoning_tokens":473,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:37:35.477253+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct a matched pair of evaluations from the same source—one a long-public static benchmark, the other a fresh set written after model training cutoffs—and run the same models on both; the claim is falsified if rank ordering and score gaps on the public benchmark closely reproduce on the fresh set across many tasks and models.","supporting_citations":[{"cited_title":"Do I mage N et classifiers generalize to I mage N et? In Chaudhuri, K","cited_arxiv_id":null,"evidence_quote":"Re-evaluation of ImageNet models on freshly collected data shows rank-order stability, the evidence that static benchmarks worked under IID assumptions."},{"cited_title":"D., Piovesan, D., Joshi, P., Reade, W., and Howard, A","cited_arxiv_id":null,"evidence_quote":"CAFA 5 uses prospective ground truth, serving as the paper's demonstration that leak-proof competitions are feasible."},{"cited_title":"S., Naidu, S","cited_arxiv_id":null,"evidence_quote":"LiveBench is the dynamic-benchmark approach the paper contrasts with competitions, showing current leakage mitigation and its limits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Konwinski Prize uses post-deadline data collection, a second feasibility proof for leak-proof competition structure."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"OpenVaccine results improved state of the art and generalized to unseen data, evidence that competition outcomes carry real-world validity."}],"review_version":1}