{"id":"927cfb78-55fa-4b46-b29f-b8d685f72a80","arxiv_id":"2506.04734","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Evaluation conditions like seed, dataset version, and answer ordering cause multi-point benchmark score swings in DeepSeek-R1-Distill and related reasoning models, undermining reliable comparison.","lead":"This paper shows that small changes in how LLM benchmarks are run, such as the random seed, how many times a question is repeated, or the version of the dataset, can shift reported scores by several percentage points for popular open-source reasoning models. It argues that many claimed performance improvements may partly come from favorable evaluation setups rather than better models, and proposes a statistical rule for choosing the number of repetitions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never audits any specific derivative model's published benchmark claim, so the central attribution that reported gains reflect favorable evaluation design is asserted, not demonstrated; Table 4 even shows official GPQA scores below the control.","rationale":"The reader's strongest_claim has two components: (a) evaluation design shifts scores, and (b) published improvements are partially attributable to those shifts. Component (a) is supported by extensive controlled experiments; component (b) is the distinctive, title-level claim and is what would need direct evidence. The paper's own Table 4 provides the only official-versus-control comparison and, for GPQA, contradicts the inflation direction. The appendix tables give fluctuations for derivatives but no official claims or original eval configurations, so the causal step from 'scores vary' to 'claimed gains are inflated' is missing. This is a correctness risk in the argument, not a disagreement with consensus: the phenomenon of evaluation sensitivity is plausible and consistent with prior work, but the paper overreaches in attributing specific published results to favorable evaluation design without auditing them. The reader's weakest_assumption (single-rerun baseline) is a real statistical weakness, but it is secondary because the raw effect sizes (e.g., up to 16 pp for GPQA option bias, 3.9 pp for dataset version) do not depend on the baseline comparison. The concrete test above would settle whether the attribution claim holds; until then, the verdict should remain CONDITIONAL, requiring revision to either add the audit or soften the central claim.","tokens_in":21739,"tokens_out":10182,"duration_ms":119398,"concrete_test":"Take three derivative models with public eval scripts and claimed scores, e.g., Skywork-OR1-32B-Preview (AIME24/AIME25), Light-R1-14B-DS (AIME24/GPQA), and DeepScaleR-1.5B-Preview (AIME24). For each, run the official released evaluation script end-to-end and record the reproduced score; then run the same model on the paper's control configuration (N=64, dynamic seed, simplescaling dataset versions, instruction after question, TP as specified). If the official-script reproduction does not match the claimed score within the baseline fluctuation range, the reproducibility claim is directly confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the paper is not merely that evaluation conditions cause score fluctuations—that part is supported by the controlled comparisons—but that 'many claimed performance gains in open-source models are partially attributable to favorable evaluation setups rather than genuine model improvement' (Section 1). This attribution requires showing, for specific published results, that the original evaluation setup inflated the score relative to a neutral or standardized setup. The paper does not do this. It reports fluctuation magnitudes for derivative models (Tables 11–13) but never compares any derivative model's officially claimed score to a reproduction under the developer's original script versus the paper's control configuration. The introduction asserts that reproduction with original evaluation code was challenging, but no reproduction attempt, diff, or score table is presented. Moreover, the one place where official and control scores are both shown (Table 4) points in the opposite direction for GPQA Diamond: official results are 33.8/49.1/59.1/62.1 versus control 40.3/54.7/61.3/67.4, i.e., the official setup produced substantially lower scores, not inflated ones. Thus the headline concept of 'strategic overclaiming through evaluation design' is not established by the evidence, even though the underlying sensitivity phenomenon is credible. The reader's baseline-fluctuation concern is valid but secondary: a single-rerun baseline makes the 'exceeds baseline' percentages statistically fragile, but the raw multi-point differences do not depend on that baseline. The missing audit is the load-bearing gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a series of controlled evaluations of DeepSeek-R1-Distill models and derivative open-source reasoning models on AIME24, AIME25, and GPQA Diamond, varying the number of samples N, random seed strategy, dataset version, instruction position, option or answer ordering, and tensor parallelism. It finds that these incidental evaluation choices can move scores by up to several percentage points (e.g., 3.9 points on AIME25 dataset versions, over 5 points in GPQA option manipulations) and argues that many published performance improvements may be partially attributable to favorable evaluation design rather than genuine model gains. The authors propose an evaluation paradigm based on full transparency and statistically grounded stability, including a CLT-based method for choosing N and reporting confidence intervals.","tokens_in":21976,"tokens_out":4409,"duration_ms":53390,"significance":"If the empirical sensitivity results are correct, the paper makes a useful and timely contribution to LLM evaluation practice. The controlled comparisons cover many models, and the appendices include extensive per-run raw tables, which is a genuine reproducibility asset. The manuscript is also careful about documenting inference settings and reports baseline reruns. However, the paper's headline attribution claim—that reported performance gains are partially due to favorable evaluation setups—is not demonstrated by the data, and the proposed statistical protocol contains internal inconsistencies. These issues are fixable but currently weaken the central narrative.","major_comments":[{"comment":"The claim that 'many claimed performance gains in open-source models are partially attributable to favorable evaluation setups rather than genuine model improvement' is not established by the presented evidence. The paper reports fluctuation magnitudes for derivative models (Tables 11–13) but never audits a specific published result by rerunning the developer's original evaluation script and comparing it against a neutral configuration. The only direct comparison of official and control scores, Table 4, shows official GPQA Diamond results below the control group (e.g., 32B: 62.1 vs 67.4), which is opposite to the inflation narrative. The manuscript should either add concrete audits of specific claims or explicitly narrow the conclusion to sensitivity of benchmark scores to evaluation design.","section":"Section 1 and Table 4"},{"comment":"The baseline fluctuation is computed from a single repeated run, so statements such as 'over 75% of experiments exhibit deviations beyond the baseline fluctuation range' have no statistical content: a single-point baseline cannot define a range. The same issue affects the 'exceeds the baseline reference' statements in Sections 2.3, 2.4, and 2.7. Please rerun the control configuration multiple times (or model the sampling distribution) and report intervals, or replace the exceedance language with effect-size comparisons and confidence intervals.","section":"Section 2.2 and Table 1"},{"comment":"The 'Estimated Interval' is not derived in the text, and the reported intervals are internally inconsistent with the control-group means. For example, for DeepSeek-R1-Distill-Qwen-32B on AIME24 the interval is reported as 73.3±1 while the control mean is 71.8, and on AIME25 the interval is 53.9±1 while the control mean is 56.6. Moreover, because s in Eq. (3) is estimated from the same data used to construct the interval, the iterative procedure may stop prematurely; no convergence or coverage analysis is provided. The proposed N-calibration method should be validated experimentally (e.g., with split-sample or bootstrap checks) before being presented as a standard.","section":"Section 3.2, Eq. (2)–(3), Table 4"}],"minor_comments":[{"comment":"Please define what the ± value represents (standard error, margin of error, or confidence-interval half-width) and state how s was computed; the notation and values are currently opaque.","section":"Table 4"},{"comment":"The sentence 'over 70% of evaluation results stabilize beyond N = 8' is ambiguous because Table 2 reports counts of experiments that have already stabilized at each minimum N; clarify the aggregation and whether the counts are cumulative.","section":"Section 2.3 and Table 2"},{"comment":"Given that baseline fluctuations can reach 1.4 percentage points (Table 1), instruction-position differences below 2 points may be within noise; consider reporting confidence intervals for these comparisons.","section":"Section 2.5"},{"comment":"The table uses footnote markers and unusual parameter formatting (e.g., the QwQ-32B row showing '1 3' and footnote references), which makes it hard to verify the exact values; please reformat for readability.","section":"Appendix A.1, Table 5"},{"comment":"Several references are to Notion or Hugging Face URLs that may break or change, and one URL is split across lines; adding DOIs, version identifiers, or stable repository links would improve reproducibility.","section":"References"},{"comment":"The figures would benefit from error bars or per-condition point markers; currently the reader cannot assess the spread underlying the plotted fluctuations.","section":"Figures 1–7"}],"recommendation":"major_revision","confidential_remarks":"The empirical sensitivity results are credible and the raw tables in the appendix are a strong asset, but the title and introduction oversell the attribution to 'strategic overclaiming.' I would encourage the authors to reframe the contribution or add direct audits of specific published claims. The self-citations to TinyR1 and Light-R1 are relevant to the model inventory and do not appear to load the central claim. The scope fits an empirical AI/ML venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look. The empirical core is real: changing N, seed, dataset version, option order, or TP moves scores by several points on DeepSeek-R1-Distill models and derivatives. The paper does careful controlled runs and publishes extensive raw tables in the appendix, which is exactly what you want for this kind of claim. If you build eval harnesses or compare open-source reasoning models, this catalog of sensitivities is useful.\n\nThe title and the intro's 'strategic overclaiming' framing overreach. The paper never audits a single published benchmark claim. It shows that evaluation conditions can move scores by several points, but it does not show that any particular model's reported gain came from favorable settings. The one place where official and control scores appear together, Table 4, points the wrong way for GPQA: official results are lower than the control, not inflated. That does not kill the sensitivity story, but it undercuts the attribution story.\n\nThe statistical grounding is thinner than it looks. The 'baseline fluctuation' is one rerun of the control setup, no distribution, no interval. So statements like 'over 75% of experiments exceed the baseline' are comparisons to a single point, not proper statistical claims. Table 4's 'estimated interval' is unexplained and internally inconsistent—the control group mean falls outside the interval in several rows. That needs a rewrite.\n\nWhat is genuinely new here is the systematic application to the R1-Distill family and the dataset-version comparison, with up to 3.9 points of shift on AIME25. That is worth having. The proposal to report confidence intervals and calibrate N is sensible, though the CLT-based formula is standard statistics, not a new derivation.\n\nWho is this for? Practitioners who trust leaderboards and need to know how much noise to expect. It deserves a serious referee, but with major revision: either audit specific published claims or reframe the claim as 'evaluation sensitivity' rather than 'strategic overclaiming,' fix the baseline statistics, and clean up Table 4. As is, I would accept it for peer review but not for publication.","headline":"Useful empirical catalog of evaluation sensitivities for R1-Distill models, but the 'strategic overclaiming' attribution is asserted, not demonstrated, and the baseline statistics are fragile.","tokens_in":22594,"tokens_out":1827,"would_cite":true,"duration_ms":21379,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Evaluation design, not model quality, can move reasoning-benchmark scores by several percentage points.","keywords":["LLM evaluation","benchmark reproducibility","evaluation design","seed sensitivity","option order bias","DeepSeek-R1-Distill","confidence intervals","reasoning models"],"falsifier":"Run the same control configuration (N=64, dynamic seed) on AIME24, AIME25, and GPQA Diamond for the four DeepSeek-R1-Distill models many times, say 30 independent reruns, and plot the spread of control scores. If that spread routinely covers the fluctuation magnitudes attributed to dataset version, seed, or option ordering, the claim that these evaluation choices are the main cause of multi-point shifts would be falsified.","tokens_in":21534,"feed_emoji":"📊","tokens_out":5468,"duration_ms":61352,"temperature":0.7,"pith_summary":"The paper sets out to show that benchmark scores for reasoning-focused LLMs are far less objective than they look: small, often unreported evaluation choices — how many times a question is sampled, which random seed is used, which version of a dataset, where an instruction sits, how multiple-choice options are ordered, and how the model is parallelized — can move scores by several percentage points. It demonstrates this on the DeepSeek-R1-Distill series and derivative open-source models across AIME24, AIME25, and GPQA Diamond. The authors argue that many published \"improvements\" over the base models are partly artifacts of favorable evaluation setups, and that the community should report confidence intervals, reproduce baseline scores under identical settings, and choose the sample size N from a statistically grounded formula. If the claim holds, comparisons between reasoning models need standardized protocols before gains can be trusted.","feed_headline":"Evaluation design can swing LLM scores by several points","feed_subtitle":"Small changes in seeds, dataset versions, and answer order shift reasoning-benchmark scores by several percentage points.","key_machinery":"The argument is carried by systematic controlled perturbation of seven evaluation variables, measured as absolute score differences against a control configuration (N=64, dynamic seeds, a fixed dataset version, instruction after the question, correct answer in option A, and a fixed tensor-parallelism setting). The quantitative instrument for the proposed standard is a central-limit-theorem sample-size formula, $N \\geq (z_{\\alpha/2} \\, s / \\epsilon)^2$, which estimates how many repeated samples are needed so the reported mean lies within a chosen error margin $\\epsilon$ at confidence level $1-\\alpha$. The paper uses this formula to show that the required N differs by model and benchmark, so fixed values like N=16, 32, or 64 are not universally justified.","core_discovery":"The central discovery is that evaluation design, not model capability, can account for a large share of apparent performance differences among reasoning models. Using controlled comparisons on four sizes of DeepSeek-R1-Distill and on other open-source reasoning models, the paper finds that switching AIME dataset versions changes scores by up to 3.9 percentage points, that seed choice alone can let a small model match or beat a larger one, and that option order and correct-answer placement in GPQA Diamond produce swings above 5 percentage points, with several models moving by 10 points or more. These magnitudes exceed the repeated-run baseline fluctuation the paper measures, leading the authors to conclude that many officially reported gains are partially attributable to evaluation setup. The paper's proposed remedy is a two-principle paradigm: transparency about every evaluation condition, and stability, meaning reporting statistically grounded confidence intervals rather than peak scores.","pith_inferences":["If these magnitudes generalize, leaderboard ranks among closely matched reasoning models may be within the noise of evaluation design; a controlled protocol could matter more than the next incremental training gain.","The option-order results suggest that GPQA Diamond partly measures position heuristics; a testable extension is to report position-balanced scores as the official metric.","The proposed N formula could be applied to any benchmark to compute minimum evaluation sizes; a practical audit is to re-run published leaderboard configurations and see how many reported gains survive.","The paper's framing implies evaluation design can be strategically tuned; a natural follow-up is an automatic audit that detects undocumented favorable configurations in released evaluation scripts."],"forward_implications":["Reported benchmark gains of open-source reasoning models should not be trusted unless the evaluation script, dataset version, seed policy, N, option ordering, and hardware settings are disclosed.","Seed and option-order effects are large enough that a small model under a favorable configuration can appear to outperform a larger model on the same benchmark.","Benchmark scores that look like training improvements may instead be dataset-version or answer-position effects; direct comparison to baselines re-run under identical conditions is required.","A statistically chosen N, rather than a round number, should become the norm, with confidence intervals reported instead of peak scores.","Multiple-choice reasoning benchmarks should randomize or counterbalance option order, since correct-answer placement alone shifts results by over five points."],"supporting_citations":[{"why":"Supplies the DeepSeek-R1-Distill model family that is the primary subject of the evaluation perturbation study.","marker":"DeepSeek-AI, 2025"},{"why":"Provides the simplescaling AIME dataset versions used as the control group and as comparison variants.","marker":"Muennighoff et al., 2025"},{"why":"Defines the GPQA Diamond benchmark used for the option-order and answer-position bias experiments.","marker":"Rein et al., 2023"},{"why":"Establishes that LLMs are not robust multiple-choice selectors, the prior result the option-order experiments extend.","marker":"Zheng et al., 2023"},{"why":"Shows that inference parameters affect reasoning results and frames why the paper instead targets overlooked evaluation variables.","marker":"Hochlehnert et al., 2025"},{"why":"Documents the inference framework whose seed-handling behavior underlies the dynamic-seed versus fixed-seed experiments.","marker":"Kwon et al., 2023"},{"why":"Provides the evaluation framework used to run the controlled benchmark comparisons.","marker":"Sheng et al., 2024"}],"fun_headline_variants":["Evaluation design swings LLM scores by up to 10 points","How evaluation choices inflate LLM benchmark results","Tiny test tweaks can flip LLM model rankings","Seed and order shift LLM scores more than skill","Evaluation artifacts explain many LLM performance claims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"All the comparison claims depend on treating a single repeated run of the control configuration as the baseline fluctuation; that one number has no confidence interval, so statements that a fluctuation \"exceeds baseline\" or that \"over 75% of experiments\" do so are not statistically supported.","fun_headline_variants_meta":{"raw":{"variants":["Evaluation design swings LLM scores by up to 10 points","How evaluation choices inflate LLM benchmark results","Tiny test tweaks can flip LLM model rankings","Seed and order shift LLM scores more than skill","Evaluation artifacts explain many LLM performance claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1161,"prompt_tokens":849,"completion_tokens":312,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":236}},"tokens_in":465,"tokens_out":312,"duration_ms":4323,"temperature":1.0,"reasoning_tokens":236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:34:14.046685+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same control configuration (N=64, dynamic seed) on AIME24, AIME25, and GPQA Diamond for the four DeepSeek-R1-Distill models many times, say 30 independent reruns, and plot the spread of control scores. If that spread routinely covers the fluctuation magnitudes attributed to dataset version, seed, or option ordering, the claim that these evaluation choices are the main cause of multi-point shifts would be falsified.","supporting_citations":[{"cited_title":"Gonzalez, Hao Zhang, and Ion Stoica","cited_arxiv_id":null,"evidence_quote":"Documents the inference framework whose seed-handling behavior underlies the dynamic-seed versus fixed-seed experiments."}],"review_version":1}