REVIEW 2 major objections 3 minor 47 references
This paper claims that in automated research, a single experimental score is unreliable evidence about an idea: re-implementing the same frozen idea reversed the one-draw winner in 25.6–43.6% of decisions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 12:45 UTC pith:UJVASX6Z
load-bearing objection A genuinely new measurement target with a careful design; the exact reversal rates are provider-conditional, but the direction is solid and the paper deserves serious refereeing. the 2 major comments →
One Run Is Not an Idea: The Implementation Lottery in Automated Research
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that idea reliability—the stability of an idea's ranking across plausible implementations—is a distinct quantity from best-of-N artifact utility. In its audit, implementation variance dominated same-artifact rerun variance by a factor of five to ten, and the winner selected by any single implementation often disagreed with the winner under the other-two implementation mean. The authors call this the implementation lottery and propose the Idea Reliability Audit as a way to measure it, reporting idea ICC and leave-one-implementation-out winner reversal.
What carries the argument
The central object is the 'idea card': a frozen statement of a mechanism that fixes the scientific intervention while leaving model, feature, preprocessing, and parameter choices open. Three fresh sessions translate each card into an implementation, an outcome-blind fidelity review checks adherence, and saved artifacts are rerun byte-identically. A balanced nested analysis of variance splits within-task score variation into idea, implementation, and rerun components, yielding an idea ICC; a leave-one-implementation-out comparison then tests whether a one-draw winner survives the other-two mean.
Load-bearing premise
The three fresh-session implementations of each frozen idea are treated as independent samples from the distribution of plausible implementations, but the provider did not expose reproducible sampling seeds, so independence rests on session nonces and identities rather than verified random draws.
What would settle it
Run the same audit with a coding agent that exposes reproducible sampling seeds; if leave-one-out winner reversal largely disappears when draws are certified independent, the measured lottery is an artifact of session dependence rather than a property of ideas.
If this is right
- If an automated research system uses a single score to prune a branch, transfer an idea, or store a result, that decision can be an artifact of which implementation was sampled.
- Best-of-N search remains valid for delivering a high-quality artifact, but it cannot justify mechanism-level claims about the underlying idea.
- Reliability is process-specific: the two coding setups did not show a consistent ordering, so each implementation process needs its own audit.
- Portfolio validity is a separate gate: outcome-blind review of candidate cards is needed before interpreting scores as evidence about ideas.
- The audit checklist—unit, process, adherence, noise, evidence—provides a concrete reporting standard for idea-level evidence.
Where Pith is reading between the lines
- If the reversal rates generalize beyond the audited tasks and provider, many existing automated-research results that credit single runs to ideas may be substantially overfit to implementation details, suggesting a correction analogous to multiple-seed reporting but applied at the idea-translation level.
- The same logic extends beyond machine learning to any automated pipeline where a specification is translated into execution—data-analysis plans, wet-lab protocols, simulation workflows—so reliability should be measured at the translation step.
- A natural testable extension: increasing the number of implementations per card should reduce reversal probability at a predictable rate if the variance decomposition is correct; the paper does not report this scaling curve.
- Since the two setups reversed ordering across task phases, a practical policy might adaptively request more implementations when candidates are close or implementation variance is high, rather than fixing a universal number of draws.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the 'implementation lottery': in automated research loops, a single run scores one implementation of an idea, and that realization-level score is often treated as evidence about the parent idea. The authors propose an 'Idea Reliability Audit' that freezes mechanism-level idea cards, samples three fresh-session implementations per card, labels fidelity outcome-blind, and reruns saved artifacts to separate implementation variance from same-artifact rerun variance. Across 312 assignments on 13 tabular tasks with two coding-agent setups, they report that implementation variance exceeds rerun variance by roughly 5x (Bounded) and 10x (Agentic), and that one-draw LOO winner reversal occurs in 25.6% and 43.6% of decisions. They also report an exploratory materials-regression diagnostic. The paper argues that idea-level reliability should be measured separately from best-of-N artifact utility, and it discusses implications for automated research pipelines.
Significance. If the empirical magnitudes are accepted, this is an important measurement contribution for the growing field of automated research systems. The audit design is careful in several ways: outcome-blind card freezing, fidelity labels with constructed controls, intention-to-treat retention of failed or drifting runs, preregistered penalty grids, leave-one-task-out estimates, cluster bootstraps, and a separate exploratory materials domain. The paper is also honest about conditionality, explicitly stating in Section 7 that the reversal rates are conditional on the task pool, implementation processes, and the absence of sampling seeds. The central limitation is that the three implementation draws per card come from a single provider without reproducible seeds, so the variance decomposition and reversal rates may be provider-specific rather than idea-level. This does not invalidate the concept, but it means the headline magnitudes need a sensitivity analysis before they can be treated as a general estimate of implementation lottery severity.
major comments (2)
- [§4, Eq. (1)–(3)] The variance decomposition and LOO reversal treat the three fresh-session implementations per card as independent/exchangeable draws from the distribution of plausible implementations. Section 4 states that 'the provider did not support sampling seeds, so independence rests on fresh sessions, nonces, and response/session identities rather than reproducible model sampling.' Fresh sessions and unique identifiers establish physical separation of contexts, not statistical independence. Because all draws come from the same model and provider, they may share session-order effects, prompt-template effects, or provider-side sampling policies. If so, the method-of-moments estimate of sigma^2_M and the reversal rates in §5 are provider- and session-process-specific rather than idea-level. The paper acknowledges conditionality in §7, but the empirical magnitude is the paper's central contribution.
- [§5, Figure 1] The claim that implementation variance is 'more than five and ten times' same-artifact rerun variance is reported as a point estimate without uncertainty, despite the fact that the rerun component is estimated from only three seeds per artifact and variance-component ratios are notoriously unstable when the denominator is small. Please report bootstrap or cluster-bootstrap confidence intervals for the ratio, and state explicitly whether the ratio uses zero-bounded or raw negative variance estimates. If the ratio cannot be estimated with acceptable precision, soften the claim to 'larger than' with the observed interval.
minor comments (3)
- [§4] The statement 'Task 18 then has three rather than four families' is unclear in relation to the earlier claim of 'Four accepted cards per task.' If Task 18 has fewer than four cards, the total of 312 assignments should be reconciled; if it still has four cards but only three distinct families, that should be stated explicitly.
- [§5] The sentence 'The equal two-to-four cards per task preclude the registered balanced method-of-moments ICC' appears after the analysis has already used the balanced assumption for the all-13 estimates. Clarify that this sentence refers to the filtered candidate sets in Table 1, not to the primary unfiltered analysis.
- [§4] The phrase 'Those identifiers verify physical independence for 311 of 312 assignments' may be misread as verifying statistical independence. Recommend rewording to 'physical separation of sessions' to avoid conflating the two notions, especially given the caveat about sampling seeds in the same paragraph.
Circularity Check
No circular derivation: the variance decomposition and LOO reversal are estimated from data, and the paper's self-citations are contextual. Seed-independence is an acknowledged validity caveat, not a circular reduction.
full rationale
The paper's headline estimates are empirical measurements, not derived predictions. Eq. (1)-(3) define a variance-components model; the finding that implementation variance exceeds rerun variance by 5x/10x could have been otherwise and is not forced by the model's definition. The LOO comparison is explicitly labeled a stability reference ('This other-two mean is a finite-budget reference, not ground truth'), so the reversal rate is a descriptive decision-stability statistic rather than a fitted parameter being called a prediction. No fitted parameter is used to generate a target quantity; fidelity labels and semantic gates are outcome-blind, and same-artifact reruns provide an independent control. The main caveat, that fresh-session draws rest on 'fresh sessions, nonces, and response/session identities rather than reproducible model sampling' (Section 4), is a threat to the independence/exchangeability assumption and therefore to how far the rates generalize; it is acknowledged in Section 7 ('25.6-43.6% is conditional rather than an error rate for arbitrary projects') and does not make the estimate equal to its input by construction. Self-citations (Ning et al. 2026a,b; Ning, Li, and Yu 2026) appear as related-work examples of automated research pipelines and are not load-bearing evidence for any of the audit's estimates. The core claim is not tautological: if implementation draws were identical to reruns, the measured ratios would have been near one and reversal would have disappeared.
Axiom & Free-Parameter Ledger
free parameters (8)
- ITT penalty for evaluation failure/invalidity =
-0.10 (primary); grid {-0.05,-0.10,-0.20}
- ITT penalty for drift/missing fidelity =
-0.02 (primary); grid {0,-0.02,-0.05}
- Number of implementation draws per card =
3
- OpenML task eligibility thresholds =
<=8000 rows, <=60 features, <=10 classes, imbalance <=10, headroom >=0.03
- Materials dielectric-task exclusion =
excluded
- SDK max_turns =
20
- Baseline hyperparameters =
lr=.08, 150 iters, 31 leaves, min_leaf 20, l2=1.0
- LOO alignment h =
1,2,3
axioms (8)
- standard math Exchangeability and zero-mean of random terms in Eq. (1) across task, idea, implementation, rerun levels
- standard math Balanced nested ANOVA subtraction of rerun mean square from implementation mean square yields unbiased variance components
- domain assumption LLM judges (DeepSeek v4 Flash/Pro) provide valid fidelity labels
- domain assumption DeepSeek v4 Pro card generation and implementations have no access to hidden outcomes
- domain assumption Relative held-out accuracy vs. a single HistGradientBoosting baseline is a meaningful utility for idea-level decisions
- domain assumption Same-provider fresh sessions yield statistically independent implementation draws
- ad hoc to paper The other-two mean is a valid reference for winner stability
- ad hoc to paper Zero-bounding negative variance estimates does not bias ICC materially
read the original abstract
Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.
Figures
Reference graph
Works this paper leans on
-
[1]
Lu, Chris and Lu, Cong and Lange, Robert Tjarko and Foerster, Jakob and Clune, Jeff and Ha, David , year =. The. 2408.06292 , archivePrefix=
-
[2]
Yamada, Yutaro and Lange, Robert Tjarko and Lu, Cong and Hu, Shengran and Lu, Chris and Foerster, Jakob and Clune, Jeff and Ha, David , year =. The. 2504.08066 , archivePrefix=
-
[3]
Towards End-to-End Automation of
Lu, Chris and Lu, Cong and Lange, Robert Tjarko and Yamada, Yutaro and Hu, Shengran and Foerster, Jakob and Ha, David and Clune, Jeff , journal =. Towards End-to-End Automation of. 2026 , doi =
2026
-
[4]
Accelerating Scientific Discovery with
Gottweis, Juraj and Weng, Wei-Hung and Daryin, Alexander and others , journal =. Accelerating Scientific Discovery with. 2026 , doi =
2026
-
[5]
Ayg. An. Nature , volume =. 2026 , doi =
2026
-
[6]
Agent Laboratory: Using
Schmidgall, Samuel and Su, Yusheng and Wang, Ze and Sun, Ximeng and Wu, Jialian and Yu, Xiaodong and Liu, Jiang and Moor, Michael and Liu, Zicheng and Barsoum, Emad , booktitle =. Agent Laboratory: Using. 2025 , doi =
2025
-
[7]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback , author =. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages =. 2025 , doi =
2025
-
[8]
Nathani, Deepak and Madaan, Lovish and Roberts, Nicholas and Bashlykov, Nikolay and Menon, Ajay and Moens, Vincent and Plekhanov, Mikhail and Budhiraja, Amar and Magka, Despoina and Vorotilov, Vladislav and Chaurasia, Gaurav and Hupkes, Dieuwke and Cabral, Ricardo Silveira and Shavrina, Tatiana and Foerster, Jakob Nicolaus and Bachrach, Yoram and Wang, Wi...
Pith/arXiv arXiv 2025
-
[9]
2025 , url =
Zhang, Yunxiang and Khalifa, Muhammad and Bhushan, Shitanshu and Murphy, Grant and Logeswaran, Lajanugen and Kim, Jaekyeom and Lee, Moontae and Lee, Honglak and Wang, Lu , booktitle =. 2025 , url =
2025
-
[10]
2025 , url =
Chen, Hui and Xiong, Miao and Lu, Yujie and Han, Wei and Deng, Ailin and He, Yufei and Wu, Jiaying and Li, Yibo and Liu, Yue and Hooi, Bryan , booktitle =. 2025 , url =
2025
-
[11]
Garikaparthi, Aniketh and Patwardhan, Manasi and Cohan, Arman , year =. 2602.15112 , archivePrefix=
-
[12]
2026 , eprint =
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes , author =. 2026 , eprint =
2026
-
[13]
2026 , eprint =
Closed-loop Auto Research for Molecular Property Prediction: Discovering and Certifying Generalizable Improvements , author =. 2026 , eprint =
2026
-
[14]
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-
Ning, Jingjie and Li, Xueqi and Yu, Chengyu , year =. Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-. 2604.01029 , archivePrefix=
-
[15]
2026 , eprint =
Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search , author =. 2026 , eprint =
2026
-
[16]
Towards Execution-Grounded Automated
Si, Chenglei and Yang, Zitong and Choi, Yejin and Cand. Towards Execution-Grounded Automated. Proceedings of the 43rd International Conference on Machine Learning , year =. 2601.14525 , archivePrefix =
-
[17]
2026 , eprint =
How Far Are We From True Auto-Research? , author =. 2026 , eprint =
2026
-
[18]
Towards a Science of
Rabanser, Stephan and Kapoor, Sayash and Kirgis, Peter and Liu, Kangheng and Utpala, Saiteja and Narayanan, Arvind , booktitle =. Towards a Science of. 2026 , eprint =
2026
-
[19]
When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for
Mehta, Aman , year =. When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for. 2602.11619 , archivePrefix=
-
[20]
2025 , eprint =
Stochasticity in Agentic Evaluations: Quantifying Inconsistency with Intraclass Correlation , author =. 2025 , eprint =
2025
-
[21]
2026 , eprint =
Rollout Cards: A Reproducibility Standard for Agent Research , author =. 2026 , eprint =
2026
-
[22]
Zhao, Xuanle and Sang, Zilin and Li, Yuxuan and Shi, Qi and Zhao, Weilun and Wang, Shuo and Zhang, Duzhen and Han, Xu and Liu, Zhiyuan and Sun, Maosong , booktitle =. 2026 , address =. doi:10.18653/v1/2026.acl-long.1001 , url =. 2505.20662 , archivePrefix =
Pith/arXiv arXiv 2026
-
[23]
International Conference on Learning Representations , year =
From Reproduction to Replication: Evaluating Research Agents with Progressive Code Masking , author =. International Conference on Learning Representations , year =. 2506.19724 , archivePrefix =
-
[24]
Si, Chenglei and Yang, Diyi and Hashimoto, Tatsunori , booktitle =. Can
-
[25]
The Ideation--Execution Gap: Execution Outcomes of
Si, Chenglei and Hashimoto, Tatsunori and Yang, Diyi , booktitle =. The Ideation--Execution Gap: Execution Outcomes of
-
[26]
and Liang, Weixin and Sun, Fan-Yun and Haber, Nick , booktitle =
Hua, Tianyu and Hua, Harper and Xiang, Violet and Klieger, Benjamin and Truong, Sang T. and Liang, Weixin and Sun, Fan-Yun and Haber, Nick , booktitle =. 2025 , eprint =
2025
-
[27]
Zhu, Minjun and Xie, Qiujie and Weng, Yixuan and Wu, Jian and Lin, Zhen and Yang, Linyi and Zhang, Yue , year =. 2506.01372 , archivePrefix=
-
[28]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Deep Reinforcement Learning That Matters , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2018 , doi =
2018
-
[29]
Proceedings of Machine Learning and Systems , volume =
Accounting for Variance in Machine Learning Benchmarks , author =. Proceedings of Machine Learning and Systems , volume =
-
[30]
Advances in Neural Information Processing Systems , volume =
Deep Reinforcement Learning at the Edge of the Statistical Precipice , author =. Advances in Neural Information Processing Systems , volume =
-
[31]
2021 , eprint =
The Benchmark Lottery , author =. 2021 , eprint =
2021
-
[32]
Improving Reproducibility in Machine Learning Research: A Report from the
Pineau, Joelle and Vincent-Lamarre, Philippe and Sinha, Koustuv and Larivi. Improving Reproducibility in Machine Learning Research: A Report from the. Journal of Machine Learning Research , volume =
-
[33]
Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =
Measurement and Fairness , author =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , doi =
2021
-
[34]
Advances in Neural Information Processing Systems , volume =
Measuring What Matters: Construct Validity in Large Language Model Benchmarks , author =. Advances in Neural Information Processing Systems , volume =
-
[35]
arXiv preprint arXiv:2107.03374 , year =
Evaluating Large Language Models Trained on Code , author =. arXiv preprint arXiv:2107.03374 , year =
-
[36]
Psychological Bulletin , volume =
Intraclass Correlations: Uses in Assessing Rater Reliability , author =. Psychological Bulletin , volume =. 1979 , doi =
1979
-
[37]
2001 , doi =
Generalizability Theory , author =. 2001 , doi =
2001
-
[38]
and Bischl, Bernd and Torgo, Luis , journal =
Vanschoren, Joaquin and van Rijn, Jan N. and Bischl, Bernd and Torgo, Luis , journal =. 2014 , doi =
2014
-
[39]
Benchmarking Materials Property Prediction Methods: The
Dunn, Alexander and Wang, Qi and Ganose, Alex and Dopp, Daniel and Jain, Anubhav , journal =. Benchmarking Materials Property Prediction Methods: The. 2020 , doi =
2020
-
[40]
Nature , volume =
Scientific Discovery in the Age of Artificial Intelligence , author =. Nature , volume =. 2023 , doi =
2023
-
[41]
Nature , volume =
Mathematical Discoveries from Program Search with Large Language Models , author =. Nature , volume =. 2024 , doi =
2024
-
[42]
2025 , eprint =
Novikov, Alexander and V. 2025 , eprint =
2025
-
[43]
International Conference on Learning Representations , year =
Self-Consistency Improves Chain of Thought Reasoning in Language Models , author =. International Conference on Learning Representations , year =
-
[44]
Zheng, Lianmin and Chiang, Wei-Lin and Sheng, Ying and others , booktitle =. Judging
-
[45]
Proceedings of the 39th International Conference on Machine Learning , series =
Model Soups: Averaging Weights of Multiple Fine-Tuned Models Improves Accuracy without Increasing Inference Time , author =. Proceedings of the 39th International Conference on Machine Learning , series =. 2022 , url =
2022
-
[46]
Computational Materials Science , volume =
Matminer: An Open Source Toolkit for Materials Data Mining , author =. Computational Materials Science , volume =. 2018 , doi =
2018
-
[47]
Transactions on Machine Learning Research , year =
Holistic Evaluation of Language Models , author =. Transactions on Machine Learning Research , year =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.