REVIEW 4 major objections 5 minor 21 references
Today's frontier AI agents can finish the engineering of open-ended AI research but still cannot answer the research questions at a publishable standard.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 10:53 UTC pith:PMIYGK5J
load-bearing objection Useful new eval design plus honest negative case studies; the n=2 non-blind bridge is thin but the paper already says so, and the released rejects/logs still earn a serious read. the 4 major comments →
Can AI agents conduct open-ended AI research? Early evidence from two case studies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Well-resourced frontier agents can autonomously perform the engineering steps of open-ended AI research, yet cannot make substantial progress toward answering those research questions at top-conference quality. In two shadow evaluations on unpublished NeurIPS submissions, original authors rejected both agent-written papers, citing weak experimental judgment, thin novelty, and unclear writing despite successful end-to-end execution.
What carries the argument
Shadow evaluations: assign a frontier agent the central open-ended research question of a high-quality unpublished paper, give it multi-day wall-clock time and large API and GPU budgets on a general-purpose scaffold, then have the paper's original authors grade the agent's output as a conference submission.
Load-bearing premise
That unambiguous rejection by the original authors on two non-blind case studies is a reliable enough signal of agent research inability, rather than mainly author preference, scaffold under-elicitation, or the particular questions chosen.
What would settle it
A shadow evaluation on a comparable unpublished top-conference question where a frontier agent under similar time and compute budgets produces a paper the original authors score as at least a weak accept, or a larger series of such runs that consistently clear that bar.
If this is right
- Claims that AI will soon automate AI R&D must separate engineering automation from open-ended scientific judgment and redesign.
- Verifier-only and blind-review evaluations will systematically miss the failure modes that shadow evaluations expose.
- More wall-clock time or raw compute alone is unlikely to fix premature commitment, weak backtracking, and unused budgets.
- Self-review loops that reliably reject weak drafts do not by themselves force creative redesign of the research approach.
- Released logs, reviews, and repositories make the method reusable as stronger models and scaffolds appear.
Where Pith is reading between the lines
- If engineering keeps improving while judgment, backtracking, and resource sense lag, recursive self-improvement may accelerate narrow optimization faster than open-ended discovery.
- Because both agent papers were rejects, the study cannot yet tell whether AI reviewers are truly discriminative or merely harsh—closing that generator-verifier gap would need accepted as well as rejected drafts.
- Instruction drift and unfinished budgets look like near-term scaffold and training targets that might raise performance without solving creativity.
- Results on two empirical NeurIPS-style questions may understate agent skill on more metric-driven or hill-climbable slices of AI R&D.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes shadow evaluations as a third method for measuring progress toward AI R&D automation: frontier agents are given the central open-ended research question from a high-quality unpublished paper, run for roughly six days with large API/GPU budgets on general-purpose scaffolds, and are graded by the original authors as conference reviewers. On two unpublished NeurIPS-2026-style questions (Personas; TabPFN), agents completed literature review, GPU debugging, large experiment suites, and camera-ready LaTeX without substantive human research help, yet both outputs were unambiguously rejected (overall 2/6 and 1/6). From logs the authors extract five recurring failure modes—judgment about the publishable bar, uncreative response to design shortcomings, ineffective project-level backtracking, poor resource/timeline awareness, and instruction drift—and report a Codex/GPT-5.6 Sol Ultra robustness rerun that largely reproduces them. Artifacts (reviews, survey priors, repos, logs) are released; limitations of n=2, non-blind grading, and open-world discretion are discussed in Table 3 and §§7–8.
Significance. If the reported pattern holds, the work supplies a concrete, complementary evaluation construct between verifier-scored R&D benchmarks and stochastic blind peer review, and it documents a generator–verifier gap plus engineering-vs-research dissociation that matters for forecasts of recursive self-improvement. Strengths that should be credited include: uncontaminated tasks from unpublished submissions; expert grading with released full reviews; pre-experiment collaborator priors; multi-source AI review trajectories that never accepted; annotated resource-use timelines (Figure 1); a cross-model/scaffold robustness check; and unusually full release of logs and repositories. As early case-study evidence the contribution is real; the broader claim that agents “struggle with critical parts of the research lifecycle” is only as strong as the n=2 non-blind bridge the authors themselves flag.
major comments (4)
- [Abstract; §1; Table 1; Table 3; §§7–8] Abstract and §1/§5 generalize from two author-rejected runs to “today’s agents … struggle with critical parts of the research lifecycle.” Table 3 and §§7–8 correctly list small sample, non-blind grading, and question selection as limitations, but the load-bearing inferential step is still thin: graders are the original authors (already committed to a framing), know the papers are AI-generated, and score 1–2/6 with high confidence. For the lifecycle claim to stand at journal strength, either (a) obtain at least one independent expert review per agent paper under the same NeurIPS rubric (blinded to AI authorship if feasible), or (b) systematically rewrite title/abstract/conclusion so the claim is explicitly scoped to “two NeurIPS-style open-ended questions under this protocol,” with the lifecycle language marked as a hypothesis for follow-up. The released author reviews look severe enough
- [§3; Figure 4; §7.1] §3 documents three human interventions: OpenClaw/Anthropic thinking-block harness patch (14 resets TabPFN, 5 Personas), a 24-hour deadline extension after self-graded Weak Reject drafts, and a mandated readability rewrite of “inscrutable” prose (Figure 4). The paper argues these are logistical and do not negate autonomous research failure. That is plausible for engineering competence, but two interventions directly shape the graded object (extra time after the agent declared completion; human-requested clarity pass) and one repeatedly wiped context. A load-bearing clarification is needed: report scores or qualitative deltas on the pre-extension / pre-rewrite drafts versus finals, and state whether author grades apply only to the post-intervention PDFs. Without that, “unambiguous rejection of autonomous open-ended research” is partly confounded with “rejection after scaffold patches and h
- [§5.3; §4.3; Figure 2] §5.3 treats failure to abandon the approach and restart as a primary failure mode, yet §5.3 and the TabPFN narrative note the design required a paper and offered no abstain option, which “may have led the agent to write a negative-results paper.” That is a protocol confound for the backtracking claim: a human might stop or pivot without delivering a forced manuscript. Either add an explicit abstain/“return null with justification” action and re-interpret logs under that counterfactual, or downgrade “ineffective backtracking” from a model failure mode to a joint agent–protocol outcome and adjust §5’s five-mode taxonomy accordingly. As written, the causal attribution to the agent alone overreaches the design.
- [§4.4; Figure 3; Table 4] §4.4 and Figure 3 show AI self- and external reviews never accepting, and Table 4 shows partial overlap with human critiques, supporting a possible generator–verifier gap. The same subsection correctly notes both agent papers were rejects, so one cannot tell whether verifiers discriminate quality or uniformly reject. This undercuts using the review stack as evidence that agents “could not make good use of feedback” in a way that implies the feedback was calibrated gold. Tighten the claim to: agents did not respond to recurring soundness critiques with redesign (supported by logs), and separately mark verifier accuracy as unestablished. Avoid implying the AI-review panel is a validated training signal for RL until accept-class agent or human papers are graded by the same stack.
minor comments (5)
- [Table 1; footnote 3] Table 1 summary scores are clear; consider adding a one-line note that Personas has since been made public (Baines et al., 2026) while TabPFN details remain restricted, so external readers cannot equally audit both agent papers.
- [Figure 1; §3] Figure 1 GPU budget annotations ($392 of $500 vs $69 of $100) are hard to compare across papers; state explicitly how GPU dollar caps were set relative to author estimates in §3.
- [§4.7] §4.7’s definition of reward hacking (deceiving the verifier into a high score) is useful; cross-reference it when collaborators’ “premature negative result ≈ reward hacking” view is mentioned so readers see the operational criterion.
- [§3; Appendix A] Appendix A research questions are appropriately detailed; ensure the main text points to them early when introducing Personas vs TabPFN so readers need not reach the appendix to understand task open-endedness.
- [passim] Minor copyediting: spacing quirks around “W e”/“T o” and similar artifacts appear throughout the compiled text; clean for camera-ready.
Circularity Check
No significant circularity: empirical case-study outcomes and author grades are not forced by definition, fit, or load-bearing self-citation.
full rationale
This paper’s load-bearing chain is observational, not definitional. Agents are given unpublished research questions, produce papers under fixed budgets, and are graded by the original authors against a NeurIPS-style rubric; the rejects, criterion scores, and five failure modes are grounded in released expert reviews, logs, and a second-scaffold robustness run. Those outcomes are not algebraically or statistically forced by the inputs: success was possible in principle, pre-experiment survey priors are reported separately from results, and the paper does not fit a parameter on the same quantity it then calls a prediction. Self-citations (e.g., Kapoor et al. 2026 on open-world evaluations; related CRUX work) supply methodological framing and limitations language, not a uniqueness theorem or ansatz that forbids alternatives or manufactures the capability claim. Overlap of one coauthor with a source paper and non-blind grading are bias/identification threats the paper itself flags (Table 3; §§7–8); they weaken external validity but do not make ‘agents failed to produce publishable research’ true by construction. No self-definitional loop, fitted-input-as-prediction, uniqueness-import, or renaming-of-a-known-result step is quotable in the required sense. Honest finding: circularity score 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- Wall-clock budget (6 days + 24h extension) =
120h + 24h extension
- API spend cap (~$3000 Anthropic credits) =
$3000
- GPU compute allowances (e.g. ~$500 / ~$100 class budgets in figures) =
paper-specific GPU credit caps
- Exploration gate (~48h before paper writing, overridable by critic) =
48-hour gate
- Model/scaffold selection (Opus 4.8 extra-high on OpenClaw; robustness GPT-5.6 Sol Ultra on Codex) =
Opus 4.8 / GPT-5.6 Sol Ultra
axioms (5)
- domain assumption Original-author conference-style grading of agent papers is a valid primary measure of open-ended research quality for the same question.
- domain assumption Unpublished NeurIPS-caliber empirical questions are representative enough of the open-ended AI R&D skills that matter for automation/RSI debates.
- domain assumption General-purpose scaffolds with broad ML-research instructions (not task-specific scaffolds) are an appropriate elicitation standard for claiming agent capability limits.
- standard math Standard conference review dimensions (quality, clarity, significance, originality, overall 1–6) transfer to grading agent research outputs.
- ad hoc to paper Limited human logistical interventions (bug patch, deadline extension, readability rewrite) do not negate the claim of autonomous research failure.
invented entities (1)
-
Shadow evaluation
independent evidence
read the original abstract
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
Reference graph
Works this paper leans on
-
[1]
No critique here, and no pasting the abstract
Summary Restate the problem, approach, and contributions in your own words — a well-written summary is one the authors would nod along to. No critique here, and no pasting the abstract. Prior Fitted Networks such as T abPFN degrade silently when deployed on data in which the query distribution is shifted away from the context distribution. This paper inve...
-
[2]
there are no signals we can use that leverage a model’s internals
Strengths and Weaknesses Think of these as your reasons to accept or reject. Touch on all four dimensions (Quality, Clarity, Significance, Originality). Be specific — cite sections, equations, tables, or figures — since vague points are unfairly hard for authors to answer. If you argue novelty is lacking, name the prior work and where the overlap is. Strengt...
-
[3]
Life After Benchmark Saturation: A Case Study of CORE-Bench
doi: 10.48550/arXiv.2606.26158. URLhttps://arxiv.org/abs/2606.26158. Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav V orotilov, Gaurav Chaurasia, Dieuwke Hupkes, Ricardo Silveira Cabral, T atiana Shavrina, Jakob Foerster, Y oram Bachrach, William Y ang W ang, and Rob...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2606.26158 2025
-
[4]
URLhttps://arxiv.org/abs/2603.08640
doi: 10.48550/arXiv.2603.08640. URLhttps://arxiv.org/abs/2603.08640. Samuel Schmidgall, Yusheng Su, Ze W ang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043. Assoc...
-
[5]
Shoshannah T ekofsky
URL https://proceedings.neurips.cc/paper_files/paper/2025/hash/ 0d904d300a105809a2114d727851e759-Abstract-Conference.html. Shoshannah T ekofsky. AIs Finetune Their Own Leader: A Barking Simpleton. AI Village Blog, Sage Future, July
2025
-
[6]
https://aivillageblog.substack.com/p/ais-finetune-their-own-leader-a-barking , accessed 2026-07-28. W eco T eam. AIDE2: The first evidence of recursive self-improvement.W eco AI Blog, July 2026. URL https: //www.weco.ai/blog/first-evidence-of-recursive-self-improvement . Published July 14, 2026. Jiaxin W en, Liang Qiu, Joe Benton, Jan Hendrik Kirchner, an...
-
[7]
Stating explicitly what would move your score makes the rebuttal and discussion far more productive
Questions Aim for roughly 3–5 focused, actionable items where an author response could genuinely change your opinion, resolve a confusion, or address a limitation. Stating explicitly what would move your score makes the rebuttal and discussion far more productive
-
[8]
Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No
Limitations If limitations and potential negative societal impact are adequately covered, “Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No” on some checklist items is typically not grounds for rejection. • Adequately addressed? Y es. The authors do an excellent job at stating (sometimes ...
-
[9]
Use the two borderline options sparingly
Overall Score Choose one. Use the two borderline options sparingly. [ ] 6 — Strong Accept.T echnically flawless; potential to reshape one or more areas; exceptional evaluation, reproducibility, and resources; no outstanding ethical concerns. [ ] 5 — Accept.T echnically solid; high impact on a subfield, or moderate-to-high impact across several; strong eva...
-
[10]
single reference-anchored benign-quantile threshold on the raw DoC gap
Section 3.2, T able 3, Appendix B: which detector did you use to produce the table? What are the differences between what you used and Pouget et. al.’s suitability filter using confidence as the statistic (which they literally test)? 3.2 says that the alarms use the suitability filter’s non-inferiority test but Appendix B says “single reference-anchored b...
-
[11]
Also distribution shifts can be induced via sampling if a held-out validation set is present
Can the authors widen their experimental grid by using real datasets with (covariate) shift? There are plenty of datasets whose test distributions are shifted from the training dataset, among which are Folktables and ACS which are present here but 2 is not satisfactory. Also distribution shifts can be induced via sampling if a held-out validation set is p...
-
[12]
What motivated your decision to re-implement D3M when the code is readily available and public? Could you at least compare the performance of your approximation with the reference implementation for validity?
-
[15]
Deeply familiar with the related work; checked the math and details carefully
Confidence [×]5— Absolutely certain. Deeply familiar with the related work; checked the math and details carefully. [ ]4— Confident but not certain. Small chance of a misunderstanding or an unfamiliar piece of related work. [ ] 3— Fairly confident. Possible gaps in my understanding or in my coverage of the literature; details not carefully verified. [ ]2—...
-
[16]
No critique here, and no pasting the abstract
Summary Restate the problem, approach, and contributions in your own words — a well-written summary is one the authors would nod along to. No critique here, and no pasting the abstract. The personas of large language models can be controlled using weight diffs between the base model and the fine-tuned model towards a certain trait. Such a diff is a coordi...
-
[17]
structured,
Strengths and Weaknesses Think of these as your reasons to accept or reject. Touch on all four dimensions (Quality, Clarity, Significance, Originality). Be specific — cite sections, equations, tables, or figures — since vague points are unfairly hard for authors to answer. If 27 you argue novelty is lacking, name the prior work and where the overlap is. Stre...
-
[18]
Stating explicitly what would move your score makes the rebuttal and discussion far more productive
Questions Aim for roughly 3–5 focused, actionable items where an author response could genuinely change your opinion, resolve a confusion, or address a limitation. Stating explicitly what would move your score makes the rebuttal and discussion far more productive. • What can be done with this weight-space representation that a capacity-matched activation ...
-
[19]
Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No
Limitations If limitations and potential negative societal impact are adequately covered, “Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No” on some checklist items is typically not grounds for rejection. •Adequately addressed? Y es •If no, what’s missing and how to fix it:
-
[20]
Use the two borderline options sparingly
Overall Score Choose one. Use the two borderline options sparingly. [ ] 6 — Strong Accept.T echnically flawless; potential to reshape one or more areas; exceptional evaluation, reproducibility, and resources; no outstanding ethical concerns. [ ] 5 — Accept.T echnically solid; high impact on a subfield, or moderate-to-high impact across several; strong eva...
-
[21]
what frontier AI agents are capable of as of June 2026
Confidence [ ]5— Absolutely certain. Deeply familiar with the related work; checked the math and details carefully. [×]4— Confident but not certain. Small chance of a misunderstanding or an unfamiliar piece of related work. [ ] 3— Fairly confident. Possible gaps in my understanding or in my coverage of the literature; details not carefully verified. [ ]2—...
2026
-
[2025]
URLhttps://openreview.net/forum?id=6s5uXNWGIh. Hui Chen, Miao Xiong, Yujie Lu, W ei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi. MLR-Bench: Evaluating AI agents on open-ended machine learning research. InAdvances in Neural Information Processing Systems, volume 38, 2025. Corinna Cortes and Neil D. Lawrence. Inconsistency in con...
-
[2026]
URLhttps://arxiv.org/abs/2606.26294
doi: 10.48550/arXiv.2606.26294. URLhttps://arxiv.org/abs/2606.26294. Intology. Zochi publishes A * paper.Intology Blog, May 2025. URL https://www.intology.ai/blog/ zochi-acl. Published May 27, 2025. Peter Jansen, Oyvind T afjord, Marissa Radensky, Pao Siangliulue, T om Hope, Bhavana Dalvi Mishra, Bod- hisattwa Prasad Majumder, Daniel S. W eld, and Peter C...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.