REVIEW 4 major objections 5 minor
Can AI agents conduct open-ended AI research? Early evidence from two case studies
T0 review · 4 major / 5 minor · reviewed 2026-07-30 · grok-4.5
Pith's one-line read Today's frontier AI agents can finish the engineering of open-ended AI research but still cannot answer the research questions at a publishable standard.
desk verdict Useful new eval design plus honest negative case studies; the n=2 non-blind bridge is thin but the paper already says so, and the released rejects/logs still earn a serious read. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Shadow evaluations: assign a frontier agent the central open-ended research question of a high-quality unpublished paper, give it multi-day wall-clock time and large API and GPU budgets on a general-purpose scaffold, then have the paper's original authors grade the agent's output as a conference submission.
What would settle it
A shadow evaluation on a comparable unpublished top-conference question where a frontier agent under similar time and compute budgets produces a paper the original authors score as at least a weak accept, or a larger series of such runs that consistently clear that bar.
Extended reading notes
Core claim
Well-resourced frontier agents can autonomously perform the engineering steps of open-ended AI research, yet cannot make substantial progress toward answering those research questions at top-conference quality. In two shadow evaluations on unpublished NeurIPS submissions, original authors rejected both agent-written papers, citing weak experimental judgment, thin novelty, and unclear writing despite successful end-to-end execution.
Load-bearing premise
That unambiguous rejection by the original authors on two non-blind case studies is a reliable enough signal of agent research inability, rather than mainly author preference, scaffold under-elicitation, or the particular questions chosen.
Editorial extensions
If this is right
- Claims that AI will soon automate AI R&D must separate engineering automation from open-ended scientific judgment and redesign.
- Verifier-only and blind-review evaluations will systematically miss the failure modes that shadow evaluations expose.
- More wall-clock time or raw compute alone is unlikely to fix premature commitment, weak backtracking, and unused budgets.
- Self-review loops that reliably reject weak drafts do not by themselves force creative redesign of the research approach.
- Released logs, reviews, and repositories make the method reusable as stronger models and scaffolds appear.
Reading between the lines
- If engineering keeps improving while judgment, backtracking, and resource sense lag, recursive self-improvement may accelerate narrow optimization faster than open-ended discovery.
- Because both agent papers were rejects, the study cannot yet tell whether AI reviewers are truly discriminative or merely harsh—closing that generator-verifier gap would need accepted as well as rejected drafts.
- Instruction drift and unfinished budgets look like near-term scaffold and training targets that might raise performance without solving creativity.
- Results on two empirical NeurIPS-style questions may understate agent skill on more metric-driven or hill-climbable slices of AI R&D.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes shadow evaluations as a third method for measuring progress toward AI R&D automation: frontier agents are given the central open-ended research question from a high-quality unpublished paper, run for roughly six days with large API/GPU budgets on general-purpose scaffolds, and are graded by the original authors as conference reviewers. On two unpublished NeurIPS-2026-style questions (Personas; TabPFN), agents completed literature review, GPU debugging, large experiment suites, and camera-ready LaTeX without substantive human research help, yet both outputs were unambiguously rejected (overall 2/6 and 1/6). From logs the authors extract five recurring failure modes—judgment about the publishable bar, uncreative response to design shortcomings, ineffective project-level backtracking, poor resource/timeline awareness, and instruction drift—and report a Codex/GPT-5.6 Sol Ultra robustness rerun that largely reproduces them. Artifacts (reviews, survey priors, repos, logs) are released; limitations of n=2, non-blind grading, and open-world discretion are discussed in Table 3 and §§7–8.
Significance. If the reported pattern holds, the work supplies a concrete, complementary evaluation construct between verifier-scored R&D benchmarks and stochastic blind peer review, and it documents a generator–verifier gap plus engineering-vs-research dissociation that matters for forecasts of recursive self-improvement. Strengths that should be credited include: uncontaminated tasks from unpublished submissions; expert grading with released full reviews; pre-experiment collaborator priors; multi-source AI review trajectories that never accepted; annotated resource-use timelines (Figure 1); a cross-model/scaffold robustness check; and unusually full release of logs and repositories. As early case-study evidence the contribution is real; the broader claim that agents “struggle with critical parts of the research lifecycle” is only as strong as the n=2 non-blind bridge the authors themselves flag.
major comments (4)
- [Abstract; §1; Table 1; Table 3; §§7–8] Abstract and §1/§5 generalize from two author-rejected runs to “today’s agents … struggle with critical parts of the research lifecycle.” Table 3 and §§7–8 correctly list small sample, non-blind grading, and question selection as limitations, but the load-bearing inferential step is still thin: graders are the original authors (already committed to a framing), know the papers are AI-generated, and score 1–2/6 with high confidence. For the lifecycle claim to stand at journal strength, either (a) obtain at least one independent expert review per agent paper under the same NeurIPS rubric (blinded to AI authorship if feasible), or (b) systematically rewrite title/abstract/conclusion so the claim is explicitly scoped to “two NeurIPS-style open-ended questions under this protocol,” with the lifecycle language marked as a hypothesis for follow-up. The released author reviews look severe enough
- [§3; Figure 4; §7.1] §3 documents three human interventions: OpenClaw/Anthropic thinking-block harness patch (14 resets TabPFN, 5 Personas), a 24-hour deadline extension after self-graded Weak Reject drafts, and a mandated readability rewrite of “inscrutable” prose (Figure 4). The paper argues these are logistical and do not negate autonomous research failure. That is plausible for engineering competence, but two interventions directly shape the graded object (extra time after the agent declared completion; human-requested clarity pass) and one repeatedly wiped context. A load-bearing clarification is needed: report scores or qualitative deltas on the pre-extension / pre-rewrite drafts versus finals, and state whether author grades apply only to the post-intervention PDFs. Without that, “unambiguous rejection of autonomous open-ended research” is partly confounded with “rejection after scaffold patches and h
- [§5.3; §4.3; Figure 2] §5.3 treats failure to abandon the approach and restart as a primary failure mode, yet §5.3 and the TabPFN narrative note the design required a paper and offered no abstain option, which “may have led the agent to write a negative-results paper.” That is a protocol confound for the backtracking claim: a human might stop or pivot without delivering a forced manuscript. Either add an explicit abstain/“return null with justification” action and re-interpret logs under that counterfactual, or downgrade “ineffective backtracking” from a model failure mode to a joint agent–protocol outcome and adjust §5’s five-mode taxonomy accordingly. As written, the causal attribution to the agent alone overreaches the design.
- [§4.4; Figure 3; Table 4] §4.4 and Figure 3 show AI self- and external reviews never accepting, and Table 4 shows partial overlap with human critiques, supporting a possible generator–verifier gap. The same subsection correctly notes both agent papers were rejects, so one cannot tell whether verifiers discriminate quality or uniformly reject. This undercuts using the review stack as evidence that agents “could not make good use of feedback” in a way that implies the feedback was calibrated gold. Tighten the claim to: agents did not respond to recurring soundness critiques with redesign (supported by logs), and separately mark verifier accuracy as unestablished. Avoid implying the AI-review panel is a validated training signal for RL until accept-class agent or human papers are graded by the same stack.
minor comments (5)
- [Table 1; footnote 3] Table 1 summary scores are clear; consider adding a one-line note that Personas has since been made public (Baines et al., 2026) while TabPFN details remain restricted, so external readers cannot equally audit both agent papers.
- [Figure 1; §3] Figure 1 GPU budget annotations ($392 of $500 vs $69 of $100) are hard to compare across papers; state explicitly how GPU dollar caps were set relative to author estimates in §3.
- [§4.7] §4.7’s definition of reward hacking (deceiving the verifier into a high score) is useful; cross-reference it when collaborators’ “premature negative result ≈ reward hacking” view is mentioned so readers see the operational criterion.
- [§3; Appendix A] Appendix A research questions are appropriately detailed; ensure the main text points to them early when introducing Personas vs TabPFN so readers need not reach the appendix to understand task open-endedness.
- [passim] Minor copyediting: spacing quirks around “W e”/“T o” and similar artifacts appear throughout the compiled text; clean for camera-ready.
Circularity Check
No significant circularity: empirical case-study outcomes and author grades are not forced by definition, fit, or load-bearing self-citation.
full rationale
This paper’s load-bearing chain is observational, not definitional. Agents are given unpublished research questions, produce papers under fixed budgets, and are graded by the original authors against a NeurIPS-style rubric; the rejects, criterion scores, and five failure modes are grounded in released expert reviews, logs, and a second-scaffold robustness run. Those outcomes are not algebraically or statistically forced by the inputs: success was possible in principle, pre-experiment survey priors are reported separately from results, and the paper does not fit a parameter on the same quantity it then calls a prediction. Self-citations (e.g., Kapoor et al. 2026 on open-world evaluations; related CRUX work) supply methodological framing and limitations language, not a uniqueness theorem or ansatz that forbids alternatives or manufactures the capability claim. Overlap of one coauthor with a source paper and non-blind grading are bias/identification threats the paper itself flags (Table 3; §§7–8); they weaken external validity but do not make ‘agents failed to produce publishable research’ true by construction. No self-definitional loop, fitted-input-as-prediction, uniqueness-import, or renaming-of-a-known-result step is quotable in the required sense. Honest finding: circularity score 0.
Assumptions & free parameters
free parameters (5)
- Wall-clock budget (6 days + 24h extension) =
120h + 24h extension
- API spend cap (~$3000 Anthropic credits) =
$3000
- GPU compute allowances (e.g. ~$500 / ~$100 class budgets in figures) =
paper-specific GPU credit caps
- Exploration gate (~48h before paper writing, overridable by critic) =
48-hour gate
- Model/scaffold selection (Opus 4.8 extra-high on OpenClaw; robustness GPT-5.6 Sol Ultra on Codex) =
Opus 4.8 / GPT-5.6 Sol Ultra
assumptions (5)
- domain assumption Original-author conference-style grading of agent papers is a valid primary measure of open-ended research quality for the same question.
- domain assumption Unpublished NeurIPS-caliber empirical questions are representative enough of the open-ended AI R&D skills that matter for automation/RSI debates.
- domain assumption General-purpose scaffolds with broad ML-research instructions (not task-specific scaffolds) are an appropriate elicitation standard for claiming agent capability limits.
- standard math Standard conference review dimensions (quality, clarity, significance, originality, overall 1–6) transfer to grading agent research outputs.
- ad hoc to paper Limited human logistical interventions (bug patch, deadline extension, readability rewrite) do not negate the claim of autonomous research failure.
invented entities (1)
-
Shadow evaluation
independent evidence
Cite this review
Pith. "Pith review of Can AI agents conduct open-ended AI research? Early evidence from two case studies." pith.science (2026). https://pith.science/paper/PMIYGK5J
@misc{pith2026260727191,
author = {Pith},
title = {Pith review of: Can AI agents conduct open-ended AI research? Early evidence from two case studies},
year = {2026},
howpublished = {\url{https://pith.science/paper/PMIYGK5J}},
note = {Machine review of arXiv:2607.27191}
}
read the original abstract
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.
Reviewed July 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.