Pith. sign in

REVIEW 4 major objections 5 minor 21 references

Today's frontier AI agents can finish the engineering of open-ended AI research but still cannot answer the research questions at a publishable standard.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 10:53 UTC pith:PMIYGK5J

load-bearing objection Useful new eval design plus honest negative case studies; the n=2 non-blind bridge is thin but the paper already says so, and the released rejects/logs still earn a serious read. the 4 major comments →

arxiv 2607.27191 v1 pith:PMIYGK5J submitted 2026-07-29 cs.AI cs.CYcs.LG

Can AI agents conduct open-ended AI research? Early evidence from two case studies

classification cs.AI cs.CYcs.LG
keywords AI agentsAI R&D automationshadow evaluationopen-ended researchresearch engineeringfailure modespeer reviewlong-horizon agents
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Forecasts of explosive AI progress often assume agents will soon automate AI research itself, yet most tests either score narrow verifiable metrics or rely on noisy conference peer review. This paper offers a third measurement: shadow evaluations, in which an agent receives the central open-ended question from a high-quality unpublished paper, multi-day time and large compute budgets, and is graded by the original authors as if for a top conference. On two unpublished NeurIPS-level questions, frontier agents completed literature review, GPU work, experiments, and paper writing without human help, but made little substantive progress on the science; both papers were unambiguously rejected. The authors document five recurring failure modes—poor judgment of the publishable bar, uncreative responses to design flaws, weak backtracking, poor resource awareness, and instruction drift—and reproduce them with a second model and scaffold. The result is early evidence that research engineering is ahead of the judgment-heavy parts of the research lifecycle that self-improving AI forecasts depend on.

Core claim

Well-resourced frontier agents can autonomously perform the engineering steps of open-ended AI research, yet cannot make substantial progress toward answering those research questions at top-conference quality. In two shadow evaluations on unpublished NeurIPS submissions, original authors rejected both agent-written papers, citing weak experimental judgment, thin novelty, and unclear writing despite successful end-to-end execution.

What carries the argument

Shadow evaluations: assign a frontier agent the central open-ended research question of a high-quality unpublished paper, give it multi-day wall-clock time and large API and GPU budgets on a general-purpose scaffold, then have the paper's original authors grade the agent's output as a conference submission.

Load-bearing premise

That unambiguous rejection by the original authors on two non-blind case studies is a reliable enough signal of agent research inability, rather than mainly author preference, scaffold under-elicitation, or the particular questions chosen.

What would settle it

A shadow evaluation on a comparable unpublished top-conference question where a frontier agent under similar time and compute budgets produces a paper the original authors score as at least a weak accept, or a larger series of such runs that consistently clear that bar.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Claims that AI will soon automate AI R&D must separate engineering automation from open-ended scientific judgment and redesign.
  • Verifier-only and blind-review evaluations will systematically miss the failure modes that shadow evaluations expose.
  • More wall-clock time or raw compute alone is unlikely to fix premature commitment, weak backtracking, and unused budgets.
  • Self-review loops that reliably reject weak drafts do not by themselves force creative redesign of the research approach.
  • Released logs, reviews, and repositories make the method reusable as stronger models and scaffolds appear.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If engineering keeps improving while judgment, backtracking, and resource sense lag, recursive self-improvement may accelerate narrow optimization faster than open-ended discovery.
  • Because both agent papers were rejects, the study cannot yet tell whether AI reviewers are truly discriminative or merely harsh—closing that generator-verifier gap would need accepted as well as rejected drafts.
  • Instruction drift and unfinished budgets look like near-term scaffold and training targets that might raise performance without solving creativity.
  • Results on two empirical NeurIPS-style questions may understate agent skill on more metric-driven or hill-climbable slices of AI R&D.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes shadow evaluations as a third method for measuring progress toward AI R&D automation: frontier agents are given the central open-ended research question from a high-quality unpublished paper, run for roughly six days with large API/GPU budgets on general-purpose scaffolds, and are graded by the original authors as conference reviewers. On two unpublished NeurIPS-2026-style questions (Personas; TabPFN), agents completed literature review, GPU debugging, large experiment suites, and camera-ready LaTeX without substantive human research help, yet both outputs were unambiguously rejected (overall 2/6 and 1/6). From logs the authors extract five recurring failure modes—judgment about the publishable bar, uncreative response to design shortcomings, ineffective project-level backtracking, poor resource/timeline awareness, and instruction drift—and report a Codex/GPT-5.6 Sol Ultra robustness rerun that largely reproduces them. Artifacts (reviews, survey priors, repos, logs) are released; limitations of n=2, non-blind grading, and open-world discretion are discussed in Table 3 and §§7–8.

Significance. If the reported pattern holds, the work supplies a concrete, complementary evaluation construct between verifier-scored R&D benchmarks and stochastic blind peer review, and it documents a generator–verifier gap plus engineering-vs-research dissociation that matters for forecasts of recursive self-improvement. Strengths that should be credited include: uncontaminated tasks from unpublished submissions; expert grading with released full reviews; pre-experiment collaborator priors; multi-source AI review trajectories that never accepted; annotated resource-use timelines (Figure 1); a cross-model/scaffold robustness check; and unusually full release of logs and repositories. As early case-study evidence the contribution is real; the broader claim that agents “struggle with critical parts of the research lifecycle” is only as strong as the n=2 non-blind bridge the authors themselves flag.

major comments (4)
  1. [Abstract; §1; Table 1; Table 3; §§7–8] Abstract and §1/§5 generalize from two author-rejected runs to “today’s agents … struggle with critical parts of the research lifecycle.” Table 3 and §§7–8 correctly list small sample, non-blind grading, and question selection as limitations, but the load-bearing inferential step is still thin: graders are the original authors (already committed to a framing), know the papers are AI-generated, and score 1–2/6 with high confidence. For the lifecycle claim to stand at journal strength, either (a) obtain at least one independent expert review per agent paper under the same NeurIPS rubric (blinded to AI authorship if feasible), or (b) systematically rewrite title/abstract/conclusion so the claim is explicitly scoped to “two NeurIPS-style open-ended questions under this protocol,” with the lifecycle language marked as a hypothesis for follow-up. The released author reviews look severe enough
  2. [§3; Figure 4; §7.1] §3 documents three human interventions: OpenClaw/Anthropic thinking-block harness patch (14 resets TabPFN, 5 Personas), a 24-hour deadline extension after self-graded Weak Reject drafts, and a mandated readability rewrite of “inscrutable” prose (Figure 4). The paper argues these are logistical and do not negate autonomous research failure. That is plausible for engineering competence, but two interventions directly shape the graded object (extra time after the agent declared completion; human-requested clarity pass) and one repeatedly wiped context. A load-bearing clarification is needed: report scores or qualitative deltas on the pre-extension / pre-rewrite drafts versus finals, and state whether author grades apply only to the post-intervention PDFs. Without that, “unambiguous rejection of autonomous open-ended research” is partly confounded with “rejection after scaffold patches and h
  3. [§5.3; §4.3; Figure 2] §5.3 treats failure to abandon the approach and restart as a primary failure mode, yet §5.3 and the TabPFN narrative note the design required a paper and offered no abstain option, which “may have led the agent to write a negative-results paper.” That is a protocol confound for the backtracking claim: a human might stop or pivot without delivering a forced manuscript. Either add an explicit abstain/“return null with justification” action and re-interpret logs under that counterfactual, or downgrade “ineffective backtracking” from a model failure mode to a joint agent–protocol outcome and adjust §5’s five-mode taxonomy accordingly. As written, the causal attribution to the agent alone overreaches the design.
  4. [§4.4; Figure 3; Table 4] §4.4 and Figure 3 show AI self- and external reviews never accepting, and Table 4 shows partial overlap with human critiques, supporting a possible generator–verifier gap. The same subsection correctly notes both agent papers were rejects, so one cannot tell whether verifiers discriminate quality or uniformly reject. This undercuts using the review stack as evidence that agents “could not make good use of feedback” in a way that implies the feedback was calibrated gold. Tighten the claim to: agents did not respond to recurring soundness critiques with redesign (supported by logs), and separately mark verifier accuracy as unestablished. Avoid implying the AI-review panel is a validated training signal for RL until accept-class agent or human papers are graded by the same stack.
minor comments (5)
  1. [Table 1; footnote 3] Table 1 summary scores are clear; consider adding a one-line note that Personas has since been made public (Baines et al., 2026) while TabPFN details remain restricted, so external readers cannot equally audit both agent papers.
  2. [Figure 1; §3] Figure 1 GPU budget annotations ($392 of $500 vs $69 of $100) are hard to compare across papers; state explicitly how GPU dollar caps were set relative to author estimates in §3.
  3. [§4.7] §4.7’s definition of reward hacking (deceiving the verifier into a high score) is useful; cross-reference it when collaborators’ “premature negative result ≈ reward hacking” view is mentioned so readers see the operational criterion.
  4. [§3; Appendix A] Appendix A research questions are appropriately detailed; ensure the main text points to them early when introducing Personas vs TabPFN so readers need not reach the appendix to understand task open-endedness.
  5. [passim] Minor copyediting: spacing quirks around “W e”/“T o” and similar artifacts appear throughout the compiled text; clean for camera-ready.

Circularity Check

0 steps flagged

No significant circularity: empirical case-study outcomes and author grades are not forced by definition, fit, or load-bearing self-citation.

full rationale

This paper’s load-bearing chain is observational, not definitional. Agents are given unpublished research questions, produce papers under fixed budgets, and are graded by the original authors against a NeurIPS-style rubric; the rejects, criterion scores, and five failure modes are grounded in released expert reviews, logs, and a second-scaffold robustness run. Those outcomes are not algebraically or statistically forced by the inputs: success was possible in principle, pre-experiment survey priors are reported separately from results, and the paper does not fit a parameter on the same quantity it then calls a prediction. Self-citations (e.g., Kapoor et al. 2026 on open-world evaluations; related CRUX work) supply methodological framing and limitations language, not a uniqueness theorem or ansatz that forbids alternatives or manufactures the capability claim. Overlap of one coauthor with a source paper and non-blind grading are bias/identification threats the paper itself flags (Table 3; §§7–8); they weaken external validity but do not make ‘agents failed to produce publishable research’ true by construction. No self-definitional loop, fitted-input-as-prediction, uniqueness-import, or renaming-of-a-known-result step is quotable in the required sense. Honest finding: circularity score 0.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

This is an empirical methods-and-measurement paper, not a first-principles derivation. The claim rests on design choices about what counts as open-ended AI research, how to resource agents, and how to interpret non-blind expert rejection—not on fitted physical constants or invented particles. The ledger below captures those load-bearing methodological commitments.

free parameters (5)
  • Wall-clock budget (6 days + 24h extension) = 120h + 24h extension
    Chosen with author input as 'enough to answer the question substantively'; not derived. Results partly depend on agents failing to use the full budget, but the horizon itself is a free design choice.
  • API spend cap (~$3000 Anthropic credits) = $3000
    Hand-set resource limit that defines the elicitation regime; agents left >50% unspent in main runs.
  • GPU compute allowances (e.g. ~$500 / ~$100 class budgets in figures) = paper-specific GPU credit caps
    Author-estimated experiment budgets; free parameters of the evaluation design.
  • Exploration gate (~48h before paper writing, overridable by critic) = 48-hour gate
    Scaffold heuristic introduced after pilots to force exploration; agents still coalesced early.
  • Model/scaffold selection (Opus 4.8 extra-high on OpenClaw; robustness GPT-5.6 Sol Ultra on Codex) = Opus 4.8 / GPT-5.6 Sol Ultra
    Chosen after dry runs; performance claims are conditional on these elicitation choices.
axioms (5)
  • domain assumption Original-author conference-style grading of agent papers is a valid primary measure of open-ended research quality for the same question.
    Core of the shadow-evaluation design (Section 2); authors acknowledge non-blind bias risk in Section 7–8 and Table 3.
  • domain assumption Unpublished NeurIPS-caliber empirical questions are representative enough of the open-ended AI R&D skills that matter for automation/RSI debates.
    Stated motivation in Introduction and Section 7.2; paper notes hill-climbing on verifiable objectives might dominate real lab progress.
  • domain assumption General-purpose scaffolds with broad ML-research instructions (not task-specific scaffolds) are an appropriate elicitation standard for claiming agent capability limits.
    Section 3 and Table 3; robustness run partially addresses scaffold overhang concerns.
  • standard math Standard conference review dimensions (quality, clarity, significance, originality, overall 1–6) transfer to grading agent research outputs.
    Uses NeurIPS-style rubrics and AI review tools as operationalizations; conventional in the field though known to be noisy.
  • ad hoc to paper Limited human logistical interventions (bug patch, deadline extension, readability rewrite) do not negate the claim of autonomous research failure.
    Section 3 documents three intervention classes; interpretation treats them as non-substantive to the scientific contribution.
invented entities (1)
  • Shadow evaluation independent evidence
    purpose: Name and operationalize a third evaluation paradigm: agent answers central question of an unpublished paper; original authors grade.
    Methodological construct introduced in Section 2; not a physical entity. Independent handle exists: others can repeat the protocol on new unpublished papers.

pith-pipeline@v1.2.0-daily-grok45 · 32723 in / 3693 out tokens · 73255 ms · 2026-07-30T10:53:55.290705+00:00 · methodology

0 comments
read the original abstract

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

21 extracted references · 1 canonical work pages · 1 internal anchor

  1. [1]

    No critique here, and no pasting the abstract

    Summary Restate the problem, approach, and contributions in your own words — a well-written summary is one the authors would nod along to. No critique here, and no pasting the abstract. Prior Fitted Networks such as T abPFN degrade silently when deployed on data in which the query distribution is shifted away from the context distribution. This paper inve...

  2. [2]

    there are no signals we can use that leverage a model’s internals

    Strengths and Weaknesses Think of these as your reasons to accept or reject. Touch on all four dimensions (Quality, Clarity, Significance, Originality). Be specific — cite sections, equations, tables, or figures — since vague points are unfairly hard for authors to answer. If you argue novelty is lacking, name the prior work and where the overlap is. Strengt...

  3. [3]

    Life After Benchmark Saturation: A Case Study of CORE-Bench

    doi: 10.48550/arXiv.2606.26158. URLhttps://arxiv.org/abs/2606.26158. Deepak Nathani, Lovish Madaan, Nicholas Roberts, Nikolay Bashlykov, Ajay Menon, Vincent Moens, Amar Budhiraja, Despoina Magka, Vladislav V orotilov, Gaurav Chaurasia, Dieuwke Hupkes, Ricardo Silveira Cabral, T atiana Shavrina, Jakob Foerster, Y oram Bachrach, William Y ang W ang, and Rob...

  4. [4]

    URLhttps://arxiv.org/abs/2603.08640

    doi: 10.48550/arXiv.2603.08640. URLhttps://arxiv.org/abs/2603.08640. Samuel Schmidgall, Yusheng Su, Ze W ang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043. Assoc...

  5. [5]

    Shoshannah T ekofsky

    URL https://proceedings.neurips.cc/paper_files/paper/2025/hash/ 0d904d300a105809a2114d727851e759-Abstract-Conference.html. Shoshannah T ekofsky. AIs Finetune Their Own Leader: A Barking Simpleton. AI Village Blog, Sage Future, July

  6. [6]

    trait space

    https://aivillageblog.substack.com/p/ais-finetune-their-own-leader-a-barking , accessed 2026-07-28. W eco T eam. AIDE2: The first evidence of recursive self-improvement.W eco AI Blog, July 2026. URL https: //www.weco.ai/blog/first-evidence-of-recursive-self-improvement . Published July 14, 2026. Jiaxin W en, Liang Qiu, Joe Benton, Jan Hendrik Kirchner, an...

  7. [7]

    Stating explicitly what would move your score makes the rebuttal and discussion far more productive

    Questions Aim for roughly 3–5 focused, actionable items where an author response could genuinely change your opinion, resolve a confusion, or address a limitation. Stating explicitly what would move your score makes the rebuttal and discussion far more productive

  8. [8]

    Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No

    Limitations If limitations and potential negative societal impact are adequately covered, “Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No” on some checklist items is typically not grounds for rejection. • Adequately addressed? Y es. The authors do an excellent job at stating (sometimes ...

  9. [9]

    Use the two borderline options sparingly

    Overall Score Choose one. Use the two borderline options sparingly. [ ] 6 — Strong Accept.T echnically flawless; potential to reshape one or more areas; exceptional evaluation, reproducibility, and resources; no outstanding ethical concerns. [ ] 5 — Accept.T echnically solid; high impact on a subfield, or moderate-to-high impact across several; strong eva...

  10. [10]

    single reference-anchored benign-quantile threshold on the raw DoC gap

    Section 3.2, T able 3, Appendix B: which detector did you use to produce the table? What are the differences between what you used and Pouget et. al.’s suitability filter using confidence as the statistic (which they literally test)? 3.2 says that the alarms use the suitability filter’s non-inferiority test but Appendix B says “single reference-anchored b...

  11. [11]

    Also distribution shifts can be induced via sampling if a held-out validation set is present

    Can the authors widen their experimental grid by using real datasets with (covariate) shift? There are plenty of datasets whose test distributions are shifted from the training dataset, among which are Folktables and ACS which are present here but 2 is not satisfactory. Also distribution shifts can be induced via sampling if a held-out validation set is p...

  12. [12]

    What motivated your decision to re-implement D3M when the code is readily available and public? Could you at least compare the performance of your approximation with the reference implementation for validity?

  13. [15]

    Deeply familiar with the related work; checked the math and details carefully

    Confidence [×]5— Absolutely certain. Deeply familiar with the related work; checked the math and details carefully. [ ]4— Confident but not certain. Small chance of a misunderstanding or an unfamiliar piece of related work. [ ] 3— Fairly confident. Possible gaps in my understanding or in my coverage of the literature; details not carefully verified. [ ]2—...

  14. [16]

    No critique here, and no pasting the abstract

    Summary Restate the problem, approach, and contributions in your own words — a well-written summary is one the authors would nod along to. No critique here, and no pasting the abstract. The personas of large language models can be controlled using weight diffs between the base model and the fine-tuned model towards a certain trait. Such a diff is a coordi...

  15. [17]

    structured,

    Strengths and Weaknesses Think of these as your reasons to accept or reject. Touch on all four dimensions (Quality, Clarity, Significance, Originality). Be specific — cite sections, equations, tables, or figures — since vague points are unfairly hard for authors to answer. If 27 you argue novelty is lacking, name the prior work and where the overlap is. Stre...

  16. [18]

    Stating explicitly what would move your score makes the rebuttal and discussion far more productive

    Questions Aim for roughly 3–5 focused, actionable items where an author response could genuinely change your opinion, resolve a confusion, or address a limitation. Stating explicitly what would move your score makes the rebuttal and discussion far more productive. • What can be done with this weight-space representation that a capacity-matched activation ...

  17. [19]

    Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No

    Limitations If limitations and potential negative societal impact are adequately covered, “Yes” suffices. If not, give constructive suggestions. Authors should be rewarded, not punished, for candor — and a “No” on some checklist items is typically not grounds for rejection. •Adequately addressed? Y es •If no, what’s missing and how to fix it:

  18. [20]

    Use the two borderline options sparingly

    Overall Score Choose one. Use the two borderline options sparingly. [ ] 6 — Strong Accept.T echnically flawless; potential to reshape one or more areas; exceptional evaluation, reproducibility, and resources; no outstanding ethical concerns. [ ] 5 — Accept.T echnically solid; high impact on a subfield, or moderate-to-high impact across several; strong eva...

  19. [21]

    what frontier AI agents are capable of as of June 2026

    Confidence [ ]5— Absolutely certain. Deeply familiar with the related work; checked the math and details carefully. [×]4— Confident but not certain. Small chance of a misunderstanding or an unfamiliar piece of related work. [ ] 3— Fairly confident. Possible gaps in my understanding or in my coverage of the literature; details not carefully verified. [ ]2—...

  20. [2025]

    Hui Chen, Miao Xiong, Yujie Lu, W ei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi

    URLhttps://openreview.net/forum?id=6s5uXNWGIh. Hui Chen, Miao Xiong, Yujie Lu, W ei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, and Bryan Hooi. MLR-Bench: Evaluating AI agents on open-ended machine learning research. InAdvances in Neural Information Processing Systems, volume 38, 2025. Corinna Cortes and Neil D. Lawrence. Inconsistency in con...

  21. [2026]

    URLhttps://arxiv.org/abs/2606.26294

    doi: 10.48550/arXiv.2606.26294. URLhttps://arxiv.org/abs/2606.26294. Intology. Zochi publishes A * paper.Intology Blog, May 2025. URL https://www.intology.ai/blog/ zochi-acl. Published May 27, 2025. Peter Jansen, Oyvind T afjord, Marissa Radensky, Pao Siangliulue, T om Hope, Bhavana Dalvi Mishra, Bod- hisattwa Prasad Majumder, Daniel S. W eld, and Peter C...