{"id":"da6e2da6-9505-4d86-8911-535b30b36806","arxiv_id":"2607.27191","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Shadow evaluations on two unpublished NeurIPS questions show today's agents can engineer AI experiments but cannot produce publishable open-ended research.","lead":"Frontier AI agents given six days and thousands of dollars failed to answer open-ended research questions from two unpublished NeurIPS papers, though they handled all the engineering. The original authors rejected both agent papers, exposing five recurring research-lifecycle failures.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"n=2 non-blind author rejects remain the load-bearing bridge from case studies to ‘agents struggle with critical parts of the research lifecycle.’","rationale":"The reader correctly locates the soft spot: not internal contradiction or hidden positive results, but whether n=2 non-blind shadow rejects plus recurring logs justify the broader ‘struggle with critical parts of the research lifecycle’ claim. Competence and honesty of the writeup are high; limitations are foregrounded; the Codex/GPT-5.6 Sol redo reduces pure OpenClaw overhang. Author reviews (Appendix B) are specific and harsh (proof-by-example, hand-picked traits, underpowered design), and agents under-used budget while AI self-reviews never accepted—so the negative case studies are substantive, not empty. That still leaves representativeness and grader bias as the load-bearing assumption for generalization. No stronger internal flaw displaces that. Verdict stays CONDITIONAL; independent blinded secondary scores (or larger n) are the natural clear/fail test. Agreement with the reader is full on the weakest assumption.","tokens_in":28625,"tokens_out":507,"duration_ms":28846,"concrete_test":"For the public Personas agent paper (and TabPFN if releasable), commission 3–5 domain experts who did not author either work to score the agent manuscript on the same NeurIPS rubric, blinded to AI authorship if feasible; pre-register that if median overall ≥4 (borderline accept) or quality/significance rise ≥1 point vs author scores, the ‘unambiguous inability’ reading weakens and CONDITIONAL should stay or soften.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is early evidence that frontier agents can do research engineering but not open-ended publishable AI research, resting on unambiguous author rejects plus five log failure modes and one cross-scaffold check. That bridge is still thin: graders are the original authors (non-blind, already committed to their own framing), the sample is two NeurIPS-style questions and a handful of runs (Table 3; §§7–8), and human interventions (deadline extension, clarity rewrite, harness patch) plus forced paper output (no abstain) shape what ‘failure’ looks like. The paper documents these limits and the rejects look severe on the released reviews, so the case studies are real; they do not yet securely underwrite the lifecycle generalization without more independent grading or broader sampling.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes shadow evaluations as a third method for measuring progress toward AI R&D automation: frontier agents are given the central open-ended research question from a high-quality unpublished paper, run for roughly six days with large API/GPU budgets on general-purpose scaffolds, and are graded by the original authors as conference reviewers. On two unpublished NeurIPS-2026-style questions (Personas; TabPFN), agents completed literature review, GPU debugging, large experiment suites, and camera-ready LaTeX without substantive human research help, yet both outputs were unambiguously rejected (overall 2/6 and 1/6). From logs the authors extract five recurring failure modes—judgment about the publishable bar, uncreative response to design shortcomings, ineffective project-level backtracking, poor resource/timeline awareness, and instruction drift—and report a Codex/GPT-5.6 Sol Ultra robustness rerun that largely reproduces them. Artifacts (reviews, survey priors, repos, logs) are released; limitations of n=2, non-blind grading, and open-world discretion are discussed in Table 3 and §§7–8.","tokens_in":28831,"tokens_out":1681,"duration_ms":43753,"significance":"If the reported pattern holds, the work supplies a concrete, complementary evaluation construct between verifier-scored R&D benchmarks and stochastic blind peer review, and it documents a generator–verifier gap plus engineering-vs-research dissociation that matters for forecasts of recursive self-improvement. Strengths that should be credited include: uncontaminated tasks from unpublished submissions; expert grading with released full reviews; pre-experiment collaborator priors; multi-source AI review trajectories that never accepted; annotated resource-use timelines (Figure 1); a cross-model/scaffold robustness check; and unusually full release of logs and repositories. As early case-study evidence the contribution is real; the broader claim that agents “struggle with critical parts of the research lifecycle” is only as strong as the n=2 non-blind bridge the authors themselves flag.","major_comments":[{"comment":"Abstract and §1/§5 generalize from two author-rejected runs to “today’s agents … struggle with critical parts of the research lifecycle.” Table 3 and §§7–8 correctly list small sample, non-blind grading, and question selection as limitations, but the load-bearing inferential step is still thin: graders are the original authors (already committed to a framing), know the papers are AI-generated, and score 1–2/6 with high confidence. For the lifecycle claim to stand at journal strength, either (a) obtain at least one independent expert review per agent paper under the same NeurIPS rubric (blinded to AI authorship if feasible), or (b) systematically rewrite title/abstract/conclusion so the claim is explicitly scoped to “two NeurIPS-style open-ended questions under this protocol,” with the lifecycle language marked as a hypothesis for follow-up. The released author reviews look severe enough","section":"Abstract; §1; Table 1; Table 3; §§7–8"},{"comment":"§3 documents three human interventions: OpenClaw/Anthropic thinking-block harness patch (14 resets TabPFN, 5 Personas), a 24-hour deadline extension after self-graded Weak Reject drafts, and a mandated readability rewrite of “inscrutable” prose (Figure 4). The paper argues these are logistical and do not negate autonomous research failure. That is plausible for engineering competence, but two interventions directly shape the graded object (extra time after the agent declared completion; human-requested clarity pass) and one repeatedly wiped context. A load-bearing clarification is needed: report scores or qualitative deltas on the pre-extension / pre-rewrite drafts versus finals, and state whether author grades apply only to the post-intervention PDFs. Without that, “unambiguous rejection of autonomous open-ended research” is partly confounded with “rejection after scaffold patches and h","section":"§3; Figure 4; §7.1"},{"comment":"§5.3 treats failure to abandon the approach and restart as a primary failure mode, yet §5.3 and the TabPFN narrative note the design required a paper and offered no abstain option, which “may have led the agent to write a negative-results paper.” That is a protocol confound for the backtracking claim: a human might stop or pivot without delivering a forced manuscript. Either add an explicit abstain/“return null with justification” action and re-interpret logs under that counterfactual, or downgrade “ineffective backtracking” from a model failure mode to a joint agent–protocol outcome and adjust §5’s five-mode taxonomy accordingly. As written, the causal attribution to the agent alone overreaches the design.","section":"§5.3; §4.3; Figure 2"},{"comment":"§4.4 and Figure 3 show AI self- and external reviews never accepting, and Table 4 shows partial overlap with human critiques, supporting a possible generator–verifier gap. The same subsection correctly notes both agent papers were rejects, so one cannot tell whether verifiers discriminate quality or uniformly reject. This undercuts using the review stack as evidence that agents “could not make good use of feedback” in a way that implies the feedback was calibrated gold. Tighten the claim to: agents did not respond to recurring soundness critiques with redesign (supported by logs), and separately mark verifier accuracy as unestablished. Avoid implying the AI-review panel is a validated training signal for RL until accept-class agent or human papers are graded by the same stack.","section":"§4.4; Figure 3; Table 4"}],"minor_comments":[{"comment":"Table 1 summary scores are clear; consider adding a one-line note that Personas has since been made public (Baines et al., 2026) while TabPFN details remain restricted, so external readers cannot equally audit both agent papers.","section":"Table 1; footnote 3"},{"comment":"Figure 1 GPU budget annotations ($392 of $500 vs $69 of $100) are hard to compare across papers; state explicitly how GPU dollar caps were set relative to author estimates in §3.","section":"Figure 1; §3"},{"comment":"§4.7’s definition of reward hacking (deceiving the verifier into a high score) is useful; cross-reference it when collaborators’ “premature negative result ≈ reward hacking” view is mentioned so readers see the operational criterion.","section":"§4.7"},{"comment":"Appendix A research questions are appropriately detailed; ensure the main text points to them early when introducing Personas vs TabPFN so readers need not reach the appendix to understand task open-endedness.","section":"§3; Appendix A"},{"comment":"Minor copyediting: spacing quirks around “W e”/“T o” and similar artifacts appear throughout the compiled text; clean for camera-ready.","section":"passim"}],"recommendation":"major_revision","confidential_remarks":"Fit is strong for an empirical methodology / evaluation venue in AI; as a pure “capabilities” claim the n=2 non-blind design will draw skepticism regardless of hedging. I would accept a revised version that either adds independent grading or sharply scopes the lifecycle language. The artifact release and author reviews already make the case studies more inspectable than most agent demos; that is the paper’s main editorial asset. No integrity red flags beyond the acknowledged interventions and forced-output design."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a real methodological contribution plus two carefully documented negative case studies, not a definitive map of agent R&D ability. Shadow evaluation—unpublished NeurIPS-quality questions, original authors as graders, full logs released—is the part that will stick. It sits cleanly between verifier benchmarks and lottery-style blind workshop submissions, and the authors are explicit about what each method can and cannot show.\n\nWhat they did well: the agents actually cleared the engineering bar (GPU debugging, literature, camera-ready LaTeX, multi-day runs) with almost no hand-holding, then still got unambiguous author rejects (2/6 and 1/6) for judgment, underpowered/synthetic design, and shallow pivots. The five failure modes are grounded in trajectories and resource plots, not slogans. The Codex/GPT-5.6 Sol check reproducing the same pattern matters. Pre-experiment survey priors, full reviews, and artifact release make the work checkable. The central claim is correctly hedged as early evidence.\n\nSoft spots, in proportion: n=2 and non-blind original-author grading are the load-bearing limits. Authors already committed to their own framings; that can bias against alternate approaches even when the agent papers look weak on the page. Human interventions (deadline extension, clarity rewrite, harness patch) and no-abstain paper requirement shape what “failure” means. Elicitation is not fully closed—OpenClaw quirks, one robustness scaffold with less tuning. None of this invents the rejects; the released reviews are severe on quality and reasoning. It does mean the leap from “these two runs failed publishably” to “today’s agents struggle with critical parts of the research lifecycle” is still a thin bridge. The paper mostly owns that in §§7–8 and Table 3.\n\nWho it’s for: people measuring AI R&D automation, lab strategy, and RSI forecasts who need open-ended evidence rather than another hill-climb score. Math/data/citations look fine for an empirical methods-and-logs paper; no circular fitting.\n\nI’d engage. Send it to peer review. I’d want larger n, secondary blinded experts, and stronger vendor scaffolds next—but this already deserves referee time and a reading-group slot.","headline":"Useful new eval design plus honest negative case studies; the n=2 non-blind bridge is thin but the paper already says so, and the released rejects/logs still earn a serious read.","tokens_in":29596,"tokens_out":570,"would_cite":true,"duration_ms":18628,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Today's frontier AI agents can finish the engineering of open-ended AI research but still cannot answer the research questions at a publishable standard.","keywords":["AI agents","AI R&D automation","shadow evaluation","open-ended research","research engineering","failure modes","peer review","long-horizon agents"],"falsifier":"A shadow evaluation on a comparable unpublished top-conference question where a frontier agent under similar time and compute budgets produces a paper the original authors score as at least a weak accept, or a larger series of such runs that consistently clear that bar.","tokens_in":29446,"feed_emoji":"🧪","tokens_out":910,"duration_ms":40210,"temperature":0.7,"pith_summary":"Forecasts of explosive AI progress often assume agents will soon automate AI research itself, yet most tests either score narrow verifiable metrics or rely on noisy conference peer review. This paper offers a third measurement: shadow evaluations, in which an agent receives the central open-ended question from a high-quality unpublished paper, multi-day time and large compute budgets, and is graded by the original authors as if for a top conference. On two unpublished NeurIPS-level questions, frontier agents completed literature review, GPU work, experiments, and paper writing without human help, but made little substantive progress on the science; both papers were unambiguously rejected. The authors document five recurring failure modes—poor judgment of the publishable bar, uncreative responses to design flaws, weak backtracking, poor resource awareness, and instruction drift—and reproduce them with a second model and scaffold. The result is early evidence that research engineering is ahead of the judgment-heavy parts of the research lifecycle that self-improving AI forecasts depend on.","feed_headline":"AI agents finish the engineering but fail the research","feed_subtitle":"On two unpublished NeurIPS questions, original authors reject both agent papers after six-day runs.","key_machinery":"Shadow evaluations: assign a frontier agent the central open-ended research question of a high-quality unpublished paper, give it multi-day wall-clock time and large API and GPU budgets on a general-purpose scaffold, then have the paper's original authors grade the agent's output as a conference submission.","core_discovery":"Well-resourced frontier agents can autonomously perform the engineering steps of open-ended AI research, yet cannot make substantial progress toward answering those research questions at top-conference quality. In two shadow evaluations on unpublished NeurIPS submissions, original authors rejected both agent-written papers, citing weak experimental judgment, thin novelty, and unclear writing despite successful end-to-end execution.","pith_inferences":["If engineering keeps improving while judgment, backtracking, and resource sense lag, recursive self-improvement may accelerate narrow optimization faster than open-ended discovery.","Because both agent papers were rejects, the study cannot yet tell whether AI reviewers are truly discriminative or merely harsh—closing that generator-verifier gap would need accepted as well as rejected drafts.","Instruction drift and unfinished budgets look like near-term scaffold and training targets that might raise performance without solving creativity.","Results on two empirical NeurIPS-style questions may understate agent skill on more metric-driven or hill-climbable slices of AI R&D."],"forward_implications":["Claims that AI will soon automate AI R&D must separate engineering automation from open-ended scientific judgment and redesign.","Verifier-only and blind-review evaluations will systematically miss the failure modes that shadow evaluations expose.","More wall-clock time or raw compute alone is unlikely to fix premature commitment, weak backtracking, and unused budgets.","Self-review loops that reliably reject weak drafts do not by themselves force creative redesign of the research approach.","Released logs, reviews, and repositories make the method reusable as stronger models and scaffolds appear."],"fun_headline_variants":["Agents nail the code, fail the science: authors reject both papers","Shadow evals: frontier agents can't clear NeurIPS research bar","Six days, full compute: agents finish engineering, stall on questions","Authors reject agent papers: judgment and novelty gaps doom both","Open-ended AI research still blocks agents despite full engineering"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That unambiguous rejection by the original authors on two non-blind case studies is a reliable enough signal of agent research inability, rather than mainly author preference, scaffold under-elicitation, or the particular questions chosen.","fun_headline_variants_meta":{"raw":{"variants":["Agents nail the code, fail the science: authors reject both papers","Shadow evals: frontier agents can't clear NeurIPS research bar","Six days, full compute: agents finish engineering, stall on questions","Authors reject agent papers: judgment and novelty gaps doom both","Open-ended AI research still blocks agents despite full engineering"]},"model":"grok-4.5","effort":"low","cost_usd":0.001928,"raw_usage":{"total_tokens":895,"prompt_tokens":806,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":19284000,"prompt_tokens_details":{"text_tokens":806,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":806,"tokens_out":69,"duration_ms":2497,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T10:53:55.290705+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A shadow evaluation on a comparable unpublished top-conference question where a frontier agent under similar time and compute budgets produces a paper the original authors score as at least a weak accept, or a larger series of such runs that consistently clear that bar.","supporting_citations":[],"review_version":2}