{"id":"240b3239-d2f6-40d8-b960-a15a12b0c424","arxiv_id":"2608.13331","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.","lead":"This paper builds Replica, a benchmark of 310 tasks that ask AI agents to re-create a missing figure from a research paper, and trains Faraday, a 27B-parameter model that directs a larger coding agent. Faraday scores above Claude Opus 4.8 and GPT-5.5 on held-out replication tasks according to the paper's rubric judge.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The rubric judge supplies both the training reward and the headline test metric, and the evidence that it tracks human replication quality is not statistically significant (p = 0.109); Faraday's advantage may reflect learned judge-pleasing rather than human-valued replication.","rationale":"The reader's weakest assumption identifies exactly the load-bearing issue: the rubric judge's validity as a measure of human-valued replication. My stress-test agrees and sharpens the point. The paper has genuine strengths that should be credited: a held-out test split, a prompt-optimization control, training/test separation of task domains, eight rollouts per task, disclosed limitations, and two human studies. However, those controls do not fully sever the circularity between the training reward and the evaluation metric. The judge-comparison study is underpowered (p = 0.109), confined to the training distribution, and selected for judge disagreement, which limits what it can establish. The prompt-optimization control is a strong and honest design choice, but it can only rule out prompt-level overfitting, not the deeper policy-level reward hacking that RL is known to amplify. The agent-comparison human study is explicitly limited to high-margin rollouts, so it cannot support the average superiority claim. If the proposed test fails, the headline claim would need to be weakened to 'better at optimizing the rubric judge,' which would not support the paper's broader conclusions about scientific generalization and human-level replication taste. Accordingly, the conditional verdict stands: the central claim is plausible and internally controlled, but the load-bearing evidence for judge validity is not yet statistically convincing. No new objection beyond the reader's was found, so no verdict change is warranted.","tokens_in":46913,"tokens_out":4640,"duration_ms":46854,"concrete_test":"Pre-register a blind human ranking study on a random sample of 16 test-split tasks (not filtered by judge margins), with at least five expert raters per task ranking eight rollouts each from Faraday, Claude, and Codex. Compute (i) the proportion of tasks where humans rank Faraday above both baselines, and (ii) per-task Kendall tau between human rankings and rubric-judge scores. Independently, re-score all evaluation rollouts with a rubric judge generated from a different meta-rubric or judge model (e.g., Claude-based instead of Codex-based) and compare the Faraday-vs-baseline gap. If humans show no average preference for Faraday on the test split, or if judge scores do not correlate with human rankings on these randomly selected tasks, the headline advantage should be attributed to judge optimization rather than replication quality.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Faraday replicates research better than Claude and Codex, but the only quantitative support is the rubric judge's scores (Section 4.3, Figure 2), and the same judge provided the GRPO reward during post-training (Section 3.2). The human-validation evidence in Section 4.1 / Appendix F.1 is weak: humans side with the rubric judge over a baseline judge on only 63% of disputed pairs, p = 0.109, and rubric-human Kendall tau is 0.19 versus 0.15 for the baseline. That study is also restricted to ten train-split tasks selected for judge disagreement, so it does not establish judge validity on the AI-for-science test split where the out-of-distribution claim is made. The prompt-optimization control (Figure 5, left) rules out prompt-level judge-pleasing, but not policy-level learned judge-pleasing: RL training can amplify subtle stylistic and procedural cues that the judge rewards and that a static prompt cannot reproduce. Because the same judge is used at train and test time, any such bias is inherited directly by the evaluation, so the reported 73%/60% task-level superiority and 6-8% average gains could be an artifact of optimizing the judge rather than a genuine human-valued replication advantage. The human-agent comparison (Section 4.4, Appendix F.2) is conditional on the judge's high-margin selection (≥0.2 above both baselines) and explicitly does not establish average human preference, as the paper itself notes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Replica, a task space of 310 figure-replication tasks derived from 100 ML and AI-for-science papers, where each task asks an agent to reproduce a redacted results figure from a paper within a one-hour budget. To supply training signal, the authors build an auto-generated, per-task rubric-based judge implemented with Codex GPT-5.5, which scores rollouts along five dimensions and produces turn-level credit weights. They post-train Faraday, a 27B-parameter model, with a modified GRPO recipe using this judge as reward, and report that Faraday outperforms Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution tasks and 60% of held-out AI-for-science tasks according to the same rubric judge, with average gains of 6% and 8% respectively. The paper also includes human studies aimed at validating the judge, a prompt-optimization control, qualitative case analyses, and generalization experiments to larger compute budgets and stronger coding tools.","tokens_in":47184,"tokens_out":2556,"duration_ms":28905,"significance":"If the central claim holds, the paper would be a meaningful step toward training small 'AI scientist' agents that can direct frontier coding agents on long-horizon, underspecified scientific tasks. The strengths of the paper are real: the Replica task space is scalable and automatically generated, the training recipe is described in unusual detail with stage-by-stage hyperparameters, the ablations (turn-level credit assignment, coding-agent tool, judge noise) are informative, and the authors include machine-checkable experiments and a prompt-optimized baseline that rules out one class of judge-pleasing. However, the significance is contingent on whether the rubric judge measures human-valued replication quality. The current evidence for that validity is thin, and because the same judge is used for both the training reward and the headline evaluation metric, the reported advantage is vulnerable to learned judge-pleasing that would not transfer to human judgment.","major_comments":[{"comment":"The rubric judge is both the GRPO training reward (Section 3.2 and 3.5) and the sole evaluation metric for the headline comparison (Section 4.3, Figure 2). Consequently, any policy-level reward hacking or judge-pleasing behavior learned during RL is inherited directly by the test evaluation, because the same judge scores the held-out rollouts. The prompt-optimization control in Figure 5 (left) rules out prompt-level judge-pleasing, but it does not rule out policy-level learned behaviors, such as producing certain procedural artifacts or stylistic traces that the judge rewards. The paper's central claim of superiority over Claude and Codex therefore rests entirely on the untested assumption that the judge is not exploiting these artifacts. To support the claim, the authors need either an independent evaluation metric (e.g., a different judge model or a non-LLM objective measure) on the test split, or a human-validated judge on the test split itself.","section":"§3.2, §3.5, §4.3"},{"comment":"The human-validation evidence for the rubric judge is not statistically significant and is restricted to the training split. Participants side with the rubric judge over the baseline judge on only 63% of disputed pairs, with p = 0.109 from a binomial mixed-effects model, and the rubric-human Kendall tau is 0.19 versus 0.15 for the baseline judge, a difference of 0.04 with no reported confidence interval. The ten tasks used in this study are all drawn from the train split and were selected specifically for judge disagreement, so the study does not establish judge validity on the AI-for-science test split where the out-of-distribution claim (60% task superiority) is made. The abstract's statement that the judge 'agrees with human assessment of replication quality' is stronger than this evidence supports.","section":"§4.1, Appendix F.1"},{"comment":"The agent-comparison human study is conditional on the rubric judge's high-margin selection: only rollouts where Faraday scores at least 0.2 above both Claude and Codex on the judge's scale are shown to humans. The paper explicitly and correctly notes that this design cannot support conclusions about average human preference. As external grounding for the judge, this study is therefore a sanity check that the judge's most confident calls align with humans, but it does not validate the judge for the typical cases that drive the aggregate results in Figure 2. The 71% preference rate among the 41 selected rollouts is compatible with a judge that is only directionally correct on extreme margins while being wrong or noisy on the majority of comparisons.","section":"§4.4, Appendix F.2"},{"comment":"The discussion of pre-training contamination in Appendix D.1 addresses the agent's exposure to the papers but does not address the judge's exposure. The judge is a Codex model that has likely seen many of the Replica papers, including their figures, in its pre-training data. Because the judge is given the gold plot and the redacted paper, it could reward rollouts that happen to match its memorized representation of the original figure, rather than rollouts that faithfully replicate the underlying experiment. Since the judge is also the training reward, this could bias Faraday toward mimicking memorized outputs rather than learning robust replication skill. The authors should at least analyze whether judge scores correlate with the judge model's (or the rubric generator's) familiarity with the paper, e.g., by publication year or by held-out citation counts, or by comparing judge scores on papers published after the judge's knowledge cutoff.","section":"§3.2, Appendix D.1"}],"minor_comments":[{"comment":"The reported Kendall tau values of 0.19 (rubric vs humans) and 0.15 (baseline vs humans) are shown as single numbers with no confidence intervals or significance tests; adding these would help readers assess the practical difference.","section":"Figure 3 (left)"},{"comment":"The qualitative case analyses appear to be based on the authors' own inspection of rollouts; please state whether the inspection was performed blind to the model identity, or at least acknowledge the potential for confirmation bias in selecting and describing these examples.","section":"Table 1"},{"comment":"The Faraday system prompt contains the literal template variable '{coding_agent_budget} tokens', which suggests an unresolved placeholder; this should be rendered as a concrete value in the published appendix.","section":"Appendix G.1"},{"comment":"The abstract claims the judge 'agrees with human assessment of replication quality' and Section 4.3 describes a 'comprehensive uplift' in performance; both statements should be tempered to reflect the non-significant human-agreement result and the conditional nature of the human preference study.","section":"Abstract and §4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful in its internal setup, but the evaluation independence is the core weakness. The same LLM-based judge is used as reward, evaluation metric, and (indirectly) as the basis for human-study selection, and the external validation is not statistically significant. I would encourage the editor to weigh whether the authors can add an independent judge or a non-LLM metric within a revision cycle; short of that, the central claim of human-valued replication superiority remains unsubstantiated. The authors' own noted limitation in Section 4.4 that the study does not establish average human preference is commendable honesty, but it does not cure the circularity in Section 4.3."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real contribution, and the authors are unusually honest about its limits, but the headline claim—that Faraday replicates better than Claude and Codex—rests on a judge that was also the training reward, and the human-validation evidence is statistically weak. The paper deserves a serious referee, but the referee should push on the validity of the judge.\n\nWhat's new: Replica, 310 figure-replication tasks from 100 papers, is a genuinely scalable task space; the auto-generated per-task rubrics are a reasonable way to grade non-verifiable replication. The CAT (coding-agent-as-tool) recipe is a clean combination: a 27B model directing GPT-5.5, post-trained with GRPO, turn-level credit assignment, and multi-sample judge aggregation. The ablations are a strength—removing turn-level credit collapses training, and prompt-optimizing the Codex baseline doesn't close the gap. The internal evaluation is carefully set up: held-out test split, 8 rollouts per task, known baselines, per-dimension score decomposition. The prompt-optimization control rules out prompt-level judge-pleasing.\n\nSoft spots, in order: (1) The judge is both the GRPO reward and the headline metric. The same judge at train and test time means the reported 73%/60% advantage could partly be learned judge-pleasing. The paper doesn't fully address that—RL can exploit procedural or stylistic cues a static prompt can't. (2) The human validation for the judge is thin. Rubric-human Kendall tau is 0.19 vs 0.15 for baseline, and humans side with the judge on only 63% of disputed pairs, p=0.109—not significant. That study is also restricted to ten train-split tasks selected for judge disagreement, so it doesn't validate the judge on the AI-for-science test split where the OOD claim is made. (3) The human preference study only covers rollouts where the judge already declared Faraday at least 0.2 above both baselines; the authors themselves note it says nothing about average human preference. (4) No artifacts released, so the judge and rubrics can't be independently audited. (5) The 'innovation' tasks are scored by the same judge, which was never validated on those.\n\nNone of this is fatal—the paper is honest about most of it, and the qualitative rollouts in Table 1 are genuinely suggestive. But the central quantitative claim is conditional on judge validity, and the evidence for that validity is currently weak.\n\nI'd bring this to a reading group and would cite it as a benchmark and training recipe. It deserves peer review: the task space alone is valuable, and the recipe is reproducible in principle. Recommend engaging, with the judge-validity issue as the central review point.","headline":"Solid, honest, partly circular: the same rubric judge supplies reward and headline metric, human validation is weak (p=0.109), but the task space and CAT training recipe are real contributions worth refereeing.","tokens_in":47816,"tokens_out":2135,"would_cite":true,"duration_ms":21650,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that Faraday, a 27B agent post-trained on Replica, a space of 310 figure-replication tasks, scores above Claude Opus 4.8 and GPT-5.5 on most tasks according to an auto-generated rubric judge, and that this advantage…","keywords":["paper replication","AI scientist agents","rubric-based reward","reinforcement learning post-training","coding agents as tools","scientific reproducibility","figure replication benchmark","GRPO"],"falsifier":"Take a random sample of Replica test tasks (not selected for large judge margins), have expert humans blindly rank rollouts from Faraday, Claude, and Codex, and compare human rankings with the rubric judge's scores; if human-judge agreement on the full sample is no better than the baseline judge's agreement, or if humans prefer the frontier baselines on average, then the rubric judge is not capturing human replication taste and Faraday's headline advantage is an artifact of the reward loop.","tokens_in":46661,"feed_emoji":"🔬","tokens_out":7127,"duration_ms":64006,"temperature":0.7,"pith_summary":"The paper tries to establish that paper replication, an underspecified and open-ended task, can be turned into a scalable training ground for AI research agents. It introduces Replica, a task space of 310 automatically generated tasks in which an agent must re-run the experiments behind a redacted results figure from one of 100 machine-learning and AI-for-science papers. To score attempts, it builds an auto-generated per-task rubric judged by a coding agent with access to the rollout, and it post-trains Faraday, a 27B model that acts as a research planner and delegates coding to a frontier coding agent. On the paper's rubric scores, Faraday beats Claude Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and on 60% of held-out AI-for-science tasks, with average gains of 6% and 8% on the test split. If the rubric captures human scientific taste, the result suggests that a small, inspectable 'scientist layer' can direct much larger coding models at tasks too underspecified for conventional verifiable benchmarks.","feed_headline":"A 27B model trained on replication beats Claude and Codex","feed_subtitle":"Trained on 310 figure-replication tasks, it wins 60% of held-out AI-for-science papers by rubric score.","key_machinery":"The load-bearing mechanism is the per-task rubric judge: an auto-generated scoring guide covering five dimensions, executed by a coding agent that inspects the rollout's container, code, git history, and the gold figure, and sampled three times to reduce noise. This judge supplies the GRPO reward, so its validity determines everything downstream. Around it sits the coding-agent-as-tool (CAT) setup, in which the 27B Faraday model plans experiments and delegates implementation to a frontier coding agent through a shell tool, plus a training recipe with turn-level credit assignment weights from the judge, normalized per token, that stabilizes long-horizon, non-verifiable reinforcement learning.","core_discovery":"The paper's central claim is that long-horizon reinforcement learning against a rubric-based judge turns a 27B base language model into an agent that replicates research figures more faithfully than frontier coding agents. Replica tasks are constructed by redacting one results plot from a paper and requiring the agent to produce the plot by running real experiments within an hour on a single GPU slice, scaling down when necessary. The judge, a coding agent scored against a per-task rubric auto-generated by Claude Opus 4.7, grades five dimensions: visual fidelity, claim reproduction, implementation fidelity, experimental depth, and scientific integrity. Faraday, post-trained with a turn-level credit variant of GRPO using three judge samples per rollout, outperforms Claude Opus 4.8 and GPT-5.5 under the same budget; qualitative inspection of rollouts shows Faraday implementing the mechanism behind the claim rather than hard-coding outputs, and it transfers to held-out AI-for-science papers and to variants with changed claims or datasets (19 of 20 'imagined' tasks). The paper is careful to note that the headline comparison is made by the rubric judge itself, with only partial human validation.","pith_inferences":["A skeptical reading: because the same rubric judge supplies the reward and the headline evaluation, some of Faraday's edge may be judge-pleasing behavior; a decisive test would rerun the main comparison on a randomly selected sample of tasks with human expert ranking as ground truth, rather than only on the strong-margin subset.","If the judge's human-agreement evidence is as weak as it appears (participants side with the rubric judge on 63% of disputed pairs, p = 0.109), averaging three judge samples could inflate Faraday's apparent advantage by rewarding consistency that humans do not value.","A natural extension is to scale Replica toward full-paper replication and multi-figure consistency, where the judge would have to check cross-figure coherence; if Faraday's skill transfers there, the paper's thesis that replication is a curriculum step toward innovation would gain support.","The CAT architecture suggests a cost-benefit question the paper leaves open: how to split size and capability between an outer judgment model and an inner execution model, and whether the outer model can remain small, open, and inspectable as coding tools improve."],"forward_implications":["If the rubric judge tracks human taste, replication skill can be trained at scale without hand-verified rewards; the paper estimates that top ML conferences alone could yield about 36,000 new tasks per year.","A small open-weights model can direct a frontier coding agent and beat the frontier agent running alone, so post-training a 'researcher layer' is a viable alternative to engineering specialized harnesses.","The learned behavior transfers to papers outside the training distribution (AI-for-science) and to altered claims or datasets, indicating generalization rather than memorization.","Swapping the inner coding agent for a stronger one at evaluation time improves performance, so the trained scientist layer can ride improvements in frontier coding models without retraining.","The training recipe with three judge samples and turn-level credit stabilizes GRPO in a long-horizon, non-verifiable setting; without it, the paper's ablation shows training collapsing after about 50 steps."],"supporting_citations":[{"why":"Supplies Gemini 2.5 Pro, the vision-language model that auto-generates Replica tasks by detecting and redacting figures from papers.","marker":"Comanici et al., 2025"},{"why":"Provides the GRPO algorithm that Faraday's post-training modifies with turn-level credit assignment and multi-sample judge aggregation.","marker":"Shao et al., 2024"},{"why":"Defines the Codex CLI schema that both the Faraday harness and the coding-agent judge are built around.","marker":"OpenAI, 2025"},{"why":"Names GPT-5.5, the frontier coding agent used as Faraday's tool and as the coding-agent judge, and one of the two main baselines.","marker":"OpenAI, 2026"},{"why":"Names Claude Opus 4.7 and 4.8, used for rubric generation and as the strongest baseline that Faraday must beat.","marker":"Anthropic, 2026"},{"why":"Provides Qwen3.6-27B, the open base model that is post-trained to become Faraday.","marker":"QwenTeam, 2026"},{"why":"Supplies LoRA, the parameter-efficient adaptation method used to update Faraday's weights during post-training.","marker":"Hu et al., 2022"},{"why":"Provides PaperBench, the closest benchmark judging agent replication from papers, which Replica extends with larger scale and per-task rubrics.","marker":"Starace et al., 2025"}],"fun_headline_variants":["RL-trained 27B agent beats Opus and GPT-5.5 at replication","27B model outwits Claude and GPT on figure replication","Replication-trained AI transfers to new papers, beats big models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire result rests on the rubric judge measuring true replication quality: the paper's own human-agreement check is weak (participants side with the judge on 63% of disputed pairs, a result not statistically significant), so if the judge rewards judge-pleasing behavior instead of faithful reproduction, Faraday's edge could vanish.","fun_headline_variants_meta":{"raw":{"variants":["RL-trained 27B agent beats Opus and GPT-5.5 at replication","27B model outwits Claude and GPT on figure replication","Replication-trained AI transfers to new papers, beats big models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000451,"raw_usage":{"total_tokens":2267,"prompt_tokens":933,"completion_tokens":1334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":1273}},"tokens_in":549,"tokens_out":1334,"duration_ms":11666,"temperature":1.0,"reasoning_tokens":1273,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:31:07.562137+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of Replica test tasks (not selected for large judge margins), have expert humans blindly rank rollouts from Faraday, Claude, and Codex, and compare human rankings with the rubric judge's scores; if human-judge agreement on the full sample is no better than the baseline judge's agreement, or if humans prefer the frontier baselines on average, then the rubric judge is not capturing human replication taste and Faraday's headline advantage is an artifact of the reward loop.","supporting_citations":[],"review_version":1}