{"id":"0118255a-01d1-4709-9250-e3c222896e18","arxiv_id":"2607.02931","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"VERITAS, a CLI-agent replication framework, leads every reported metric on CORE-Bench Hard and ReplicationBench against matched Claude Code baselines across 65 papers in four domains.","lead":"VERITAS is an end-to-end, domain-agnostic pipeline that uses CLI coding agents to extract a paper’s claims, re-run its methods while fixing failures, and score each claim against run evidence. If it generalizes beyond the two benchmarks and one model family tested, it would give reviewers and labs a practical tool for automated scientific verification.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Native-score “leads every metric / 97.8% SOTA” rests on single-run point estimates with tiny margins and a σ-rule that re-credits only VERITAS.","rationale":"The reader correctly flags that ANALYZE, the importance-weighted Replication Score, and the fix log are not what the tables measure, and that same-family single-run baselines limit the generality claim—hence CONDITIONAL is right. That concern is load-bearing for “general-purpose tool,” but for the stated strongest claim (SOTA / leads every metric on the two benchmarks) the sharper soft spot is statistical and presentational: native leads are on the order of a few capsules/tasks under single-run stochastic agents, and the 97.8% figure in Figure 1 is produced by a uniform rule whose empirical effect is entirely VERITAS-sided. Faithfulness and cheating still favor VERITAS and are less sensitive to one-bit pass/fail thresholds, so the systems contribution is not empty. Multi-run official-scorer means would settle whether “leads on every metric” survives noise; if it does, confidence can rise without changing the conditional posture on product-scope evaluation. No internal contradiction or circular scoring of the pipeline against itself appears. Verdict stays CONDITIONAL; no move to REJECT or ACCEPT.","tokens_in":21920,"tokens_out":648,"duration_ms":33189,"concrete_test":"Run VERITAS, ZEROSHOT, and RETRY for 3 independent seeds each on CORE-Bench Hard (45 capsules) and ReplicationBench paper-only (20 papers), same model/images. Score only with official graders (no σ-correction); report mean±std pass/match rates with and without cheating correction. If VERITAS’s mean does not strictly exceed both baselines on both benchmarks’ primary native metrics, the strongest claim needs qualification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that VERITAS leads every reported metric and is SOTA vs matched Claude Code baselines. Under the official CORE-Bench scorer the lead is 88.9% vs 84.4% (RETRY)—about two capsules on n=45. On ReplicationBench paper-only the lead is 0.333 vs 0.315 (~2 of 111 tasks); full mode is similarly thin. All cells are single-run; Claude Code exposes no pinable temperature, and several papers are stochastic (MCMC, unseeded training). Figure 1’s 97.8% further applies a zero-variance re-scoring rule that flips four VERITAS failures and zero baseline failures (§D), so the headline gap is partly rule-driven. Trajectory metrics (faithfulness, cheating) separate systems more cleanly and are better support; the native-score SOTA wording is the fragile part. The ANALYZE/Replication-Score bypass (§3.1) is appropriate for matched claims but means the product-facing pipeline is not what the tables validate—secondary to the robustness issue for the claim as written.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces VERITAS, a domain-agnostic scientific replication framework built around CLI coding agents (Claude Code, with support for Codex and Gemini). Given a paper, a repository, or both, the system extracts structured claims (ANALYZE), optionally generates code (CODEGEN), plans and executes the methodology while actively fixing failures under a bounded manager loop (PLAN/REPLICATE), rates applied fixes (ASSESS FIXES), and judges each claim against run evidence (VERIFY), returning an importance-weighted Replication Score, a severity-rated fix log, and a patched codebase. The authors evaluate the replication core on CORE-Bench Hard (45 capsules; CS, social science, medicine) and ReplicationBench (20 astrophysics papers / 111 tasks), comparing against two Claude Code baselines (ZEROSHOT, RETRY) on the same model (Claude Opus 4.8) and host environment. After a uniform cheating correction, VERITAS leads reported native and trajectory metrics; under a zero-variance re-scoring rule it reaches 97.8% per-capsule pass rate on CORE-Bench Hard (88.9% under the official scorer).","tokens_in":22208,"tokens_out":1401,"duration_ms":24196,"significance":"If the empirical picture holds under more robust evaluation, this is a useful systems contribution: prior replication agents are largely benchmark-tied or domain-specific, whereas VERITAS targets flexible inputs (paper+code, paper-only, repo-only) and end-to-end claim verification. Strengths include matched model/environment controls, system-blind adapted cheating and faithfulness metrics from prior work, transparent cheating correction, per-cell tables, and substantive case studies and error analysis (e.g., instructed external loads, paper-only data upper bounds, multi-value all-or-nothing scoring). The trajectory-level separation—especially faithfulness and lower cheating—is more informative than the thin native-score margins and is a genuine process-quality contribution for agentic replication research.","major_comments":[{"comment":"Abstract, Figure 1, and Table 1: the headline 97.8% CORE-Bench Hard pass rate is produced by a zero-variance re-scoring rule (§D) that flips four VERITAS failures and zero baseline failures. Under the official scorer VERITAS is 88.9% vs RETRY 84.4% and ZEROSHOT 80.0%—a lead of roughly two capsules on n=45. Leading the abstract and Figure 1 with the re-scored number, and the unqualified claim that VERITAS “achieves state-of-the-art performance and leads on every metric,” overstates a fragile single-run margin. Primary claims should use the official scorer; the σ-rule may appear as a secondary, uniformly applied analysis with explicit cross-system comparability caveats already noted in §D.","section":null},{"comment":"§3.1 “Scope and scoring”: benchmark task questions are supplied as claims, bypassing ANALYZE and the importance-weighted Replication Score (Eq. 1); the authors state that neither the Replication Score nor the severity-rated fix log is used to compute any reported number. Yet the abstract and contribution bullets present claim extraction, the Replication Score, and the fix log as core delivered outputs of a general-purpose tool. This is a load-bearing gap between product framing and what Tables 1–2 validate (PLAN/REPLICATE/VERIFY). Either evaluate ANALYZE quality and score calibration on a held-out claim set, or reframe the paper’s claims around the replication core that was actually measured.","section":null},{"comment":"§3 Experiments and Limitations (“Run-to-run variance”): all native scores are single-run point estimates; Claude Code exposes no pinable temperature, and both benchmarks include stochastic computations (MCMC, unseeded training). On ReplicationBench paper-only the match-rate lead is 0.333 vs 0.315 (~2 of 111 tasks); full-mode gaps are similarly small before trajectory metrics. Given that native “leads every metric” is central to the abstract, the manuscript needs multi-run averages, uncertainty (e.g., bootstrap over capsules/tasks), or explicitly softer language that privileges faithfulness/cheating—where separation is large and better supported—over native SOTA wording.","section":null}],"minor_comments":[{"comment":"Figure 1 caption and §3.2: make the dual reporting (official 88.9% vs σ-corrected 97.8%) visually primary rather than star-noted, so readers cannot miss which number is comparable to the HAL leaderboard.","section":null},{"comment":"Eq. (1) and §2.2: importance weights (headline=3, supporting=2) and verdict values (match=1, partial=0.5) are free parameters; a short sensitivity check or justification would help even though the score is not used in the tables.","section":null},{"comment":"§B.1 ReplicationBench shape-coercion patch: state clearly in the main text (not only the appendix) that a deterministic scorer patch was applied, and confirm it was applied identically to all systems.","section":null},{"comment":"§F.3 / Limitations: the hardcoding blind spot of the access-only cheating metric is well illustrated by fable_mps; consider promoting a one-sentence main-text caveat next to Table 2’s cheating column so readers do not over-read the count.","section":null},{"comment":"Related Work: a compact comparison table (inputs, domain, execution, claim grading, standalone vs benchmark-tied) would make the “no general-purpose tool” positioning easier to audit against Kohler et al., AutoReproduce, PaperRepro, and PaperBench agents.","section":null},{"comment":"Minor consistency: “VERITAS” sometimes loses spacing in the PDF text (“VERITASextracts”, “VERITASon”); clean typesetting throughout.","section":null}],"recommendation":"major_revision","confidential_remarks":"Solid systems paper with careful trajectory evaluation; the main risk for the venue is overclaim relative to single-run thin native margins and unevaluated product surface (ANALYZE / Replication Score). I would not reject on novelty or soundness of the pipeline design. If the authors temper native SOTA language, lead with official scorers, and either evaluate or demote unevaluated components, this could become a clear accept after revision. Fit is appropriate for an AI/systems venue interested in scientific automation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: VERITAS is the first domain-agnostic, CLI-agent replication framework that actually ships flexible inputs (paper+code, paper-only, repo-only), claim withholding, a manager loop, fix severity rating, and a verification stage, and it beats two strong same-model Claude Code baselines on CORE-Bench Hard and ReplicationBench across native scores and trajectory metrics.\n\nWhat is new is the packaging, not a new model. Prior work is mostly benchmark-bound companion agents or domain-specific tools. They run the methodology, fix failures, and judge claims against run evidence, with careful matched evaluation: same Opus 4.8 host, native scorers, uniform cheating correction, and system-blind faithfulness/cheating judges adapted from prior work. Trajectory metrics separate systems cleanly—faithfulness is the real win, and cheating is near zero for them. The case studies and error analysis are honest about near-misses, multi-value all-or-nothing scoring, and paper-only data bounds. Math and scoring are transparent weighted averages; citations look appropriate.\n\nSoft spots, in proportion: the headline “leads every metric / 97.8% SOTA” rests on single-run point estimates with thin native margins (88.9 vs 84.4 under the official CORE scorer; ~2 tasks on ReplicationBench paper-only). The σ-rule that re-credits four VERITAS capsules and zero baselines is documented and uniform, but it inflates the Figure 1 gap. They also bypass ANALYZE and the user-facing Replication Score for the tables, so the product story is broader than what the numbers validate. Baselines stay inside one model family. None of that is a load-bearing flaw; it is evaluation scope and robustness, which they mostly flag in Limitations.\n\nThis is for people building AI-for-science tooling and research-integrity automation. A serious editor should send it to referees. I would engage with the work, cite the pipeline design and trajectory metrics, and treat the native SOTA claim as directionally right but not yet settled until multi-run and broader baselines exist.","headline":"Solid systems paper: a real general-purpose replication pipeline that beats matched Claude Code baselines, with the SOTA native-score wording a bit ahead of the single-run margins.","tokens_in":22841,"tokens_out":538,"would_cite":true,"duration_ms":5024,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"VERITAS is a domain-agnostic coding-agent pipeline that replicates scientific papers from a paper, code, or both, and leads strong same-model baselines on two multi-domain replication benchmarks.","keywords":["scientific replication","coding agents","computational reproducibility","claim verification","agent pipelines","paper-to-code","automated verification","replication benchmarks"],"falsifier":"Re-run VERITAS and the same baselines on both benchmarks under multi-seed conditions with the same cheating correction: if VERITAS no longer leads on native pass or match rates, faithfulness, and cheating counts, the performance claim fails; if the full claim-extraction and scoring path systematically fails on papers outside those suites, the general-purpose claim fails.","tokens_in":22772,"feed_emoji":"🔁","tokens_out":1002,"duration_ms":20511,"temperature":0.7,"pith_summary":"Independent verification of published research is growing more important as AI speeds publication and review systems fall behind, yet manual replication remains slow and costly. Most existing agent-based efforts ship as fixed benchmarks whose companion agents only run inside those pipelines, and no general-purpose replication tool has been available. This paper presents VERITAS, a domain-agnostic framework built around command-line coding agents. Given a paper, a repository, or both, it extracts structured claims, plans and runs the methodology while repairing failures as they arise, rates every fix, and judges each claim against run evidence. The pipeline returns an importance-weighted Replication Score, a severity-rated fix log, and a patched codebase. On CORE-Bench and ReplicationBench—65 papers across computer science, social science, medicine, and astrophysics—VERITAS leads two strong coding-agent baselines run on the same model and host on every reported metric, including a high per-capsule pass rate on CORE-Bench Hard and the best paper-only and full-mode match rates on ReplicationBench after a uniform cheating correction.","feed_headline":"Agent pipeline leads on 65-paper scientific replication","feed_subtitle":"One tool runs papers with or without code and tops same-model baselines across four fields.","key_machinery":"The six-phase pipeline: ANALYZE extracts typed, importance-weighted claims (withheld from the replicating agent); optional CODEGEN builds a codebase from the paper alone; PLAN drafts claim-linked steps without paper values; REPLICATE executes with active fixing under a bounded manager loop; ASSESS FIXES rates each repair; VERIFY extracts replicated values and grades them into an importance-weighted Replication Score.","core_discovery":"The authors claim that a single multi-phase replication framework around CLI coding agents can handle flexible inputs (paper plus code, paper only, or code only), actively fix execution failures, and verify paper claims against experimental evidence in a domain-agnostic way—and that this design reaches state-of-the-art results and leads strong same-model baselines on every metric across two multi-domain replication benchmarks covering 65 papers.","pith_inferences":["As submission volumes rise, automated claim-level replication could become a routine pre- or post-review check rather than a rare manual project.","Withholding reported values from the executing agent is a reusable pattern for reducing leakage in any agentic evaluation of scientific results.","The same plan–replicate–verify loop the authors sketch for extension and rediscovery tasks would be a natural next stress test of whether the pipeline is truly domain-agnostic.","Binary benchmark scorers can hide near-complete replications; tools that report partial matches and fix severity may change how the community reads a ‘failed’ replication score."],"forward_implications":["Computational claims in published papers can be checked end-to-end without building a new agent for each benchmark or domain.","Paper-only versus code-available replication can be compared under one tool to measure how much authors’ source code closes the gap.","Severity-rated fix logs make the hidden cost of getting released code to run visible as part of a replication report.","Trajectory checks for forbidden access and execution-grounded faithfulness can sit beside native scores when judging agent replications.","A single patched codebase plus per-claim verdicts can be handed to reviewers or auditors as a concrete verification artifact."],"fun_headline_variants":["VERITAS tops same-model baselines across 65 papers","One agent framework replicates claims in four fields","CLI agents fix runs and score paper claims end-to-end","Domain-agnostic tool leads every metric on two benches","Paper-or-code input: agents verify science at SOTA level"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The general-purpose claim rests mainly on beating same-family coding-agent baselines on two fixed benchmarks, with claim extraction and the user-facing Replication Score bypassed in the scored runs and results reported from single trajectories.","fun_headline_variants_meta":{"raw":{"variants":["VERITAS tops same-model baselines across 65 papers","One agent framework replicates claims in four fields","CLI agents fix runs and score paper claims end-to-end","Domain-agnostic tool leads every metric on two benches","Paper-or-code input: agents verify science at SOTA level"]},"model":"grok-4.5","effort":"low","cost_usd":0.00305,"raw_usage":{"total_tokens":1077,"prompt_tokens":757,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":30500000,"prompt_tokens_details":{"text_tokens":757,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":235,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":757,"tokens_out":85,"duration_ms":3087,"temperature":1.0,"reasoning_tokens":235,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T06:00:42.032973+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run VERITAS and the same baselines on both benchmarks under multi-seed conditions with the same cheating correction: if VERITAS no longer leads on native pass or match rates, faithfulness, and cheating counts, the performance claim fails; if the full claim-extraction and scoring path systematically fails on papers outside those suites, the general-purpose claim fails.","supporting_citations":[],"review_version":1}