{"id":"dec00b29-a00d-4ef3-acf0-91237352b36e","arxiv_id":"2608.12788","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"ARAC-Bench introduces a three-stage researcher-mimicking benchmark that scores autonomous research systems on proposal, experiment, and synthesis quality; evaluated systems reach at most 67.9 out of 100, and scores correlate with PhD rankings at 0.81 on average.","lead":"Researchers built ARAC-Bench, a test that checks whether AI research tools plan, run, and explain research the way human scientists do, instead of only checking the final answer. Across 11 systems, the best score was 67.9 out of 100, and the test's results roughly match rankings from PhD students.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"ACS rubric corpus may include the same ICLR 2026 papers used as gold references, making the 0.8141 human-alignment validation potentially circular; the paper never states an exclusion, so the benchmark's central claim is not yet supported.","rationale":"Read in good faith, ARAC-Bench is a serious attempt at process-level evaluation, and the paper supplies a public dataset, a modular scoring protocol, and a sanity check against DeepReviewBench. What would have to be true for the central claim is that ACS encodes transferable expert cognition rather than the particular accepted papers being evaluated. That condition is the least secure: the text of Sections 3.1 and 3.2 makes overlap appear likely, and no exclusion or decontamination statement appears anywhere in the paper or appendices. I checked whether the 'physically preventing the model from directly accessing the real papers' sentence in Section 3.2 could cover this; it does not, because it constrains the evaluated frameworks, not the construction of ACS. The correlation with PhD rankings, while suggestive, does not independently resolve the issue, because the raters were asked to rank the same frameworks on the same human-research-alignment construct that ACS is supposed to measure. The reader's weakest assumption identifies exactly this point, so I agree. A conditional verdict is appropriate: the benchmark may be valid if the gold references were excluded from ACS construction, but that fact is not yet established and is easily checkable from the public dataset. I do not move the verdict; I would keep it conditional pending the overlap check.","tokens_in":17366,"tokens_out":6740,"duration_ms":71652,"concrete_test":"Download the released ARAC-Bench GitHub dataset and compute the set intersection between (i) the 7,000 paper IDs used to construct ACS and (ii) the 200 ICLR 2026 gold-reference IDs. If the intersection is non-empty, rebuild ACS from the corpus with all ICLR 2026 papers removed, re-derive the rubrics, and re-score at least the Proposal and Synthesis stages for the 11 evaluated frameworks; then recompute the Pearson correlation against the PhD rankings. If the best total score or the correlation changes materially (e.g., by more than 0.05), the reported alignment is inflated by rubric leakage. If the intersection is empty, the concern is resolved, but the paper should still document the exclusion criterion explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 says ACS is 'derived from deep mining of large-scale, high-quality academic corpora' of 7,000 AI papers 'accepted at NeurIPS, ICLR, and ICML in 2025 and 2026,' including rebuttals and review discussions. Section 3.2 says ARAC-Bench's gold references are '200 accepted papers from ICLR 2026.' Unless the 200 gold papers were explicitly withheld from the 7,000-paper ACS source corpus, the scoring rubric for a given gold paper can contain cognitive skills extracted from that same paper's own review thread. The paper's only stated isolation mechanism is the mid-2025 literature-search cutoff for the evaluated frameworks, which stops models from reading the gold papers but does nothing to decontaminate the ACS library. The Proposal Details score (25/100) retrieves the Top-5 ACS skills for the paper's sub-theme, and the Synthesis Methodological Deep Analysis score (15/100) uses ACS attribution skills; per-paper leakage into those skills means the benchmark may be grading against an answer key rather than an independent expert standard. The reported 0.8141 correlation with PhD rankings then cannot establish construct validity: the human raters may be rewarding the same content dimensions that the ACS rubric was built to reward. This is a dataset-design flaw, not a statistical reporting issue, and it is directly load-bearing for the 'reliable proxy for expert judgment' claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ARAC-Bench, a benchmark for evaluating autonomous research (Auto-Research) frameworks by measuring how closely their research processes align with human expert behavior. The benchmark decomposes research into three stages—Proposal, Experiment, and Synthesis—and scores each stage using rubrics derived from an Academic Cognition Skills (ACS) library. The ACS library is constructed from 7,000 papers accepted at NeurIPS, ICLR, and ICML in 2025-2026, including rebuttals and review discussions. ARAC-Bench's gold references are 200 accepted ICLR 2026 papers. The authors evaluate 11 frameworks, all mounted on Kimi-K2.6, and report that the best framework achieves only 67.9/100. They validate ARAC-Bench against rankings from ten PhD candidates and report an average Pearson correlation of 0.8141 across the three stages (0.8788, 0.6606, 0.9030). They also present ablations showing the importance of inspiration signals and the limited benefit of integrating existing academic skill packages.","tokens_in":17592,"tokens_out":5370,"duration_ms":48328,"significance":"If the validity concerns are resolved, ARAC-Bench addresses a genuine gap: current Auto-Research evaluation emphasizes outcome metrics such as code pass rates or paper completeness, whereas ARAC-Bench attempts a process-level, stage-resolved diagnosis of methodological alignment. The paper's strengths include a clearly described three-stage decomposition, a controlled-variable evaluation design with a temporal cutoff, a public dataset release, and practical ablations on inspiration signals and skill packages. The design direction—validating a rubric-based benchmark against human expert rankings—is sensible and the reported low scores for state-of-the-art frameworks are interesting. However, the central claim that ARAC-Bench is a 'reliable proxy for expert judgment' is not yet supported because of a potential overlap between the ACS source corpus and the gold-reference papers, and because the human validation rests on only ten raters with no reported uncertainty. The significance of the contribution will depend on whether the authors can establish that the rubric is not circularly derived from the evaluation targets.","major_comments":[{"comment":"The ACS library is constructed from 7,000 papers accepted at NeurIPS, ICLR, and ICML in 2025-2026, including rebuttals and review discussions, while the 200 gold-reference papers are all accepted ICLR 2026 papers. The manuscript nowhere states that the 200 gold papers were excluded from the 7,000-paper source corpus, and the temporal isolation described in Section 3.2 (search cutoff at mid-2025) does not decontaminate the rubric source. If the gold papers' own review threads contributed skills to the ACS library, then Proposal Details Scoring (25/100) and Methodological Deep Analysis (15/100) are partially grading against an answer key derived from the very papers being evaluated. The reported average correlation of 0.8141 with PhD rankings therefore cannot, as reported, establish that ARAC-Bench is an independent proxy for expert judgment; it may instead reflect that the rubric and the raters reward the same content dimensions. This is a dataset-design issue that must be resolved, for example by explicitly confirming an exclusion protocol or by re-constructing ACS without the gold papers and re-running the validation.","section":"Section 3.1 and Section 3.2"},{"comment":"The validation of the central reliability claim rests on ten PhD raters ranking ten frameworks. The paper reports Pearson correlations of 0.8788, 0.6606, and 0.9030 for the three stages and calls their average 0.8141, but it provides no confidence intervals, significance tests, or inter-rater agreement measures. With n=10, the 95% confidence interval for the Experiment-stage correlation alone would be very wide and may include values that are not 'moderately strong.' The paper also does not report the correlation between the overall ARAC total score and the human comprehensive ranking, which is the quantity most relevant to the claim that ARAC-Bench is a reliable proxy for expert judgment. Without these statistics, the validation evidence is too thin to support the abstract's 'strong average correlation' claim.","section":"Section 4.1, Table 4"},{"comment":"The total score is a weighted sum with weights 40/35/25 for the three stages, and the Proposal stage further assigns 5/25/10 to its sub-modules, with the 25-point Proposal Details score depending on a Top-5 ACS skill retrieval and three-level anchor scores (0/1.5/3). These weights and anchors are introduced without any sensitivity analysis or justification, yet they determine the framework ranking in Table 3 and the 'best alignment score of only 67.9' headline result. A small perturbation of the stage weights could reorder frameworks in the 50-60 range and change the reported gap. The authors should provide a sensitivity analysis or justify the weights from data (e.g., fitted to the human rankings) rather than asserting them.","section":"Section 3.2, stage weights"},{"comment":"The paper claims that ACS achieves 'the SOTA 63% consistency rate' on '200 paper queries of unseen ICLR papers' and that it 'significantly improved the accuracy of the scores' on DeepReviewBench-2025, but no definition of consistency rate, no baseline comparisons, and no full results are given. Moreover, because the ACS is built from 2025-2026 papers, the figure caption itself states results on Bench-2024 were 'not satisfactory,' which is a stated limitation that the main text does not discuss. For a benchmark whose contribution is the ACS rubric, these omitted experimental details prevent the reader from assessing the rubric's quality independently of the 0.8141 correlation.","section":"Section 3.1 and Figure 1"}],"minor_comments":[{"comment":"The title contains a typo: 'Researchs' should be 'Research' or 'Research Processes'.","section":"Title"},{"comment":"There are minor language issues, including 'AI uto-Research' in the introduction and 'Ph.D. Candidates rankings' in Section 4.1, which should read 'Ph.D. candidates' rankings.'","section":"Abstract and Introduction"},{"comment":"Table 1 lists AstaBench twice with different properties (one with R=✓, C=✓, A=✓, another with R=✓, A=✓); the duplication appears accidental and should be resolved.","section":"Table 1"},{"comment":"The text says 'All AI scoring is done using GPT-5.2' for Basic Analysis, while Section 4 says all frameworks use Kimi-K2.6; clarify the judge model and whether the judge has access to the ACS library.","section":"Section 3.2"},{"comment":"The paper describes the three stages as 'mutually independent' but Section 3.2 says 'ground-truth information from preceding stages is fed as known conditions'; this tension should be reconciled, since feeding ground truth from prior stages contradicts strict independence.","section":"Section 1 and Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the corpus overlap between the 7,000-paper ACS source and the 200 ICLR 2026 gold references. The authors should be asked to disclose exactly whether any of the 200 gold papers or their review threads appear in the ACS source corpus; if they cannot confirm exclusion, a revised version should include a decontamination analysis or a re-construction of ACS. The stage correlation of 0.66 with n=10 is also a concern for the editor: the validation evidence is thinner than the abstract suggests. The paper is within scope for the journal, but the construct-validity claim must be substantiated before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is the short version. ARAC-Bench is a genuine attempt to evaluate auto-research systems on process rather than final output, and the stage-wise rubric design is a real step forward. But the headline validation correlation of 0.8141 is not yet supported: the ACS rubric library is built from the same 2025-2026 NeurIPS/ICLR/ICML pool that includes the 200 ICLR 2026 gold-reference papers, and the paper never states that the gold papers were excluded. Until that is resolved, the benchmark may be grading against a rubric extracted from the very papers it evaluates.\n\nWhat is actually new and useful: distilling reviewer expertise into structured, stage-calibrated Academic Cognition Skills; the three-stage Proposal/Experiment/Synthesis protocol with controlled base model and hard temporal cutoff; and the 2,869-module standard library for code scoring. The leaderboard of 11 frameworks is useful data, especially the finding that the Experiment stage is the main bottleneck and the best system only reaches 67.9/100. The ablations on Inspiration and on external skill packages are suggestive and give practitioners concrete signals. If the decontamination issue is fixed, this would be a usable diagnostic tool and potentially a reward signal.\n\nThe soft spots, in order of severity. First and load-bearing: Section 3.1 says ACS is mined from 7,000 papers accepted at NeurIPS/ICLR/ICML in 2025-2026, including rebuttals and review discussions. Section 3.2 says the gold references are 200 accepted ICLR 2026 papers. The temporal isolation only stops evaluated frameworks from searching post-mid-2025; it does nothing to keep gold papers out of the ACS construction. Since Proposal Details scoring and Methodological Deep Analysis scoring are ACS-driven, a gold paper's own review thread could have contributed the skills used to score it. This is a dataset-design flaw, not a statistical nitpick. The authors need to state and verify the exclusion, and ideally release the list of gold papers and the ACS entries used for each.\n\nSecond, the human validation is thin: ten PhD raters, no error bars, no sensitivity analysis on the 40/35/25 stage weights, and one stage correlation of 0.66. That is a secondary concern if the decontamination problem is fixed, but it should be addressed with more raters and variance reporting. Minor issues: the abstract's \"first benchmark\" framing is overstated because several cited benchmarks already evaluate parts of the research process, and the claimed 63% ACS consistency on unseen papers is reported without detail.\n\nWho this is for: people building or evaluating autonomous research systems, and anyone designing process-level benchmarks. It deserves a serious referee, but only with an explicit request to resolve the corpus-overlap question and to release the decontamination evidence. I would not cite the 0.8141 number in my own work until that is on the table; I would cite the framework and the leaderboard. Send it to review, conditional.","headline":"ARAC-Bench is a serious attempt at process-level evaluation of auto-research systems, but its headline 0.8141 validation number is not yet trustworthy because the rubric source and gold references may overlap.","tokens_in":18212,"tokens_out":2185,"would_cite":false,"duration_ms":24206,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ARAC-Bench shifts evaluation of auto-research from matching final answers to reproducing human research processes, and reports that current systems top out at 67.9/100 while correlating with PhD expert rankings at 0.8141.","keywords":["Auto-Research","Research agent evaluation","Process-level benchmark","Academic Cognition Skills","Scientific discovery automation","Reviewer skill distillation","AI alignment with expert judgment"],"falsifier":"Check whether any of the 200 gold-reference papers used in Section 3.2 appear in the 7,000-paper corpus from which the ACS rubric was mined in Section 3.1. If overlap exists, re-score the frameworks on a held-out set of gold references that were never touched by the rubric construction; a large drop in the 0.8141 correlation would show the benchmark rewards reproducing its own review criteria rather than human research process.","tokens_in":17083,"feed_emoji":"🧪","tokens_out":8115,"duration_ms":75615,"temperature":0.7,"pith_summary":"Auto-research systems—AI pipelines that carry out studies end to end—are usually scored by whether their final output works, not by whether their process resembles how a human researcher thinks. This paper introduces ARAC-Bench, a benchmark that grades them on reproducing high-quality human research processes, stage by stage: Proposal, Experiment, and Synthesis. Its core move is to convert the tacit judgment of expert reviewers into a structured skill rubric, then score each stage against that rubric. The paper reports that the best of 11 current frameworks reaches only 67.9 out of 100, and that its scores correlate with PhD-candidate rankings at 0.8141. If the correlation holds, ARAC-Bench offers a scalable way both to diagnose where research agents go wrong and to reward faithful methodology in training.","feed_headline":"Best auto-research framework scores only 67.9/100","feed_subtitle":"Process-level evaluation agrees with PhD experts (r = 0.8141) and shows even the top agent is far from human methodology.","key_machinery":"The central machinery is the Academic Cognition Skills (ACS) library: a stage-aware set of quantifiable rubrics distilled from how reviewers and authors argue in rebuttals, organized into five themes and 121 sub-themes, with each rubric expressed as a checklist of core review checkpoints. ARAC-Bench uses ACS to score the Proposal and Synthesis stages, while the Experiment stage is scored against a standard module library of 2,869 functional units. The three-stage diagnostic protocol (Proposal 40, Experiment 35, Synthesis 25) enforces modular isolation, feeds ground truth from earlier stages into later ones, and truncates literature search to mid-2025 so the evaluated agents cannot retrieve the gold-reference papers. The result is a scoring system in which every lost point can be traced to a missing skill, a missing module, or a broken logical link.","core_discovery":"The paper claims that ARAC-Bench is the first benchmark for Auto-Research that evaluates the reproduction of high-quality human research processes rather than matching final answers, and that this process-level alignment can be made both quantitative and faithful. It builds the Academic Cognition Skills (ACS) library from reviewer discussions and rebuttals across a large corpus of top-conference papers, decomposes research into three mutually independent stages, and scores each stage under controlled conditions. On 200 accepted top-conference papers used as gold references, the best of 11 evaluated frameworks scores 67.9/100, while validation against rankings by PhD candidates yields a mean correlation of 0.8141. The authors conclude that ARAC-Bench captures the dimensions researchers value and exposes a large gap in current autonomous research methodology, with the Experiment stage the main bottleneck.","pith_inferences":["Beyond the paper, the ACS rubric could be inverted into a dense reward signal for training research agents, an application the authors mention as future motivation but do not implement.","Beyond the paper, the stage-isolation protocol allows the same benchmark to be re-run on different base models, separating framework architecture from raw model capability; the paper fixes the base model rather than testing this.","Beyond the paper, a direct corpus-overlap audit is a testable guardrail: if any gold-reference paper appears in the ACS mining corpus, the reported correlation should be recomputed on disjoint data.","Beyond the paper, the paper's own note that its 2025-2026 rubric data performs poorly on 2024 benchmark material implies ARAC-Bench will require periodic skill-library refresh to stay valid."],"forward_implications":["ARAC-Bench turns research quality into per-stage diagnostics: a low Experiment score is attributable to missing modules or weak parameter priors, not to an aggregate failure.","Because the best available framework scores only 67.9/100, the benchmark implies current auto-research lags human methodology by more than 30 points, with Experiment replication the weakest stage.","Strong engineering ability alone does not produce scientific reasoning depth: a general coding tool with research skills scored near the bottom despite top coding-module scores.","Removing the Inspiration signal lowers proposal scores by roughly 20 percent on average, and more so for the strongest frameworks, implying that a concrete starting clue is what triggers structured reasoning.","Off-the-shelf academic skill packages improve literature and synthesis tasks but not core experimental reasoning, pointing to where the next generation of auto-research systems must improve."],"supporting_citations":[{"why":"Supplies a prior outcome-based benchmark for scientific agents that ARAC contrasts with process-level evaluation.","marker":"[2]"},{"why":"Provides an inspiration-based task decomposition benchmark whose outcome metrics motivates ARAC's inspiration ablation study.","marker":"[13]"},{"why":"Offers a replication-focused evaluation baseline that ARAC claims to go beyond by grading process rather than final reproduction.","marker":"[19]"},{"why":"The best-performing framework in ARAC's evaluation, establishing the 67.9/100 upper bound.","marker":"[12]"},{"why":"One of the 11 evaluated state-of-the-art frameworks, used in the main comparison table.","marker":"[32]"},{"why":"One of the evaluated frameworks and a representative adversarial multi-agent research pipeline.","marker":"[33]"}],"fun_headline_variants":["Best auto-research agent musters 67.9/100 on ARAC-Bench","New benchmark: auto-research tops out at 67.9/100 in mimicking human methods","ARAC-Bench: auto-research process alignment maxes at 67.9, validation r=0.81","ARAC-Bench: Top auto-research still 32 points short of human-level process","Benchmark confirms auto-research lags: best score 67.9, expert correlation 0.81"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"ARAC-Bench's scores are only meaningful if the 7,000-paper source set used to build its skill rubric does not include the 200 accepted papers used as gold references; otherwise the rubric would encode the answers it later grades.","fun_headline_variants_meta":{"raw":{"variants":["Best auto-research agent musters 67.9/100 on ARAC-Bench","New benchmark: auto-research tops out at 67.9/100 in mimicking human methods","ARAC-Bench: auto-research process alignment maxes at 67.9, validation r=0.81","ARAC-Bench: Top auto-research still 32 points short of human-level process","Benchmark confirms auto-research lags: best score 67.9, expert correlation 0.81"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001531,"raw_usage":{"total_tokens":6121,"prompt_tokens":928,"completion_tokens":5193,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":5076}},"tokens_in":544,"tokens_out":5193,"duration_ms":35131,"temperature":1.0,"reasoning_tokens":5076,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:17:05.762670+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether any of the 200 gold-reference papers used in Section 3.2 appear in the 7,000-paper corpus from which the ACS rubric was mined in Section 3.1. If overlap exists, re-score the frameworks on a held-out set of gold references that were never touched by the rubric construction; a large drop in the 0.8141 correlation would show the benchmark rewards reproducing its own review criteria rather than human research process.","supporting_citations":[],"review_version":1}