Pith. sign in

REVIEW 4 major objections 5 minor 46 references

ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read ARAC-Bench shifts evaluation of auto-research from matching final answers to reproducing human research processes, and reports that current systems top out at 67.9/100 while correlating with PhD expert rankings at 0.8141.

desk verdict ARAC-Bench is a serious attempt at process-level evaluation of auto-research systems, but its headline 0.8141 validation number is not yet trustworthy because the rubric source and gold references may overlap. read the letter →

arxiv 2608.12788 v1 pith:O6W7NSRR submitted 2026-08-13 cs.AI

classification cs.AI
keywords Auto-ResearchResearchagentevaluationProcess-levelbenchmarkAcademicCognitionSkillsScientificdiscoveryautomationReviewerskilldistillationAIalignmentwithexpertjudgment
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Auto-research systems—AI pipelines that carry out studies end to end—are usually scored by whether their final output works, not by whether their process resembles how a human researcher thinks. This paper introduces ARAC-Bench, a benchmark that grades them on reproducing high-quality human research processes, stage by stage: Proposal, Experiment, and Synthesis. Its core move is to convert the tacit judgment of expert reviewers into a structured skill rubric, then score each stage against that rubric. The paper reports that the best of 11 current frameworks reaches only 67.9 out of 100, and that its scores correlate with PhD-candidate rankings at 0.8141. If the correlation holds, ARAC-Bench offers a scalable way both to diagnose where research agents go wrong and to reward faithful methodology in training.

What carries the argument

The central machinery is the Academic Cognition Skills (ACS) library: a stage-aware set of quantifiable rubrics distilled from how reviewers and authors argue in rebuttals, organized into five themes and 121 sub-themes, with each rubric expressed as a checklist of core review checkpoints. ARAC-Bench uses ACS to score the Proposal and Synthesis stages, while the Experiment stage is scored against a standard module library of 2,869 functional units. The three-stage diagnostic protocol (Proposal 40, Experiment 35, Synthesis 25) enforces modular isolation, feeds ground truth from earlier stages into later ones, and truncates literature search to mid-2025 so the evaluated agents cannot retrieve the gold-reference papers. The result is a scoring system in which every lost point can be traced to a missing skill, a missing module, or a broken logical link.

What would settle it

Check whether any of the 200 gold-reference papers used in Section 3.2 appear in the 7,000-paper corpus from which the ACS rubric was mined in Section 3.1. If overlap exists, re-score the frameworks on a held-out set of gold references that were never touched by the rubric construction; a large drop in the 0.8141 correlation would show the benchmark rewards reproducing its own review criteria rather than human research process.

Watch

Extended reading notes

Core claim

The paper claims that ARAC-Bench is the first benchmark for Auto-Research that evaluates the reproduction of high-quality human research processes rather than matching final answers, and that this process-level alignment can be made both quantitative and faithful. It builds the Academic Cognition Skills (ACS) library from reviewer discussions and rebuttals across a large corpus of top-conference papers, decomposes research into three mutually independent stages, and scores each stage under controlled conditions. On 200 accepted top-conference papers used as gold references, the best of 11 evaluated frameworks scores 67.9/100, while validation against rankings by PhD candidates yields a mean correlation of 0.8141. The authors conclude that ARAC-Bench captures the dimensions researchers value and exposes a large gap in current autonomous research methodology, with the Experiment stage the main bottleneck.

Load-bearing premise

ARAC-Bench's scores are only meaningful if the 7,000-paper source set used to build its skill rubric does not include the 200 accepted papers used as gold references; otherwise the rubric would encode the answers it later grades.

Editorial extensions

If this is right

  • ARAC-Bench turns research quality into per-stage diagnostics: a low Experiment score is attributable to missing modules or weak parameter priors, not to an aggregate failure.
  • Because the best available framework scores only 67.9/100, the benchmark implies current auto-research lags human methodology by more than 30 points, with Experiment replication the weakest stage.
  • Strong engineering ability alone does not produce scientific reasoning depth: a general coding tool with research skills scored near the bottom despite top coding-module scores.
  • Removing the Inspiration signal lowers proposal scores by roughly 20 percent on average, and more so for the strongest frameworks, implying that a concrete starting clue is what triggers structured reasoning.
  • Off-the-shelf academic skill packages improve literature and synthesis tasks but not core experimental reasoning, pointing to where the next generation of auto-research systems must improve.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the ACS rubric could be inverted into a dense reward signal for training research agents, an application the authors mention as future motivation but do not implement.
  • Beyond the paper, the stage-isolation protocol allows the same benchmark to be re-run on different base models, separating framework architecture from raw model capability; the paper fixes the base model rather than testing this.
  • Beyond the paper, a direct corpus-overlap audit is a testable guardrail: if any gold-reference paper appears in the ACS mining corpus, the reported correlation should be recomputed on disjoint data.
  • Beyond the paper, the paper's own note that its 2025-2026 rubric data performs poorly on 2024 benchmark material implies ARAC-Bench will require periodic skill-library refresh to stay valid.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes ARAC-Bench, a benchmark for evaluating autonomous research (Auto-Research) frameworks by measuring how closely their research processes align with human expert behavior. The benchmark decomposes research into three stages—Proposal, Experiment, and Synthesis—and scores each stage using rubrics derived from an Academic Cognition Skills (ACS) library. The ACS library is constructed from 7,000 papers accepted at NeurIPS, ICLR, and ICML in 2025-2026, including rebuttals and review discussions. ARAC-Bench's gold references are 200 accepted ICLR 2026 papers. The authors evaluate 11 frameworks, all mounted on Kimi-K2.6, and report that the best framework achieves only 67.9/100. They validate ARAC-Bench against rankings from ten PhD candidates and report an average Pearson correlation of 0.8141 across the three stages (0.8788, 0.6606, 0.9030). They also present ablations showing the importance of inspiration signals and the limited benefit of integrating existing academic skill packages.

Significance. If the validity concerns are resolved, ARAC-Bench addresses a genuine gap: current Auto-Research evaluation emphasizes outcome metrics such as code pass rates or paper completeness, whereas ARAC-Bench attempts a process-level, stage-resolved diagnosis of methodological alignment. The paper's strengths include a clearly described three-stage decomposition, a controlled-variable evaluation design with a temporal cutoff, a public dataset release, and practical ablations on inspiration signals and skill packages. The design direction—validating a rubric-based benchmark against human expert rankings—is sensible and the reported low scores for state-of-the-art frameworks are interesting. However, the central claim that ARAC-Bench is a 'reliable proxy for expert judgment' is not yet supported because of a potential overlap between the ACS source corpus and the gold-reference papers, and because the human validation rests on only ten raters with no reported uncertainty. The significance of the contribution will depend on whether the authors can establish that the rubric is not circularly derived from the evaluation targets.

major comments (4)
  1. [Section 3.1 and Section 3.2] The ACS library is constructed from 7,000 papers accepted at NeurIPS, ICLR, and ICML in 2025-2026, including rebuttals and review discussions, while the 200 gold-reference papers are all accepted ICLR 2026 papers. The manuscript nowhere states that the 200 gold papers were excluded from the 7,000-paper source corpus, and the temporal isolation described in Section 3.2 (search cutoff at mid-2025) does not decontaminate the rubric source. If the gold papers' own review threads contributed skills to the ACS library, then Proposal Details Scoring (25/100) and Methodological Deep Analysis (15/100) are partially grading against an answer key derived from the very papers being evaluated. The reported average correlation of 0.8141 with PhD rankings therefore cannot, as reported, establish that ARAC-Bench is an independent proxy for expert judgment; it may instead reflect that the rubric and the raters reward the same content dimensions. This is a dataset-design issue that must be resolved, for example by explicitly confirming an exclusion protocol or by re-constructing ACS without the gold papers and re-running the validation.
  2. [Section 4.1, Table 4] The validation of the central reliability claim rests on ten PhD raters ranking ten frameworks. The paper reports Pearson correlations of 0.8788, 0.6606, and 0.9030 for the three stages and calls their average 0.8141, but it provides no confidence intervals, significance tests, or inter-rater agreement measures. With n=10, the 95% confidence interval for the Experiment-stage correlation alone would be very wide and may include values that are not 'moderately strong.' The paper also does not report the correlation between the overall ARAC total score and the human comprehensive ranking, which is the quantity most relevant to the claim that ARAC-Bench is a reliable proxy for expert judgment. Without these statistics, the validation evidence is too thin to support the abstract's 'strong average correlation' claim.
  3. [Section 3.2, stage weights] The total score is a weighted sum with weights 40/35/25 for the three stages, and the Proposal stage further assigns 5/25/10 to its sub-modules, with the 25-point Proposal Details score depending on a Top-5 ACS skill retrieval and three-level anchor scores (0/1.5/3). These weights and anchors are introduced without any sensitivity analysis or justification, yet they determine the framework ranking in Table 3 and the 'best alignment score of only 67.9' headline result. A small perturbation of the stage weights could reorder frameworks in the 50-60 range and change the reported gap. The authors should provide a sensitivity analysis or justify the weights from data (e.g., fitted to the human rankings) rather than asserting them.
  4. [Section 3.1 and Figure 1] The paper claims that ACS achieves 'the SOTA 63% consistency rate' on '200 paper queries of unseen ICLR papers' and that it 'significantly improved the accuracy of the scores' on DeepReviewBench-2025, but no definition of consistency rate, no baseline comparisons, and no full results are given. Moreover, because the ACS is built from 2025-2026 papers, the figure caption itself states results on Bench-2024 were 'not satisfactory,' which is a stated limitation that the main text does not discuss. For a benchmark whose contribution is the ACS rubric, these omitted experimental details prevent the reader from assessing the rubric's quality independently of the 0.8141 correlation.
minor comments (5)
  1. [Title] The title contains a typo: 'Researchs' should be 'Research' or 'Research Processes'.
  2. [Abstract and Introduction] There are minor language issues, including 'AI uto-Research' in the introduction and 'Ph.D. Candidates rankings' in Section 4.1, which should read 'Ph.D. candidates' rankings.'
  3. [Table 1] Table 1 lists AstaBench twice with different properties (one with R=✓, C=✓, A=✓, another with R=✓, A=✓); the duplication appears accidental and should be resolved.
  4. [Section 3.2] The text says 'All AI scoring is done using GPT-5.2' for Basic Analysis, while Section 4 says all frameworks use Kimi-K2.6; clarify the judge model and whether the judge has access to the ACS library.
  5. [Section 1 and Section 3.2] The paper describes the three stages as 'mutually independent' but Section 3.2 says 'ground-truth information from preceding stages is fed as known conditions'; this tension should be reconciled, since feeding ground truth from prior stages contradicts strict independence.

Circularity Check

2 steps flagged · score 6.0 of 10

ACS rubric corpus may include the same ICLR 2026 gold papers, so Proposal and Synthesis scores can grade against criteria distilled from the very papers being scored; the 0.8141 PhD correlation is external but does not decontaminate the rubric.

  1. self definitional [Section 3.1 'Academic Cognition Skills' and Section 3.2 'Researcher-Mimicking: Capability Diagnosis of Proposal, Experiment, and Synthesis' (Proposal Details Scoring 25)]
    "From 7,000 top AI conference papers accepted at NeurIPS, ICLR, and ICML in 2025 and 2026, we first extracted the core scientific question each paper aimed to address; subsequently, we targeted the corresponding rebuttals and review discussions to extract the cognitive strategies employed by authors and reviewers in defending, clarifying, or revising those questions."

    Section 3.2 builds ARAC-Bench on '200 accepted papers from ICLR 2026', a conference explicitly inside the 7,000-paper ACS source set ('NeurIPS, ICLR, and ICML in 2025 and 2026'). No exclusion of the 200 gold papers from the ACS corpus is stated. The Proposal Details score (25/100) is 'entirely driven by ACS' and 'dynamically retrieves the Top-5 most relevant core cognition skills from the ACS library'; if a gold paper's own rebuttal contributed those skills, the rubric is defined from the target paper's review thread, so a high proposal score measures recapitulation of the answer key rather than an independent human standard.

  2. self definitional [Section 3.2 'Synthesis Stage scoring 25' (Methodological Deep Analysis Scoring 15)]
    "Driven by attribution reasoning and criticalsynthesis skills from the ACS library, operationalized via a three-layer progressive verification protocol. Key evaluation criteria include: ... (3) Conclusion Extrapolation Boundary: whether the Limitations claims of the original paper are faithfully reproduced."

    This module grades frameworks on reproducing the original paper's mechanism explanation, ablation counterfactuals, and Limitations. Since ACS attribution/critical-synthesis skills are distilled from rebuttals and review discussions of the 2025-2026 NeurIPS/ICLR/ICML corpus, which may include these same ICLR 2026 gold papers, the score is defined against a criterion that can be generated from the target paper itself. The 'gold' reference thereby doubles as the rubric source, making the 'alignment' measure circular at the construction level.

full rationale

ARAC-Bench contains genuinely independent components: Related Work recall against the paper's cited references, Benchmark Selection via executable sandbox checks, and Code Implementation against a 2,869-module standard library. The 0.8141 correlation with PhD rankings is an external sanity check and does not by itself constitute circularity. However, the two ACS-driven modules (Proposal Details, 25/100, and Methodological Deep Analysis, 15/100) are constructed from the same conference-year corpus as the 200 gold references. The paper never states that the 200 gold papers were excluded from the 7,000-paper ACS mining set; its only stated isolation is the mid-2025 search cutoff for evaluated frameworks, which prevents model access but not ACS contamination. If a gold paper's review thread contributed ACS skills, the corresponding score reduces to checking against criteria extracted from that same paper. Thus the central claim that ARAC-Bench reliably measures human-like research process is partially circular, though the external PhD validation and the non-ACS modules keep it from being wholly forced. The appropriate score is 6, reflecting partial circularity in the construction of the central process-alignment measure.

Assumptions & free parameters 3 free parameters · 5 assumptions · 2 invented entities

The benchmark makes no mathematical derivation; it rests on design choices and domain assumptions. The most consequential choices are the 40/35/25 stage weights, the top-5 ACS skill retrieval, and the three-level anchor scoring, none of which are sensitivity-tested. The central validity assumption is that accepted conference papers and reviewer-rebuttal patterns define a transferable gold standard for human research process, and that the ACS corpus does not leak the 200 ICLR 2026 evaluation papers into the rubric construction.

free parameters (3)
  • Stage weights = 40/35/25
    Chosen by authors to define the total alignment score; the main result of 67.9 is a weighted sum, and no sensitivity analysis is reported for these weights.
  • Top-5 ACS skill selection = 5
    Proposal and Synthesis scoring retrieve the top five matching ACS skills; the number and the matching method are hand-chosen and affect which rubric dimensions are graded.
  • Anchor scoring levels = 3, 1.5, 0
    Three-level anchor scoring for each ACS skill point in the Proposal stage; the choice of half-point midlevel affects Idea scores and all downstream rankings.
assumptions (5)
  • domain assumption Accepted ICLR 2026 papers are valid gold standards for high-quality human research process.
    All three stages score relative to these papers: related-work recall, hyperparameter settings, golden skeleton structure, and causal explanations. If accepted papers are not representative of ideal research process, the scores do not measure what is claimed.
  • domain assumption The ACS corpus built from 7,000 accepted papers does not overlap with the 200 ICLR 2026 gold-reference papers, or any overlap does not inflate scores.
    No exclusion of the 200 gold papers from the ACS construction corpus is stated in Section 3.1; the validity of the alignment scores depends on this being true or harmless.
  • domain assumption Evaluating only final stabilized outputs captures research process quality.
    The Figure 2 caption says 'In every stages, we only evaluate the final output, even if the framework refines,' yet the paper claims process-level diagnosis; internal search, failure, and iteration behavior are not directly observed.
  • domain assumption Related-work quality is measured by recall against the original paper's cited reference set.
    Citations are not an exhaustive ground truth and no precision or relevance judgment is used, so a framework retrieving extra relevant works is penalized and citation quality is only approximated.
  • domain assumption The 2,869-unit standard module library is an objective yardstick for coding quality.
    The module library is constructed from Proposal designs; if its decomposition is arbitrary or built around gold implementations, coding scores reward matching that decomposition rather than general engineering quality.
invented entities (2)
  • Academic Cognition Skills (ACS) library independent evidence
    purpose: A constructed rubric knowledge base that distills reviewer expertise into stage-calibrated scoring checkpoints used for Proposal and Synthesis evaluation.
    The paper reports 63% consistency on unseen ICLR queries and improved accuracy on DeepReviewBench-2025 in Figure 1, which gives an external falsifiable handle, though no error bars or comparison protocol are given.
  • Standard module library of 2,869 functional units
    purpose: An objective scoring yardstick for code implementation in the Experiment stage; missing modules are scored as zero.
    No external validation or release of the module library is provided, and its construction from Proposal designs makes it an internal artifact of the benchmark.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs." pith.science (2026). https://pith.science/paper/O6W7NSRR

@misc{pith2026260812788,
  author       = {Pith},
  title        = {Pith review of: ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O6W7NSRR}},
  note         = {Machine review of arXiv:2608.12788}
}
read the original abstract

The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.

Figures

Figures reproduced from arXiv: 2608.12788 by the authors.

Figure 1
Figure 1. To validate its reliability, we evaluated ACS on 200 paper queries of unseen ICLR papers above, [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. To ensure evaluation fairness and ecological validity, we enforce strict temporal isolation: the [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 2
Figure 2. ARAC-Bench decomposes the research workflow into three independent stages, while ground [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 24 canonical work pages

  1. [1]

    Evaluating sakana’s ai scientist: Bold claims, mixed results, and a promising future? InACM SIGIR Forum, volume 59, pages 1–20

    Joeran Beel, Min-Yen Kan, and Moritz Baumgart. Evaluating sakana’s ai scientist: Bold claims, mixed results, and a promising future? InACM SIGIR Forum, volume 59, pages 1–20. ACM New York, NY , USA, 2025

  2. [2]

    Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, et al. Scienceagentbench: Toward rigorous assessment of language agents for data-driven scientific discovery. InInternational Conference on Learning Representations, volume 2025, pages 96934–96990, 2025

  3. [3]

    Re- viewagents: Bridging the gap between human and ai-generated paper reviews.arXiv preprint arXiv:2503.08506, 2025

    Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao, Ting Liu, and Yuzhuo Fu. Re- viewagents: Bridging the gap between human and ai-generated paper reviews.arXiv preprint arXiv:2503.08506, 2025

  4. [4]

    Towards an ai co-scientist

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, et al. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2025

  5. [5]

    Accelerating scien- tific discovery with co-scientist.Nature, pages 1–3, 2026

    Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, et al. Accelerating scien- tific discovery with co-scientist.Nature, pages 1–3, 2026

  6. [6]

    Scholarpeer: A context-aware multi-agent framework for automated peer review.arXiv preprint arXiv:2601.22638, 2026

    Palash Goyal, Mihir Parmar, Yiwen Song, Hamid Palangi, Tomas Pfister, and Jinsung Yoon. Scholarpeer: A context-aware multi-agent framework for automated peer review.arXiv preprint arXiv:2601.22638, 2026

  7. [7]

    Openreviewer: A specialized large language model for gen- erating critical scientific paper reviews

    Maximilian Idahl and Zahra Ahmadi. Openreviewer: A specialized large language model for gen- erating critical scientific paper reviews. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (System Demonstrations), pages 550–562, 2025

  8. [8]

    Toward generalist autonomous research via hypothesis-tree refinement, 2026

    Jiajie Jin, Yuyang Hu, Kai Qiu, Qi Dai, Chong Luo, Guanting Dong, Xiaoxi Li, Tong Zhao, Xiao- long Ma, Gongrui Zhang, Zhirong Wu, Bei Liu, Zhengyuan Yang, Linjie Li, Lijuan Wang, Hongjin Qian, Yutao Zhu, and Zhicheng Dou. Toward generalist autonomous research via hypothesis-tree refinement, 2026. URLhttps://arxiv.org/abs/2606.11926

Show all 46 references
  1. [9]

    Scientific agent skills: A comprehensive collection of scientific tools for ai agents,

    K-Dense Inc. Scientific agent skills: A comprehensive collection of scientific tools for ai agents,

  2. [10]

    Ai for auto-research: Roadmap & user guide.arXiv preprint arXiv:2605.18661, 2026

    Lingdong Kong, Xian Sun, Wei Chow, Linfeng Li, Kevin Qinghong Lin, Xuan Billy Zhang, Song Wang, Rong Li, Qing Wu, Wei Gao, et al. Ai for auto-research: Roadmap & user guide.arXiv preprint arXiv:2605.18661, 2026

  3. [11]

    Autosota: An end-to-end automated research system for state-of-the-art ai model discovery.arXiv preprint arXiv:2604.05550, 2026

    Yu Li, Chenyang Shao, Xinyang Liu, Ruotong Zhao, Peijie Liu, Hongyuan Su, Zhibin Chen, Qin- glong Yang, Anjie Xu, Yi Fang, et al. Autosota: An end-to-end automated research system for state-of-the-art ai model discovery.arXiv preprint arXiv:2604.05550, 2026

  4. [12]

    Autoresearchclaw: Self-reinforcing autonomous research with human- ai collaboration.arXiv preprint arXiv:2605.20025, 2026

    Jiaqi Liu, Shi Qiu, Mairui Li, Bingzhou Li, Haonian Ji, Siwei Han, Xinyu Ye, Peng Xia, Zihan Dong, Congyu Zhang, et al. Autoresearchclaw: Self-reinforcing autonomous research with human- ai collaboration.arXiv preprint arXiv:2605.20025, 2026. 12

  5. [13]

    Researchbench: Benchmarking llms in scientific discovery via inspiration-based task decomposition.arXiv preprint arXiv:2503.21248, 2025

    Yujie Liu, Zonglin Yang, Tong Xie, Jinjie Ni, Ben Gao, Yuqiang Li, Shixiang Tang, Wanli Ouyang, Erik Cambria, and Dongzhan Zhou. Researchbench: Benchmarking llms in scientific discovery via inspiration-based task decomposition.arXiv preprint arXiv:2503.21248, 2025

  6. [14]

    Evoscientist: Towards multi-agent evolving ai scientists for end-to-end scientific discovery.arXiv preprint arXiv:2603.08127, 2026

    Yougang Lyu, Xi Zhang, Xinhao Yi, Yuyue Zhao, Shuyu Guo, Wenxiang Hu, Jan Piotrowski, Jakub Kaliski, Jacopo Urbani, Zaiqiao Meng, et al. Evoscientist: Towards multi-agent evolving ai scientists for end-to-end scientific discovery.arXiv preprint arXiv:2603.08127, 2026

  7. [15]

    Ai research skills library, 2025

    Orchestra Research. Ai research skills library, 2025. URLhttps://github.com/ orchestra-research/AI-research-SKILLs. Open-source skills library enabling AI agents to autonomously conduct AI research

  8. [16]

    Autosci: A memory-centric agentic system for the full scientific research lifecycle.arXiv preprint arXiv:2605.31468, 2026

    Weitong Qian, Beicheng Xu, Zhongao Xie, Bowen Fan, Guozheng Tang, Jiale Chen, Xinzhe Wu, Mingtian Yang, Chenyang Di, Jiajun Li, et al. Autosci: A memory-centric agentic system for the full scientific research lifecycle.arXiv preprint arXiv:2605.31468, 2026

  9. [17]

    Agent laboratory: Using llm agents as research assistants.Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977– 6043, 2025

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using llm agents as research assistants.Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977– 6043, 2025

  10. [18]

    Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, and Lichao Sun. Dr. claw: An ai research workspace from idea to paper, 2026. URLhttps: //github.com/OpenLAIR/dr-claw

  11. [19]

    Paperbench: Evaluating ai’s ability to replicate ai research.arXiv preprint arXiv:2504.01848, 2025

    Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, et al. Paperbench: Evaluating ai’s ability to replicate ai research.arXiv preprint arXiv:2504.01848, 2025

  12. [20]

    Ai-researcher: Autonomous scientific innovation.Advances in Neural Information Processing Systems, 38:9481–9520, 2026

    Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. Ai-researcher: Autonomous scientific innovation.Advances in Neural Information Processing Systems, 38:9481–9520, 2026

  13. [21]

    Autoresearch ai: Towards ai-powered research automation for scientific discovery.arXiv preprint arXiv:2605.23204, 2026

    Guiyao Tie, Jiawen Shi, Dingjie Song, Yixiao Huang, Ziji Sheng, Xueyang Zhou, Daizong Liu, Pan Zhou, Yongchao Chen, Ran Xu, et al. Autoresearch ai: Towards ai-powered research automation for scientific discovery.arXiv preprint arXiv:2605.23204, 2026

  14. [22]

    Autonomous ‘self-driving’laboratories: a review of tech- nology and policy implications.Royal Society Open Science, 12(7):250646, 2025

    Alexander V Tobias and Adam Wahab. Autonomous ‘self-driving’laboratories: a review of tech- nology and policy implications.Royal Society Open Science, 12(7):250646, 2025

  15. [23]

    Deep research arena: The first exam of llms’ research abilities via seminar- grounded tasks

    Haiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu, Jiaxuan Lu, Di Yu, Jianbao Cao, Ben Gao, Jiaqing Xie, Aoran Wang, et al. Deep research arena: The first exam of llms’ research abilities via seminar- grounded tasks. InProceedings of the AAAI Conference on Artificial Intelligence, v...

  16. [24]

    Fire-bench: Evaluating agents on the rediscovery of scientific insights.arXiv preprint arXiv:2602.02905, 2026

    Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, et al. Fire-bench: Evaluating agents on the rediscovery of scientific insights.arXiv preprint arXiv:2602.02905, 2026

  17. [25]

    From ai for science to agentic science: A survey on autonomous scientific discovery.arXiv preprint arXiv:2508.14111, 2025

    Jiaqi Wei, Yuejin Yang, Xiang Zhang, Yuhan Chen, Xiang Zhuang, Zhangyang Gao, Dongzhan Zhou, Guangshuai Wang, Zhiqiang Gao, Juntai Cao, et al. From ai for science to agentic science: A survey on autonomous scientific discovery.arXiv preprint arXiv:2508.14111, 2025. 13

  18. [26]

    Cycleresearcher: Improving automated research via automated review

    Yixuan Weng, Minjun Zhu, Guangsheng Bao, Hongbo Zhang, Jindong Wang, Yue Zhang, and Linyi Yang. Cycleresearcher: Improving automated research via automated review. InInternational Conference on Learning Representations, volume 2025, pages 3669–3709, 2025

  19. [27]

    Deepscien- tist: Advancing frontier-pushing scientific findings progressively.arXiv preprint arXiv:2509.26603, 2025

    Yixuan Weng, Minjun Zhu, Qiujie Xie, Qiyao Sun, Zhen Lin, Sifan Liu, and Yue Zhang. Deepscien- tist: Advancing frontier-pushing scientific findings progressively.arXiv preprint arXiv:2509.26603, 2025

  20. [28]

    Deepreviewer 2.0: A traceable agentic system for auditable scientific peer review.arXiv preprint arXiv:2604.09590, 2026

    Yixuan Weng, Minjun Zhu, Qiujie Xie, Zhiyuan Ning, Shichen Li, Panzhong Lu, Zhen Lin, En- hao Gu, Qiyao Sun, and Yue Zhang. Deepreviewer 2.0: A traceable agentic system for auditable scientific peer review.arXiv preprint arXiv:2604.09590, 2026

  21. [29]

    Claw ai lab: An autonomous multi-agent research team.arXiv preprint arXiv:2605.22662, 2026

    Fan Wu, Cheng Chen, Zhenshan Tan, Taiyu Zhang, Xinzhen Xu, Yanyu Qian, Dingcheng Gao, Lanyun Zhu, Qi Zhu, Yi Tan, et al. Claw ai lab: An autonomous multi-agent research team.arXiv preprint arXiv:2605.22662, 2026

  22. [30]

    Autoresearchbench: Benchmarking ai agents on complex scientific literature discovery.arXiv preprint arXiv:2604.25256, 2026

    Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang, Jin-Ge Yao, Zheng Liu, Jingying Shao, Jianlyu Chen, Hongjin Qian, Xi Yang, et al. Autoresearchbench: Benchmarking ai agents on complex scientific literature discovery.arXiv preprint arXiv:2604.25256, 2026

  23. [31]

    Nanoresearch: Co-evolving skills, memory, and policy for personalized research automation.arXiv preprint arXiv:2605.10813, 2026

    Jinhang Xu, Qiyuan Zhu, Yujun Wu, Zirui Wang, Dongxu Zhang, Jianxin Tang, Marcia Tian, Yiling Duan, Siyuan Li, Jingxuan Wei, et al. Nanoresearch: Co-evolving skills, memory, and policy for personalized research automation.arXiv preprint arXiv:2605.10813, 2026

  24. [32]

    The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search.arXiv preprint arXiv:2504.08066, 2025

    Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob Foerster, Jeff Clune, and David Ha. The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search.arXiv preprint arXiv:2504.08066, 2025

  25. [33]

    Aris: Autonomous research via adversarial multi-agent collaboration.arXiv preprint arXiv:2605.03042, 2026

    Ruofeng Yang, Yongcan Li, and Shuai Li. Aris: Autonomous research via adversarial multi-agent collaboration.arXiv preprint arXiv:2605.03042, 2026

  26. [34]

    Autoreproduce: Automatic ai experiment reproduction with paper lineage.arXiv preprint arXiv:2505.20662, 2025

    Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi, Weilun Zhao, Shuo Wang, Duzhen Zhang, Xu Han, Zhiyuan Liu, and Maosong Sun. Autoreproduce: Automatic ai experiment reproduction with paper lineage.arXiv preprint arXiv:2505.20662, 2025

  27. [35]

    From automation to autonomy: A survey on large language models in scientific discovery

    Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. From automation to autonomy: A survey on large language models in scientific discovery. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, ...

  28. [36]

    Deepreview: Improving llm-based paper review with human-like deep thinking process

    Minjun Zhu, Yixuan Weng, Linyi Yang, and Yue Zhang. Deepreview: Improving llm-based paper review with human-like deep thinking process. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 29330–29355, 2025. 1...

  29. [38]

    Functional Correctness Mapping 10 • Do the core classes/functions in the code strictly correspond to the algorithmic pseudocode or mathematical formulas in the proposal? • Do the transformations of key tensor dimensions remain consistent with formula derivations? • Are there o...

  30. [39]

    Module Completeness & Decoupling 10 • Are Dataloader, Model, Trainer, Logger reasonably separated? • For any missing modules, are there clear interfaces or TODO comments indicating that the framework is aware of the omission but has reserved extension points? • Is there basic ...

  31. [40]

    • Dependency Check: Inspect ‘requirements.txt‘ or ‘environment.yml‘ to verify whether special dependency libraries mentioned in the proposal are included

    Engineering Robustness 10 • Are hyperparameters managed centrally via configuration files (Config), rather than scattered throughout the code? • Is the random seed fixed to ensure reproducibility? • Are there simple unit tests or assertions to verify the most basic I/O of core...

  32. [41]

    Introduction (3 Points) • Are the Research Gap and Contributions list proposed in the original paper fully covered by the test text? • If the test text supplements new contribution points not explicitly mentioned in the original paper but logically reasonable, it is regarded a...

  33. [42]

    Related Work (4 Points) • Are the academic schools of thought divided in the original paper reproduced by the test text? • Does the test text cite new SOTA works not covered by the original paper but more relevant in the past two years? If so, it is regarded as a positive devi...

  34. [43]

    Preliminaries (3 Points) • Are the mathematical notations and foundational formulas defined in the original paper (e.g., loss function definitions) fully preserved by the test text? • If the test text deletes complex formula derivations from the original paper, rendering subse...

  35. [44]

    Mechanism Explanation Depth (Relative Baseline, 5 Points) • How does the original paper explain the intrinsic mechanism by which method works? • Does the test text reproduce this explanation? If the test text employs more mathematical descriptions, it is regarded as a positive...

  36. [45]

    Counterfactual Reasoning Coverage (Relative Baseline, 5 Points) • What kind of why it degrades analysis does the original paper provide for ablation studies? • Does the test text cover the analysis ofallkey ablation items in the original paper? If the test text additionally ex...

  37. [46]

    Conclusion Extrapolation Boundary (Relative Baseline, 5 Points) • How does the original paper limit the scope of applicability of its conclusions (Limitations) in the Conclusion section? • Does the test text faithfully reproduce the original paper’s limitation claims? If the t...

  38. [2026]

    148 skills cov- ering databases, packages, integrations, and analysis tools

    URLhttps://github.com/K-Dense-AI/scientific-agent-skills. 148 skills cov- ering databases, packages, integrations, and analysis tools

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.