Pith. sign in

REVIEW 2 major objections 6 cited by

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement

T0 review · 2 major / 0 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Arbor maintains a persistent Hypothesis Tree that lets an AI coordinator accumulate and refine research insights across many iterations instead of restarting each time.

desk verdict The claimed gains look like they could come from running on GPT-5.5 while baselines used older models, so the HTR contribution is not yet isolated. read the letter →

arxiv 2606.11926 v1 pith:Z3SZXXLM submitted 2026-06-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords autonomousresearchhypothesistreerefinementAIagentsautomationcumulativelearningmodeloptimizationscientificdiscovery
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces Arbor as a framework that runs the scientific loop of exploration, experimentation, and abstraction autonomously over long horizons. It pairs a long-lived coordinator that manages global strategy with short-lived executors that test individual hypotheses, all organized through Hypothesis Tree Refinement. The tree stores hypotheses, artifacts, evidence, and distilled lessons so that verified improvements and reusable insights propagate forward rather than being lost. Evaluation on six concrete tasks in model training, harness engineering, and data synthesis shows Arbor reaching the best held-out result on every task and more than 2.5 times the average relative gain of Codex and Claude Code under identical interfaces and budgets. On MLE-Bench Lite the system reaches 86.36 percent Any Medal with GPT-5.5.

What carries the argument

Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights so the coordinator can propagate lessons and refine the search frontier across iterations.

What would settle it

Running the same six tasks with an otherwise identical Arbor variant that replaces the Hypothesis Tree with a flat chronological log and measuring whether the 2.5x relative gain disappears.

Watch

Extended reading notes

Core claim

Arbor is a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. Acr

Load-bearing premise

The six chosen tasks plus MLE-Bench Lite are representative of general autonomous research and the reported gains arise from the Hypothesis Tree Refinement mechanism rather than from differences in prompting, model access, or task-specific engineering.

Editorial extensions

If this is right

  • Research agents can improve an initial artifact through iterative experimentation without step-level human supervision.
  • Lessons from failed or successful hypotheses become reusable across later attempts rather than being discarded.
  • The same task interface and resource budget produce substantially higher held-out performance than prior agent baselines.
  • The framework scales to multiple domains including model training, harness engineering, and data synthesis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the tree structure continues to grow without becoming intractable, the approach could support research campaigns lasting hundreds of iterations.
  • The coordinator-executor split may allow future versions to run many executors in parallel while the coordinator maintains a single coherent research plan.
  • Distilled insights stored in the tree could be extracted and reused by other agents or even by human researchers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript introduces Arbor, a framework for autonomous research that integrates a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR) to maintain a persistent tree linking hypotheses, artifacts, evidence, and distilled insights across iterations. It evaluates the system in an Autonomous Optimization setting on six real research tasks (model training, harness engineering, data synthesis) plus MLE-Bench Lite, claiming that Arbor achieves the best held-out result on all six tasks, more than 2.5x the average relative held-out gain of Codex and Claude Code under identical task interface and resource budget, and 86.36% Any Medal on MLE-Bench Lite when using GPT-5.5.

Significance. If the reported gains prove robust and attributable to the HTR mechanism rather than model or implementation differences, the work would offer a concrete approach to cumulative long-horizon autonomous research, moving beyond isolated attempts to a process that propagates lessons across time; the persistent tree structure is a clear conceptual contribution.

major comments (2)
  1. [Abstract] Abstract: the central performance claims (best result on all six tasks, >2.5x average relative held-out gain, 86.36% Any Medal) are presented without any description of experimental controls, statistical significance testing, or confirmation that the baselines (Codex, Claude Code) used the identical base model GPT-5.5 rather than weaker models; this directly affects whether gains can be attributed to HTR versus model access.
  2. [Abstract] Abstract and evaluation description: the claim that results arise from the Hypothesis Tree Refinement mechanism (persistent linking of hypotheses/artifacts/evidence) rather than task-specific engineering or prompting differences is not supported by any ablation or controlled comparison that isolates the tree component while holding model, prompt template, and executor fixed.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback highlighting the need for greater clarity in the abstract regarding experimental controls and the attribution of gains to the HTR mechanism. We address each major comment below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central performance claims (best result on all six tasks, >2.5x average relative held-out gain, 86.36% Any Medal) are presented without any description of experimental controls, statistical significance testing, or confirmation that the baselines (Codex, Claude Code) used the identical base model GPT-5.5 rather than weaker models; this directly affects whether gains can be attributed to HTR versus model access.

    Authors: We agree that the abstract should explicitly describe the experimental controls to support attribution. The manuscript already states that comparisons used the same task interface and resource budget, with the MLE-Bench Lite result reported using GPT-5.5 for Arbor. We will revise the abstract to add a concise description of these controls and to note that statistical significance testing across multiple runs was not performed due to computational cost. This revision will make the shared setup clearer without altering the reported numbers. revision: yes

  2. Referee: [Abstract] Abstract and evaluation description: the claim that results arise from the Hypothesis Tree Refinement mechanism (persistent linking of hypotheses/artifacts/evidence) rather than task-specific engineering or prompting differences is not supported by any ablation or controlled comparison that isolates the tree component while holding model, prompt template, and executor fixed.

    Authors: The referee correctly observes that the manuscript contains no ablation that removes only the HTR component while holding the model, prompt templates, and executor implementation fixed. The reported comparisons evaluate the complete Arbor system against other agent frameworks under matched model and interface conditions, but do not isolate the persistent tree. We will revise the evaluation description to more explicitly articulate the intended contribution of HTR and to acknowledge this limitation. A dedicated ablation would require additional controlled experiments that are outside the scope of the current results. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; empirical framework evaluation with no derivations or self-referential reductions.

full rationale

The paper introduces the Arbor framework and Hypothesis Tree Refinement for autonomous research, then reports empirical results on six tasks and MLE-Bench Lite. No equations, fitted parameters, or mathematical derivations are present. Performance claims rest on held-out comparisons under a stated task interface and budget, without any step that reduces by construction to inputs, self-citations, or ansatzes. The central attribution to HTR is an empirical claim open to experimental scrutiny rather than a definitional or fitted tautology. This is a standard non-circular empirical paper.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract-only review supplies no information on free parameters, axioms, or invented entities; the ledger is therefore empty.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward Generalist Autonomous Research via Hypothesis-Tree Refinement." pith.science (2026). https://pith.science/paper/Z3SZXXLM

@misc{pith2026260611926,
  author       = {Pith},
  title        = {Pith review of: Toward Generalist Autonomous Research via Hypothesis-Tree Refinement},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z3SZXXLM}},
  note         = {Machine review of arXiv:2606.11926}
}
read the original abstract

Scientific progress depends on a repeated loop of exploration, experimentation, and abstraction. Researchers test candidate directions, interpret the evidence, and carry the resulting lessons into later attempts. We study how an AI agent can run this loop autonomously over long horizons. We introduce Arbor, a general framework for autonomous research that combines a long-lived coordinator, short-lived executors, and Hypothesis Tree Refinement (HTR), a persistent tree that links hypotheses, artifacts, evidence, and distilled insights across time. The coordinator manages global research strategy over the tree, while executors implement and test individual hypotheses in isolated worktrees. As results return, Arbor updates the tree, propagates reusable lessons, refines the search frontier, and admits verified improvements. This design turns autonomous research from a sequence of local attempts into a cumulative process in which strategy, execution, and evidence are carried across time. We evaluate Arbor under Autonomous Optimization (AO), an operational setting where an agent improves an initial research artifact through iterative experimentation without step-level human supervision. Across six real research tasks in model training, harness engineering, and data synthesis, Arbor achieves the best held-out result on all six tasks, attaining more than 2.5x the average relative held-out gain of Codex and Claude Code under the same task interface and resource budget. On MLE-Bench Lite, Arbor reaches 86.36% Any Medal with GPT-5.5, the strongest result in our comparison.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory

    math.FA 2026-07 conditional novelty 8.0 of 10

    AI-generated, human-verified proofs of five open Banach-space problems, including primariness of Lp(L1) and a unital Banach algebra that is not any Calkin algebra.

  2. Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Argus demonstrates that a fixed-weight, self-evolving multi-role agentic runtime with verification-gated persistence can achieve competitive benchmark results and retain reusable state across long-horizon tasks.

  3. Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Fetch-then-Explore, which stores fetched pages in a per-question workspace and extracts evidence on demand with grep/read, beats visit-and-read and browsing baselines on BrowseComp across three LLM backbones.

  4. Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

    cs.CL 2026-08 conditional novelty 6.0 of 10

    LLM judges of scientific ideas are measurably swayed by writing style; a style-detecting module reduces but does not remove the bias.

  5. Towards Autonomous and Auditable Medical Imaging Model Development

    cs.CV 2026-07 conditional novelty 6.0 of 10

    AMID, a verification-guided multi-agent MLE system for medical imaging, outperforms general MLE agents on 20 ReX-MLE challenges and approaches human challenge solutions on several tasks.

  6. Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    HEP externalizes hypothesis generation, evidence-driven belief updates, and lifecycle verdicts so LLM agents run an auditable hypothesis-test-evidence-belief cycle on materials research questions.

Reference graph

Works this paper leans on

127 extracted references · 92 canonical work pages · cited by 6 Pith papers

  1. [1]

    ACM Transactions on Information Systems , author =

    William Webber and Alistair Moffat and Justin Zobel , title =. 2010 , url =. doi:10.1145/1852102.1852106 , timestamp =

  2. [2]

    Widesearch: Benchmarking agentic broad info-seeking.arXiv preprint arXiv:2508.07999, 2025

    Ryan Wong and Jiawei Wang and Junjie Zhao and Li Chen and Yan Gao and Long Zhang and Xuan Zhou and Zuo Wang and Kai Xiang and Ge Zhang and Wenhao Huang and Yang Wang and Ke Wang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.07999 , eprinttype =. 2508.07999 , timestamp =

  3. [4]

    2025 , howpublished =

    Keller Jordan and contributors , title =. 2025 , howpublished =

  4. [6]

    BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese

    Peilin Zhou and Bruce Leon and Xiang Ying and Can Zhang and Yifan Shao and Qichen Ye and Dading Chong and Zhiling Jin and Chenxuan Xie and Meng Cao and Yuxin Gu and Sixin Hong and Jing Ren and Jian Chen and Chao Liu and Yining Hua , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2504.19314 , eprinttype =. 2504.19314 , timestamp =

  5. [7]

    , year =

    Kidd, Celeste and Hayden, Benjamin Y. , year =. The Psychology and Neuroscience of Curiosity , volume =. Neuron , publisher =. doi:10.1016/j.neuron.2015.09.010 , number =

  6. [8]

    The Twelfth International Conference on Learning Representations,

    Gr. The Twelfth International Conference on Learning Representations,. 2024 , url =

  7. [9]

    Infodeepseek: Benchmarking agentic information seeking for retrieval- augmented generation,

    Yunjia Xi and Jianghao Lin and Menghui Zhu and Yongzhao Xiao and Zhuoying Ou and Jiaqi Liu and Tong Wan and Bo Chen and Weiwen Liu and Yasheng Wang and Ruiming Tang and Weinan Zhang and Yong Yu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.15872 , eprinttype =. 2505.15872 , timestamp =

  8. [10]

    CoRR , volume =

    Tian Lan and Bin Zhu and Qianghuai Jia and Junyang Ren and Haijun Li and Longyue Wang and Zhao Xu and Weihua Luo and Kaifu Zhang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2510.20168 , eprinttype =. 2510.20168 , timestamp =

Show all 127 references
  1. [11]

    CoRR , volume =

    Junting Zhou and Wang Li and Yiyan Liao and Nengyuan Zhang and Tingjia Miao and Zhihui Qi and Yuhan Wu and Tong Yang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2506.13784 , eprinttype =. 2506.13784 , timestamp =

  2. [12]

    xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations , journal =

    Kaiyuan Chen and Yixin Ren and Yang Liu and Xiaobo Hu and Haotong Tian and Tianbao Xie and Fangfu Liu and Haoye Zhang and Hongzhang Liu and Yuan Gong and Chen Sun and Han Hou and Hui Yang and James Pan and Jianan Lou and Jiayi Mao and Jizheng Liu and Jinpeng Li and Kangyi Liu ...

  3. [13]

    SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models , journal =

    Thinh Pham and Nguyen Nguyen and Pratibha Zunjare and Weiyuan Chen and Yu. SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2506.01062 , eprinttype =. 2506.01062 , timestamp =

  4. [14]

    CoRR , volume =

    Yilong Xu and Xiang Long and Zhi Zheng and Jinhua Gao , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2507.16725 , eprinttype =. 2507.16725 , timestamp =

  5. [15]

    CoRR , volume =

    Tomer Wolfson and Harsh Trivedi and Mor Geva and Yoav Goldberg and Dan Roth and Tushar Khot and Ashish Sabharwal and Reut Tsarfaty , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.11133 , eprinttype =. 2508.11133 , timestamp =

  6. [16]

    CoRR , volume =

    Heng Zhou and Ao Yu and Yuchen Fan and Jianing Shi and Li Kang and Hejia Geng and Yongting Zhang and Yutao Fan and Yuhao Wu and Tiancheng He and Yiran Qin and Lei Bai and Zhenfei Yin , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2511.01409 , eprinttype =. 2511.0...

  7. [17]

    Large Language Models for Information Retrieval:

    Yutao Zhu and Huaying Yuan and Shuting Wang and Jiongnan Liu and Wenhan Liu and Chenlong Deng and Zhicheng Dou and Ji. Large Language Models for Information Retrieval:. CoRR , volume =. 2023 , url =. doi:10.48550/ARXIV.2308.07107 , eprinttype =. 2308.07107 , timestamp =

  8. [18]

    CoRR , volume =

    Yunjia Xi and Jianghao Lin and Yongzhao Xiao and Zheli Zhou and Rong Shan and Te Gao and Jiachen Zhu and Weiwen Liu and Yong Yu and Weinan Zhang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.05668 , eprinttype =. 2508.05668 , timestamp =

  9. [19]

    Aggarwal and Hui Liu and Xiang Zhang and Suhang Wang , title =

    Minhua Lin and Zongyu Wu and Zhichao Xu and Hui Liu and Xianfeng Tang and Qi He and Charu C. Aggarwal and Hui Liu and Xiang Zhang and Suhang Wang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2510.16724 , eprinttype =. 2510.16724 , timestamp =

  10. [20]

    CoRR , volume =

    Xiaoxi Li and Guanting Dong and Jiajie Jin and Yuyao Zhang and Yujia Zhou and Yutao Zhu and Peitian Zhang and Zhicheng Dou , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2501.05366 , eprinttype =. 2501.05366 , timestamp =

  11. [21]

    WebThinker: Empowering Large Reasoning Models with Deep Research Capability , journal =

    Xiaoxi Li and Jiajie Jin and Guanting Dong and Hongjin Qian and Yutao Zhu and Yongkang Wu and Ji. WebThinker: Empowering Large Reasoning Models with Deep Research Capability , journal =. 2025 , url =. doi:10.48550/ARXIV.2504.21776 , eprinttype =. 2504.21776 , timestamp =

  12. [22]

    CoRR , volume =

    Bowen Jin and Hansi Zeng and Zhenrui Yue and Dong Wang and Hamed Zamani and Jiawei Han , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.09516 , eprinttype =. 2503.09516 , timestamp =

  13. [23]

    CoRR , volume =

    Kuan Li and Zhongwang Zhang and Huifeng Yin and Liwen Zhang and Litu Ou and Jialong Wu and Wenbiao Yin and Baixuan Li and Zhengwei Tao and Xinyu Wang and Weizhou Shen and Junkai Zhang and Dingchu Zhang and Xixi Wu and Yong Jiang and Ming Yan and Pengjun Xie and Fei Huang and J...

  14. [24]

    CoRR , volume =

    Rui Lu and Zhenyu Hou and Zihan Wang and Hanchen Zhang and Xiao Liu and Yujiang Li and Shi Feng and Jie Tang and Yuxiao Dong , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2509.10446 , eprinttype =. 2509.10446 , timestamp =

  15. [25]

    CoRR , volume =

    Dayoon Ko and Jihyuk Kim and Haeju Park and Sohyeon Kim and Dahyun Lee and Yongrae Jo and Gunhee Kim and Moontae Lee and Kyungjae Lee , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2508.19113 , eprinttype =. 2508.19113 , timestamp =

  16. [26]

    CoRR , volume =

    Baixuan Li and Dingchu Zhang and Jialong Wu and Wenbiao Yin and Zhengwei Tao and Yida Zhao and Liwen Zhang and Haiyang Shen and Runnan Fang and Pengjun Xie and Jingren Zhou and Yong Jiang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2510.24698 , eprinttype =. 2...

  17. [27]

    CoRR , volume =

    Lisheng Huang and Yichen Liu and Jinhao Jiang and Rongxiang Zhang and Jiahao Yan and Junyi Li and Wayne Xin Zhao , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.18105 , eprinttype =. 2505.18105 , timestamp =

  18. [28]

    Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking , journal =

    Ruiyang Ren and Yuhao Wang and Junyi Li and Jinhao Jiang and Wayne Xin Zhao and Wenjie Wang and Tat. Holistically Guided Monte Carlo Tree Search for Intricate Information Seeking , journal =. 2025 , url =. doi:10.48550/ARXIV.2502.04751 , eprinttype =. 2502.04751 , timestamp =

  19. [30]

    Yangyang Yu and Zhiyuan Yao and Haohang Li and Zhiyang Deng and Yuechen Jiang and Yupeng Cao and Zhi Chen and Jordan W. Suchow and Zhenyu Cui and Rong Liu and Zhaozhuo Xu and Denghui Zhang and Koduvayur Subbalakshmi and Guojun Xiong and Yueru He and Jimin Huang and Dong Li and...

  20. [31]

    ChatDev: Communicative Agents for Software Development , booktitle =

    Chen Qian and Wei Liu and Hongzhang Liu and Nuo Chen and Yufan Dang and Jiahao Li and Cheng Yang and Weize Chen and Yusheng Su and Xin Cong and Juyuan Xu and Dahai Li and Zhiyuan Liu and Maosong Sun , editor =. ChatDev: Communicative Agents for Software Development , booktitle...

  21. [32]

    CoRR , volume =

    Shanghua Gao and Ada Fang and Yepeng Huang and Valentina Giunchiglia and Ayush Noori and Jonathan Richard Schwarz and Yasha Ektefaie and Jovana Kondic and Marinka Zitnik , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2404.02831 , eprinttype =. 2404.02831 , timestamp =

  22. [33]

    WebWalker: Benchmarking LLMs in Web Traversal , booktitle =

    Jialong Wu and Wenbiao Yin and Yong Jiang and Zhenglin Wang and Zekun Xi and Runnan Fang and Linhai Zhang and Yulan He and Deyu Zhou and Pengjun Xie and Fei Huang , editor =. WebWalker: Benchmarking LLMs in Web Traversal , booktitle =. 2025 , url =

  23. [34]

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , journal =

    DeepSeek. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2512.02556 , eprinttype =. 2512.02556 , timestamp =

  24. [35]

    Narasimhan and Yuan Cao , title =

    Shunyu Yao and Jeffrey Zhao and Dian Yu and Nan Du and Izhak Shafran and Karthik R. Narasimhan and Yuan Cao , title =. The Eleventh International Conference on Learning Representations,. 2023 , url =

  25. [36]

    CoRR , volume =

    An Yang and Anfeng Li and Baosong Yang and Beichen Zhang and Binyuan Hui and Bo Zheng and Bowen Yu and Chang Gao and Chengen Huang and Chenxu Lv and Chujie Zheng and Dayiheng Liu and Fan Zhou and Fei Huang and Feng Hu and Hao Ge and Haoran Wei and Huan Lin and Jialong Tang and...

  26. [37]

    CoRR , volume =

    Aohan Zeng and Xin Lv and Qinkai Zheng and Zhenyu Hou and Bin Chen and Chengxing Xie and Cunxiang Wang and Da Yin and Hao Zeng and Jiajie Zhang and Kedong Wang and Lucen Zhong and Mingdao Liu and Rui Lu and Shulin Cao and Xiaohan Zhang and Xuancheng Huang and Yao Wei and Yean ...

  27. [38]

    DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments , booktitle =

    Yuxiang Zheng and Dayuan Fu and Xiangkun Hu and Xiaojie Cai and Lyumanshan Ye and Pengrui Lu and Pengfei Liu , editor =. DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments , booktitle =. 2025 , url =. doi:10.18653/V1/2025.EMNLP-MAIN.22 ...

  28. [39]

    R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning , journal =

    Huatong Song and Jinhao Jiang and Yingqian Min and Jie Chen and Zhipeng Chen and Wayne Xin Zhao and Lei Fang and Ji. R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning , journal =. 2025 , url =. doi:10.48550/ARXIV.2503.05592 , eprinttype =. 250...

  29. [40]

    CoRR , volume =

    Baixuan Li and Bo Zhang and Dingchu Zhang and Fei Huang and Guangyu Li and Guoxin Chen and Huifeng Yin and Jialong Wu and Jingren Zhou and Kuan Li and Liangcai Su and Litu Ou and Liwen Zhang and Pengjun Xie and Rui Ye and Wenbiao Yin and Xinmiao Yu and Xinyu Wang and Xixi Wu a...

  30. [41]

    2026 , url =

    Harness design for long-running application development , author =. 2026 , url =

  31. [43]

    2026 , url =

    Harness engineering: leveraging Codex in an agent-first world , author =. 2026 , url =

  32. [44]

    CoRR , volume =

    Jialong Wu and Baixuan Li and Runnan Fang and Wenbiao Yin and Liwen Zhang and Zhengwei Tao and Dingchu Zhang and Zekun Xi and Yong Jiang and Pengjun Xie and Fei Huang and Jingren Zhou , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.22648 , eprinttype =. 2505....

  33. [45]

    CoRR , volume =

    Zhengwei Tao and Jialong Wu and Wenbiao Yin and Junkai Zhang and Baixuan Li and Haiyang Shen and Kuan Li and Liwen Zhang and Xinyu Wang and Yong Jiang and Pengjun Xie and Fei Huang and Jingren Zhou , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2507.15061 , eprin...

  34. [46]

    CoRR , volume =

    Hao Sun and Zile Qiao and Jiayan Guo and Xuanbo Fan and Yingyan Hou and Yong Jiang and Pengjun Xie and Yan Zhang and Fei Huang and Jingren Zhou , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.04588 , eprinttype =. 2505.04588 , timestamp =

  35. [47]

    Kanell and Peter Xu and Omar Khattab and Monica S

    Yijia Shao and Yucheng Jiang and Theodore A. Kanell and Peter Xu and Omar Khattab and Monica S. Lam , editor =. Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models , booktitle =. 2024 , url =. doi:10.18653/V1/2024.NAACL-LONG.347 , timestamp =

  36. [48]

    Daya Guo and Dejian Yang and Haowei Zhang and Junxiao Song and Peiyi Wang and Qihao Zhu and Runxin Xu and Ruoyu Zhang and Shirong Ma and Xiao Bi and Xiaokang Zhang and Xingkai Yu and Yu Wu and Z. F. Wu and Zhibin Gou and Zhihong Shao and Zhuoshu Li and Ziyi Gao and Aixin Liu a...

  37. [49]

    Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

    Timo Schick and Jane Dwivedi. Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =. 2023 , url =

  38. [50]

    The Twelfth International Conference on Learning Representations,

    Yujia Qin and Shihao Liang and Yining Ye and Kunlun Zhu and Lan Yan and Yaxi Lu and Yankai Lin and Xin Cong and Xiangru Tang and Bill Qian and Sihan Zhao and Lauren Hong and Runchu Tian and Ruobing Xie and Jie Zhou and Mark Gerstein and Dahai Li and Zhiyuan Liu and Maosong Sun...

  39. [51]

    HuggingGPT: Solving

    Yongliang Shen and Kaitao Song and Xu Tan and Dongsheng Li and Weiming Lu and Yueting Zhuang , editor =. HuggingGPT: Solving. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, U...

  40. [52]

    Reflexion: language agents with verbal reinforcement learning , booktitle =

    Noah Shinn and Federico Cassano and Ashwin Gopinath and Karthik Narasimhan and Shunyu Yao , editor =. Reflexion: language agents with verbal reinforcement learning , booktitle =. 2023 , url =

  41. [57]

    MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation , booktitle =

    Qian Huang and Jian Vora and Percy Liang and Jure Leskovec , editor =. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation , booktitle =. 2024 , url =

  42. [58]

    CoRR , volume =

    Reiichiro Nakano and Jacob Hilton and Suchir Balaji and Jeff Wu and Long Ouyang and Christina Kim and Christopher Hesse and Shantanu Jain and Vineet Kosaraju and William Saunders and Xu Jiang and Karl Cobbe and Tyna Eloundou and Gretchen Krueger and Kevin Button and Matthew Kn...

  43. [63]

    arXiv preprint arXiv:2503.18102 , year =

    Samuel Schmidgall and Michael Moor , title =. arXiv preprint arXiv:2503.18102 , year =

  44. [64]

    arXiv preprint arXiv:2505.18705 , year =

    Jiabin Tang and Lianghao Xia and Zhonghang Li and Chao Huang , title =. arXiv preprint arXiv:2505.18705 , year =

  45. [65]

    bioRxiv , year =

    Kexin Huang and Serena Zhang and Hanchen Wang and Yuanhao Qu and Yingzhou Lu and Yusuf Roohani and Ryan Li and Lin Qiu and Junze Zhang and Yin Di and others , title =. bioRxiv , year =

  46. [67]

    2026 , howpublished =

    Jiaqi Liu and Peng Xia and Siwei Han and Shi Qiu and Letian Zhang and Guiming Chen and Haoqin Tu and Xinyu Yang and Jiawei Zhou and Hongtu Zhu and Yun Li and Jiaheng Zhang and Yuyin Zhou and Zeyu Zheng and Cihang Xie and Mingyu Ding and Huaxiu Yao , title =. 2026 , howpublished =

  47. [68]

    2026 , howpublished =

    Andrej Karpathy , title =. 2026 , howpublished =

  48. [79]

    arXiv preprint arXiv:2503.21248 , year =

    Yujie Liu and Zonglin Yang and Tong Xie and Jinjie Ni and Ben Gao and Yuqiang Li and Shixiang Tang and Wanli Ouyang and Erik Cambria and Dongzhan Zhou , title =. arXiv preprint arXiv:2503.21248 , year =

  49. [81]

    Siegel and Sayash Kapoor and Nitya Nadgir and Benedikt Stroebl and Arvind Narayanan , title =

    Zachary S. Siegel and Sayash Kapoor and Nitya Nadgir and Benedikt Stroebl and Arvind Narayanan , title =. arXiv preprint arXiv:2409.11363 , year =

  50. [82]

    arXiv preprint arXiv:2407.01725 , year =

    Bodhisattwa Prasad Majumder and Harshit Surana and Dhruv Agarwal and Bhavana Dalvi Mishra and Abhijeetsingh Meena and Aryan Prakhar and Tirth Vora and Tushar Khot and Ashish Sabharwal and Peter Clark , title =. arXiv preprint arXiv:2407.01725 , year =

  51. [84]

    arXiv preprint arXiv:2505.19955 , year =

    Hui Chen and Miao Xiong and Yujie Lu and Wei Han and Ailin Deng and Yang He and Jiaying Wu and Kai Wang and Yibo Wang and Shen Li and Jiani Yu and Bryan Hooi , title =. arXiv preprint arXiv:2505.19955 , year =

  52. [85]

    Browne and Edward Powley and Daniel Whitehouse and Simon M

    Cameron B. Browne and Edward Powley and Daniel Whitehouse and Simon M. Lucas and Peter I. Cowling and Philipp Rohlfshagen and Stephen Tavener and Diego Perez and Spyridon Samothrakis and Simon Colton , title =. IEEE Transactions on Computational Intelligence and AI in Games , volume =

  53. [91]

    2026 , eprint=

    DataMaster: Data-Centric Autonomous AI Research , author=. 2026 , eprint=

  54. [93]

    2025 , month = mar, organization =

    Automated Researchers Can Subtly Sandbag , author =. 2025 , month = mar, organization =

  55. [97]

    2026 , eprint=

    AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning , author=. 2026 , eprint=

  56. [107]

    2026 , eprint=

    SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent , author=. 2026 , eprint=

  57. [109]

    CoRR , volume =

    Jiafeng Liang and Hao Li and Chang Li and Jiaqi Zhou and Shixin Jiang and Zekun Wang and Changkai Ji and Zhihao Zhu and Runxuan Liu and Tao Ren and Jinlan Fu and See. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2512.23343 , eprinttype =. 2512.23343 , timestamp =

  58. [110]

    The Thirteenth International Conference on Learning Representations,

    Jiayi Zhang and Jinyu Xiang and Zhaoyang Yu and Fengwei Teng and Xionghui Chen and Jiaqi Chen and Mingchen Zhuge and Xin Cheng and Sirui Hong and Jinlin Wang and Bingnan Zheng and Bang Liu and Yuyu Luo and Chenglin Wu , title =. The Thirteenth International Conference on Learn...

  59. [111]

    2025 , howpublished =

  60. [112]

    Xu and Xiangru Tang and Mingchen Zhuge and Jiayi Pan and Yueqi Song and Bowen Li and Jaskirat Singh and Hoang H

    Xingyao Wang and Boxuan Li and Yufan Song and Frank F. Xu and Xiangru Tang and Mingchen Zhuge and Jiayi Pan and Yueqi Song and Bowen Li and Jaskirat Singh and Hoang H. Tran and Fuqiang Li and Ren Ma and Mingzhang Zheng and Bill Qian and Yanjun Shao and Niklas Muennighoff and Y...

  61. [113]

    Claude Code

    Anthropic . Claude Code . https://github.com/anthropics/claude-code, 2025. Agentic coding tool for terminal, IDE, and GitHub workflows. Accessed: 2026-06-02

  62. [114]

    Scimaster: Towards general-purpose scientific AI agents, part i

    Jingyi Chai, Shuo Tang, Rui Ye, Yuwen Du, Xinyu Zhu, Mengcheng Zhou, Yanfeng Wang, Weinan E, Yuzhi Zhang, Linfeng Zhang, and Siheng Chen. Scimaster: Towards general-purpose scientific AI agents, part i. x-master as foundation: Can we lead on humanity's last exam? CoRR, abs/250...

  63. [115]

    Mle-bench: Evaluating machine learning agents on machine learning engineering

    Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, Dane Sherburn, Evan Mays, Giulio Starace, Kevin Liu, Leon Maksin, Tejal Patwardhan, Lilian Weng, and Aleksander Madry. Mle-bench: Evaluating machine learning agents on machine learning engineering. CoRR, abs/2410.07095,...

  64. [116]

    Toward autonomous long-horizon engineering for ML research

    Guoxin Chen, Jie Chen, Lei Chen, Jiale Zhao, Fanzhe Meng, Wayne Xin Zhao, Ruihua Song, Cheng Chen, Ji - Rong Wen, and Kai Jia. Toward autonomous long-horizon engineering for ML research. CoRR, abs/2604.13018, 2026 a . doi:10.48550/ARXIV.2604.13018. https://doi.org/10.48550/arX...

  65. [117]

    MARS: modular agent with reflective search for automated AI research

    Jiefeng Chen, Bhavana Dalvi Mishra, Jaehyun Nam, Rui Meng, Tomas Pfister, and Jinsung Yoon. MARS: modular agent with reflective search for automated AI research. CoRR, abs/2602.02660, 2026 b . doi:10.48550/ARXIV.2602.02660. https://doi.org/10.48550/arXiv.2602.02660

  66. [118]

    Baker, Benjamin Burns, Daniel Adu - Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun

    Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, Vishal Dey, Mingyi Xue, Frazier N. Baker, Benjamin Burns, Daniel Adu - Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun. Scienceagentbench: Towar...

  67. [119]

    Frontier-eng: Benchmarking self-evolving agents on real-world engineering tasks with generative optimization

    Yizhe Chi, Deyao Hong, Dapeng Jiang, Tianwei Luo, Kaisen Yang, Boshi Zhang, Zhe Cao, Xiaoyan Fan, Bingxiang He, Han Hao, Weiyang Jin, Dianqiao Lei, Qingle Liu, Houde Qian, Bowen Wang, Situ Wang, Youjie Zheng, Yifan Zhou, Calvin Xiao, Eren Cai, and Qinhuai Na. Frontier-eng: Ben...

  68. [120]

    Datamaster: Data-centric autonomous ai research, 2026

    Yaxin Du, Xiyuan Yang, Zhifan Zhou, Wanxu Liu, Zixing Lei, Zimeng Chen, Fenyi Liu, Haotian Wu, Yuzhu Cai, Zexi Liu, Xinyu Zhu, WenHao Wang, Linfeng Zhang, Chen Qian, and Siheng Chen. Datamaster: Data-centric autonomous ai research, 2026. https://arxiv.org/abs/2605.10906

  69. [121]

    Automated researchers can subtly sandbag, March 2025

    Johannes Gasteiger, Akbir Khan, Sam Bowman, Vladimir Mikulik, Ethan Perez, and Fabien Roger. Automated researchers can subtly sandbag, March 2025. https://alignment.anthropic.com/2025/automated-researchers-sandbag/

  70. [122]

    Memory in the age of AI agents

    Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang, Boyang Liu, Fangyi Zhu, Jiahang Lin, Honglin Guo, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiejun Tan, Yanbin Yin, Jiongnan Liu, Zeyu Zhang, Zhongxiang Sun, Yutao Zhu, Hao Sun, Boci Peng, Zhenrong Cheng, Xuanbo Fan, Jiaxin Guo, Xinl...

  71. [123]

    Agentfugue: Agent scaling for long-horizon tasks through collective reasoning, 2026 a

    Yuyang Hu, Hongjin Qian, Shuting Wang, Jiongnan Liu, Tong Zhao, Xiaoxi Li, Zheng Liu, and Zhicheng Dou. Agentfugue: Agent scaling for long-horizon tasks through collective reasoning, 2026 a . https://arxiv.org/abs/2605.24486

  72. [124]

    Sam: State-adaptive memory for long-horizon reasoning agent, 2026 b

    Yuyang Hu, Hongjin Qian, Shuting Wang, Jiongnan Liu, Ziliang Zhao, Jiejun Tan, Zheng Liu, and Zhicheng Dou. Sam: State-adaptive memory for long-horizon reasoning agent, 2026 b . https://arxiv.org/abs/2605.24468

  73. [125]

    Mlagentbench: Evaluating language agents on machine learning experimentation

    Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec. Mlagentbench: Evaluating language agents on machine learning experimentation. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors, Forty-...

  74. [126]

    AIDE: ai-driven exploration in the space of code

    Zhengyao Jiang, Dominik Schmidt, Dhruv Srikanth, Dixing Xu, Ian Kaplan, Deniss Jacenko, and Yuxiang Wu. AIDE: ai-driven exploration in the space of code. CoRR, abs/2502.13138, 2025. doi:10.48550/ARXIV.2502.13138. https://doi.org/10.48550/arXiv.2502.13138

  75. [127]

    NanoGPT-Bench : NanoGPT training speedrun benchmark

    Keller Jordan and contributors. NanoGPT-Bench : NanoGPT training speedrun benchmark. https://github.com/KellerJordan/modded-nanogpt, 2025

  76. [128]

    autoresearch: AI agents running research on single- GPU nanochat training automatically

    Andrej Karpathy. autoresearch: AI agents running research on single- GPU nanochat training automatically. https://github.com/karpathy/autoresearch, 2026

  77. [129]

    Ziegler, Elizabeth Barnes, and Lawrence Chan

    Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Neev Parikh, David Rein, Lucas...

  78. [130]

    Meta-harness: End-to-end optimization of model harnesses

    Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn. Meta-harness: End-to-end optimization of model harnesses. CoRR, abs/2603.28052, 2026. doi:10.48550/ARXIV.2603.28052. https://doi.org/10.48550/arXiv.2603.28052

  79. [131]

    The FM agent

    Annan Li, Chufan Wu, Zengle Ge, Yee Hin Chong, Zhinan Hou, Lizhe Cao, Cheng Ju, Jianmin Wu, Huaiming Li, Haobo Zhang, Shenghao Feng, Mo Zhao, Fengzhi Qiu, Rui Yang, Mengmeng Zhang, Wenyi Zhu, Yingying Sun, Quan Sun, Shunhao Yan, Danyu Liu, Dawei Yin, and Dou Shen. The FM agent...

  80. [132]

    Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses

    Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Xuanjing Huang, Hang Yan, Zhenhua Han, and Tao Gui. Agentic harness engineering: Observability-driven automatic evolution of coding-agent harnesses. CoRR, abs/2604.25850, 2026. doi:10.48550/ARXIV.2604.25850. https:...

  81. [133]

    A comprehensive survey on long context language modeling

    Jiaheng Liu, Dawei Zhu, Zhiqi Bai, Yancheng He, Huanxuan Liao, Haoran Que, Zekun Wang, Chenchen Zhang, Ge Zhang, Jiebin Zhang, Yuanxing Zhang, Zhuo Chen, Hangyu Guo, Shilong Li, Ziqiang Liu, Yong Shan, Yifan Song, Jiayi Tian, Wenhao Wu, Zhejian Zhou, Ruijie Zhu, Junlan Feng, Y...

  82. [134]

    Ml-master: Towards ai-for-ai via integration of exploration and reasoning

    Zexi Liu, Yuzhu Cai, Xinyu Zhu, Yujie Zheng, Runkun Chen, Ying Wen, Yanfeng Wang, Weinan E, and Siheng Chen. Ml-master: Towards ai-for-ai via integration of exploration and reasoning. CoRR, abs/2506.16499, 2025 b . doi:10.48550/ARXIV.2506.16499. https://doi.org/10.48550/arXiv....

  83. [135]

    Xinghua Lou, Miguel L \' a zaro - Gredilla, Antoine Dedieu, Carter Wendelken, Wolfgang Lehrach, and Kevin P. Murphy. Autoharness: improving LLM agents by automatically synthesizing a code harness. CoRR, abs/2603.03329, 2026. doi:10.48550/ARXIV.2603.03329. https://doi.org/10.48...

  84. [136]

    Foerster, Jeff Clune, and David Ha

    Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob N. Foerster, Jeff Clune, and David Ha. The AI scientist: Towards fully automated open-ended scientific discovery. CoRR, abs/2408.06292, 2024. doi:10.48550/ARXIV.2408.06292. https://doi.org/10.48550/arXiv.2408.06292

  85. [137]

    Mike A. Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, Estefany Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Jenia Jitsev, Di Lu, Orfeas Men...

  86. [138]

    KAPSO: A knowledge-grounded framework for autonomous program synthesis and optimization

    Alireza Nadafian, Alireza Mohammadshahi, and Majid Yazdani. KAPSO: A knowledge-grounded framework for autonomous program synthesis and optimization. CoRR, abs/2601.21526, 2026. doi:10.48550/ARXIV.2601.21526. https://doi.org/10.48550/arXiv.2601.21526

  87. [139]

    Alexander Novikov, Ng \^ a n Vu, Marvin Eisenberger, Emilien Dupont, Po - Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin,...

  88. [140]

    Codex CLI

    OpenAI . Codex CLI . https://github.com/openai/codex, 2025. Lightweight coding agent that runs locally on a user's computer. Accessed: 2026-06-02

  89. [141]

    O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S

    Joon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Sean Follmer, Jeff Han, J \" u rgen Steimle, and Nathalie Henry Riche, editors, Proceedings of the 3...

  90. [142]

    Ori Press, Brandon Amos, Haoyu Zhao, Yikai Wu, Samuel K. Ainsworth, Dominik Krupke, Patrick Kidger, Touqir Sajed, Bartolomeo Stellato, Jisun Park, Nathanael Bosch, Eli Meril, Albert Steppi, Arman Zharmagambetov, Fangzhao Zhang, David Perez - Pineiro, Alberto Mercurio, Ni Zhan,...

  91. [143]

    K, Rongzhi Zhang, Changhao Li, Ian Shu - Hei Wong, Sherry Yang, Percy Liang, Chao Zhang, and Bo Dai

    Rushi Qiang, Yuchen Zhuang, Yinghao Li, Dingu Sagar V. K, Rongzhi Zhang, Changhao Li, Ian Shu - Hei Wong, Sherry Yang, Percy Liang, Chao Zhang, and Bo Dai. Mle-dojo: Interactive environments for empowering LLM agents in machine learning engineering. CoRR, abs/2505.07782, 2025....

  92. [144]

    Toolllm: Facilitating large language models to master 16000+ real-world apis

    Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. Toolllm: Facilitating large language models to m...

  93. [145]

    Posttrainbench: Can LLM agents automate LLM post-training? CoRR, abs/2603.08640, 2026

    Ben Rank, Hardik Bhatnagar, Ameya Prabhu, Shira Eisenberg, Karina Nguyen, Matthias Bethge, and Maksym Andriushchenko. Posttrainbench: Can LLM agents automate LLM post-training? CoRR, abs/2603.08640, 2026. doi:10.48550/ARXIV.2603.08640. https://doi.org/10.48550/arXiv.2603.08640

  94. [146]

    HCAST: human-calibrated autonomy software tasks

    David Rein, Joel Becker, Amy Deng, Seraphina Nix, Chris Canal, Daniel O'Connel, Pip Arnott, Ryan Bloom, Thomas Broadley, Katharyn Garcia, Brian Goodrich, Max Hasin, Sami Jawhar, Megan Kinniment, Thomas Kwa, Aron Lajko, Nate Rush, Lucas Jun Koba Sato, Sydney von Arx, Ben West, ...

  95. [147]

    Pawan Kumar, Emilien Dupont, Francisco J

    Bernardino Romera - Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi. Mathematical discoveries from program search with la...

  96. [148]

    Toolformer: Language models can teach themselves to use tools

    Timo Schick, Jane Dwivedi - Yu, Roberto Dess \` , Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Morit...

  97. [149]

    Agent laboratory: Using LLM agents as research assistants

    Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Zicheng Liu, and Emad Barsoum. Agent laboratory: Using LLM agents as research assistants. CoRR, abs/2501.04227, 2025. doi:10.48550/ARXIV.2501.04227. https://doi.org/10.48550/arXiv.2501.04227

  98. [150]

    Reflexion: language agents with verbal reinforcement learning

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. Reflexion: language agents with verbal reinforcement learning. In Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine, editors, Advances in Neural Information...

  99. [151]

    The illusion of diminishing returns: Measuring long horizon execution in llms

    Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping. The illusion of diminishing returns: Measuring long horizon execution in llms. CoRR, abs/2509.09677, 2025. doi:10.48550/ARXIV.2509.09677. https://doi.org/10.48550/arXiv.2509.09677

  100. [152]

    Paperbench: Evaluating ai's ability to replicate AI research

    Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. Paperbench: Evaluating ai's ability to replicate AI research. CoRR, abs/2504...

  101. [153]

    Internagent-1.5: A unified agentic framework for long-horizon autonomous scientific discovery

    InternScience Team. Internagent-1.5: A unified agentic framework for long-horizon autonomous scientific discovery. CoRR, abs/2602.08990, 2026. doi:10.48550/ARXIV.2602.08990. https://doi.org/10.48550/arXiv.2602.08990

  102. [154]

    A survey of AI scientists

    Guiyao Tie, Pan Zhou, and Lichao Sun. A survey of AI scientists. CoRR, abs/2510.23045, 2025. doi:10.48550/ARXIV.2510.23045. https://doi.org/10.48550/arXiv.2510.23045

  103. [155]

    Miller, Abhishek Charnalia, Derek Dunfield, Carole - Jean Wu, Pontus Stenetorp, Nicola Cancedda, Jakob Nicolaus Foerster, and Yoram Bachrach

    Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Rishi Hazra, Nicolas Mario Baldwin, Alexis Audran - Reiss, Michael Kuchnik, Despoina Magka, Minqi Jiang, Alisia Maria Lupidi, Andrei Lupu, Roberta Raileanu, Kelvin Niu, Tatiana Shavrina, Jean - Christophe Gagnon - Audet, Mich...

  104. [156]

    Loongflow: Directed evolutionary search via a cognitive plan-execute-summarize paradigm

    Chunhui Wan, Xunan Dai, Zhuo Wang, Minglei Li, Yanpeng Wang, Yinan Mao, Yu Lan, and Zhiwen Xiao. Loongflow: Directed evolutionary search via a cognitive plan-execute-summarize paradigm. CoRR, abs/2512.24077, 2025. doi:10.48550/ARXIV.2512.24077. https://doi.org/10.48550/arXiv.2...

  105. [157]

    Frontierscience: Evaluating ai's ability to perform expert-level scientific tasks

    Miles Wang, Robi Lin, Kat Hu, Joy Jiao, Neil Chowdhury, Ethan Chang, and Tejal Patwardhan. Frontierscience: Evaluating ai's ability to perform expert-level scientific tasks. CoRR, abs/2601.21165, 2026. doi:10.48550/ARXIV.2601.21165. https://doi.org/10.48550/arXiv.2601.21165

  106. [158]

    Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

    Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, and et al. Op...

  107. [159]

    Browsecomp: A simple yet challenging benchmark for browsing agents

    Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. CoRR, abs/2504.12516, 2025. doi:10.48550/ARXIV.2504.1251...

  108. [160]

    Re-bench: Evaluating frontier AI r & d capabilities of language model agents against human experts

    Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Joshua Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Sato, William Saunders, Ma...

  109. [161]

    Foerster, Jeff Clune, and David Ha

    Yutaro Yamada, Robert Tjarko Lange, Cong Lu, Shengran Hu, Chris Lu, Jakob N. Foerster, Jeff Clune, and David Ha. The AI scientist-v2: Workshop-level automated scientific discovery via agentic tree search. CoRR, abs/2504.08066, 2025. doi:10.48550/ARXIV.2504.08066. https://doi.o...

  110. [162]

    R & d-agent: Automating data-driven AI solution building through llm-powered automated research, development, and evolution

    Xu Yang, Xiao Yang, Shikai Fang, Bowen Xian, Yuante Li, Jian Wang, Minrui Xu, Haoran Pan, Xinpeng Hong, Weiqing Liu, Yelong Shen, Weizhu Chen, and Jiang Bian. R & d-agent: Automating data-driven AI solution building through llm-powered automated research, development, and evol...

  111. [163]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. CoRR, abs/2305.10601, 2023 a . doi:10.48550/ARXIV.2305.10601. https://doi.org/10.48550/arXiv.2305.10601

  112. [164]

    Narasimhan, and Yuan Cao

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023 . OpenRevie...

  113. [165]

    Lange, and Jeff Clune

    Jenny Zhang, Shengran Hu, Cong Lu, Robert T. Lange, and Jeff Clune. Darwin godel machine: Open-ended evolution of self-improving agents. CoRR, abs/2505.22954, 2025 a . doi:10.48550/ARXIV.2505.22954. https://doi.org/10.48550/arXiv.2505.22954

  114. [166]

    Aflow: Automating agentic workflow generation

    Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, Bingnan Zheng, Bang Liu, Yuyu Luo, and Chenglin Wu. Aflow: Automating agentic workflow generation. In The Thirteenth International Conference on ...

  115. [167]

    Agentic context engineering: Evolving contexts for self-improving language models

    Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, Urmish Thakker, James Zou, and Kunle Olukotun. Agentic context engineering: Evolving contexts for self-improving language models. CoRR, abs...

  116. [168]

    Aibuildai: An AI agent for automatically building AI models

    Ruiyi Zhang, Peijia Qin, Qi Cao, Li Zhang, and Pengtao Xie. Aibuildai: An AI agent for automatically building AI models. CoRR, abs/2604.14455, 2026. doi:10.48550/ARXIV.2604.14455. https://doi.org/10.48550/arXiv.2604.14455

  117. [169]

    Miller, Oisin Mac Aodha, Jakob N

    Bingchen Zhao, Despoina Magka, Minqi Jiang, Xian Li, Roberta Raileanu, Tatiana Shavrina, Jean - Christophe Gagnon - Audet, Kelvin Niu, Shagun Sodhani, Michael Shvartsman, Andrei Lupu, Alisia Maria Lupidi, Edan Toledo, Karen Hambardzumyan, Martin Josifoski, Thomas Foster, Lucia...

  118. [170]

    Language agent tree search unifies reasoning acting and planning in language models

    Andy Zhou, Kai Yan, Michal Shlapentokh - Rothman, Haohan Wang, and Yu - Xiong Wang. Language agent tree search unifies reasoning acting and planning in language models. CoRR, abs/2310.04406, 2023. doi:10.48550/ARXIV.2310.04406. https://doi.org/10.48550/arXiv.2310.04406

  119. [171]

    Toward ultra-long-horizon agentic science: Cognitive accumulation for machine learning engineering

    Xinyu Zhu, Yuzhu Cai, Zexi Liu, Bingyang Zheng, Cheng Wang, Rui Ye, Jiaao Chen, Hanrui Wang, Wei - Chen Wang, Yuzhi Zhang, Linfeng Zhang, Weinan E, Di Jin, Siheng Chen, and Yanfeng Wang. Toward ultra-long-horizon agentic science: Cognitive accumulation for machine learning eng...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.