ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Bin Wang; Bo Zhang; Chaofan Hu; Chunfeng Song; Dongzhan Zhou; Fangchen Yu; Fenghua Ling; Guangtao Zhai; Haoxiang Yin; Haoxuan Li

arxiv: 2606.07591 · v1 · pith:IIHFIAMAnew · submitted 2026-05-28 · 💻 cs.LG · cs.AI· cs.CL

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

Wanghan Xu , Shuo Li , Tianlin Ye , Qinglong Cao , Yixin Chen , Hengjian Gao , Yiheng Wang , Qi Li

show 42 more authors

Kun Li Sheng Xu Shengdu Chai Fangchen Yu Xiangyu Zhao Zhangrui Zhao Weijie Ma Zijie Guo Haoyu Zhou Haoxiang Yin Lixue Cheng Chaofan Hu Haoxuan Li Lu Mi Xuxuan Xie Yifan Zhou Ruizhe Chen Zhiwang Zhou Xingjian Guo Yuhao Zhou Xuming He Shengyuan Xu Xinyu Gu Jiamin Wu Mianxin Liu Chunfeng Song Fenghua Ling Dongzhan Zhou Shixiang Tang Yuqiang Li Mao Su Peng Ye Siqi Sun Bin Wang Xue Yang Zhenfei Yin Tianfan Fu Guangtao Zhai Wanli Ouyang Bo Zhang Lei Bai Wenlong Zhang

This is my paper

classification 💻 cs.LG cs.AIcs.CL

keywords scientificautonomousresearchevaluationresearchclawbenchagentsaveragesbenchmark

0 comments

read the original abstract

AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating autonomous scientific research across 40 tasks from 10 scientific domains. Each task is grounded in a real published paper, provides related literature and raw data, and hides the target paper during evaluation. Expert-curated multimodal rubrics decompose the target scientific artifacts into weighted criteria, enabling evaluation of target-paper-level re-discovery while leaving room for new discovery. We evaluate seven autonomous research (auto-research) agents under a unified protocol and seventeen native LLMs through the lightweight ResearchHarness. Current systems remain far from reliable re-discovery: the strongest autonomous agent, Claude Code, averages 21.5, and the strongest ResearchHarness LLM, Claude-Opus-4.7, averages 20.7, with an LLM frontier mean of only 26.5. Error analysis shows that failures concentrate in experimental protocol mismatch, evidence mismatch, and missing scientific core. ResearchClawBench provides a reproducible evaluation frontier for measuring progress toward autonomous scientific research.

This paper has not been read by Pith yet.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

discussion (0)