A 40-task, 10-domain benchmark finds current auto-research agents and LLMs score only ~20–26 on re-discovering real published scientific artifacts.
Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , pages=
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
years
2026 3representative citing papers
SkillReranker decomposes tasks and skills into state transitions, builds an execution graph, and adaptively selects skills per task stage, improving agent performance on ALFWorld and ScienceWorld.
citing papers explorer
-
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
A 40-task, 10-domain benchmark finds current auto-research agents and LLMs score only ~20–26 on re-discovering real published scientific artifacts.
-
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
SkillReranker decomposes tasks and skills into state transitions, builds an execution graph, and adaptively selects skills per task stage, improving agent performance on ALFWorld and ScienceWorld.
- DORA Explorer: Improving the Exploration Ability of LLMs Without Training