AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
https: //arxiv.org/abs/2512.07921
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 10roles
background 3polarities
background 3representative citing papers
An LLM-powered agentic framework autonomously designs competitive and sometimes superior explainable algorithms for wireless PHY and MAC layer tasks.
Distilling GitHub repositories and papers into verified, loadable skills raises a fixed Codex agent's score on MLE-bench from 31.1% to 72.9%, with smaller gains on three other research benchmarks.
Mixed-methods study creates taxonomy of AI IDE rules from 7310 instances, analyzes evolution drivers, and reports that rule updates raise average artifact compliance from 49.14% to 72.13%.
Proposes agentic framework-based reproduction with a slot-binding interface to turn 16 PHM papers into standardized, assumption-aware benchmark implementations.
ARA uses LLMs to build workflow graphs linking sources, methods, and outputs in papers, then scores reproducibility, reaching ~61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldStandardDB.
QMP-Bench supplies a realistic test set for AI on quantum many-body problems while PhysVEC uses integrated verifiers to turn unreliable LLM generations into code that passes both syntax and physics checks, outperforming baselines.
A hierarchical multi-agent system with File-as-Bus durable state beats matched baselines on PaperBench and MLE-Bench Lite, and ablations show project-state continuity drives later-round gains.
State-aware iterative subplanning in DeepRepro outperforms static-planning coding agents on PaperBench Code-Dev paper-to-code reproduction, averaging 84.2 on a five-paper subset.
DietDelta uses vision-language prompts on paired before-and-after RGB images to localize food items, estimate their weights, and compute consumption differences, reporting better results than prior single-image methods on three public datasets.
citing papers explorer
-
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
-
The AI Telco Engineer: Toward Autonomous Discovery of Wireless Communications Algorithms
An LLM-powered agentic framework autonomously designs competitive and sometimes superior explainable algorithms for wireless PHY and MAC layer tasks.
-
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
Distilling GitHub repositories and papers into verified, loadable skills raises a fixed Codex agent's score on MLE-bench from 31.1% to 72.9%, with smaller gains on three other research benchmarks.
-
Rule Taxonomy and Evolution in AI IDEs: A Mining and Survey Study
Mixed-methods study creates taxonomy of AI IDE rules from 7310 instances, analyzes evolution drivers, and reports that rule updates raise average artifact compliance from 49.14% to 72.13%.
-
From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence
Proposes agentic framework-based reproduction with a slot-binding interface to turn 16 PHM papers into standardized, assumption-aware benchmark implementations.
-
ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review
ARA uses LLMs to build workflow graphs linking sources, methods, and outputs in papers, then scores reproducibility, reaching ~61% accuracy on 213 ReScience C articles and outperforming priors on ReproBench and GoldStandardDB.
-
Towards Verifiable and Self-Correcting AI Physicists for Quantum Many-Body Simulations
QMP-Bench supplies a realistic test set for AI on quantum many-body problems while PhysVEC uses integrated verifiers to turn unreliable LLM generations into code that passes both syntax and physics checks, outperforming baselines.
-
Toward Autonomous Long-Horizon Engineering for ML Research
A hierarchical multi-agent system with File-as-Bus durable state beats matched baselines on PaperBench and MLE-Bench Lite, and ablations show project-state continuity drives later-round gains.
-
DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories
State-aware iterative subplanning in DeepRepro outperforms static-planning coding agents on PaperBench Code-Dev paper-to-code reproduction, averaging 84.2 on a five-paper subset.
-
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
DietDelta uses vision-language prompts on paired before-and-after RGB images to localize food items, estimate their weights, and compute consumption differences, reporting better results than prior single-image methods on three public datasets.