AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
Infoflow: Reinforcing search agent via reward density optimization
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
Retrievers trained on agent trajectories via the LRAT framework improve evidence recall, task success, and efficiency in agentic search benchmarks.
CAPF improves Qwen3-4B exact-match scores from 44.7% to 48.5% on seven QA benchmarks by allowing privileged verifier feedback during RLVR training with attenuated credit for the feedback step.
citing papers explorer
-
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
-
Learning to Retrieve from Agent Trajectories
Retrievers trained on agent trajectories via the LRAT framework improve evidence recall, task success, and efficiency in agentic search benchmarks.
-
CAPF: Guiding Search-Agent Rollouts with Credit-Attenuated Privileged Feedback
CAPF improves Qwen3-4B exact-match scores from 44.7% to 48.5% on seven QA benchmarks by allowing privileged verifier feedback during RLVR training with attenuated credit for the feedback step.