A new benchmark evaluates AI deep research systems on 65 frontier AI questions, finding OpenAI and Gemini lead on rubric coverage while all systems cite accurately but leave much content unsupported.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
A new benchmark evaluates AI deep research systems on 65 frontier AI questions, finding OpenAI and Gemini lead on rubric coverage while all systems cite accurately but leave much content unsupported.