Pith. sign in

REVIEW 7 cited by

Deep Research Bench: Evaluating AI Web Research Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.06287 v1 pith:2I66YXJV submitted 2025-05-06 cs.AI

classification cs.AI
keywords researchagentsdeepevaluationssearchagentbenchmajor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench, consisting of 89 multi-step web research task instances of varying difficulty across 8 diverse task categories, with the answers carefully worked out by skilled humans. We provide a "RetroSearch" environment with a large frozen set of scraped web pages, and demonstrate that offline "RetroSearch" agents perform comparably to "live web" agents, enabling reliable evaluations of models over time. We provide robust agent tooling and scaffolding to benchmark major LLMs as they are released, including "thinking" models like o3 and Gemini 2.5 Pro. We include automated evaluations of the lengthy agent traces to report progress over time in hallucinations, tool use, and forgetting. Finally, we evaluate the major web research products branded as "Deep Research", "Deep Search", "Search", or "Research." Results are available on a public leaderboard at https://drb.futuresearch.ai/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new benchmark, LegalCiteTrust, measures citation existence, fidelity, and applicability in Chinese legal research reports and shows that more legal retrieval does not automatically make citations more trustworthy.

  2. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A pre-registered benchmark of 13 LLM systems on all 104 World Cup 2026 matches shows fine-grained predictions expose differences that result accuracy hides.

  3. LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services

    cs.AI 2025-12 conditional novelty 6.0 of 10

    LocalSearchBench—1.3M merchant records and 900 multi-hop local-life QA tasks across 9 Chinese cities—shows the best reasoning agent reaches only 35.6% correctness.

  4. DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence

    cs.CL 2025-09 conditional novelty 6.0 of 10

    An audit framework and empirical study showing that generative search engines and deep research agents frequently produce one-sided answers and weakly supported citations, with citation accuracy between 40 and 80%.

  5. Characterizing Deep Research: A Benchmark and Formal Definition

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Deep research is characterized by high search and reasoning intensity; the new LiveDRBench measures claim-level precision and recall, where the best current model scores 0.55 F1.

  6. Bench to the Future: A Pastcasting Benchmark for Forecasting Agents

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Bench to the Future is a pastcasting benchmark: 299 already-resolved forecasting questions, each paired with a frozen corpus of about 20,000 web pages, on which newer LLMs and agentic search score better.

  7. Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A position paper advocating large-scale training of event forecasting LLMs, with proposals for label selection, counterfactual training data, auxiliary rewards, and multi-source datasets.

Pith tools