DeepWeb-Bench is a benchmark requiring massive cross-source evidence collection and long-horizon derivation, with evaluations on nine frontier models showing derivation and calibration as primary failure modes.
Browsecomp: A simple yet challenging benchmark for browsing agents
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.AI 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
LLM agents improve adaptability by first using an interaction budget for systematic exploration measured via Exploration Checkpoint Coverage before executing tasks.
citing papers explorer
-
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
DeepWeb-Bench is a benchmark requiring massive cross-source evidence collection and long-horizon derivation, with evaluations on nine frontier models showing derivation and calibration as primary failure modes.
-
Look Before You Leap: Autonomous Exploration for LLM Agents
LLM agents improve adaptability by first using an interaction budget for systematic exploration measured via Exploration Checkpoint Coverage before executing tasks.