A new benchmark shows search-augmented LLMs correctly refuse only up to 42.9% of unanswerable multi-hop questions, with failures split between hallucinated answers and search exhaustion.
United States
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning
A new benchmark shows search-augmented LLMs correctly refuse only up to 42.9% of unanswerable multi-hop questions, with failures split between hallucinated answers and search exhaustion.