A new benchmark shows search-augmented LLMs correctly refuse only up to 42.9% of unanswerable multi-hop questions, with failures split between hallucinated answers and search exhaustion.
16 Topology-Controlled Query Fusion (continued) The first thing the agent must resolve is already unanswerable
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning
A new benchmark shows search-augmented LLMs correctly refuse only up to 42.9% of unanswerable multi-hop questions, with failures split between hallucinated answers and search exhaustion.