AR-Bench shows that LLMs struggle with active reasoning: even GPT-4o solves only 35% of the guessing-number task and 54% of detective cases, while humans reach 80-100%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
AR-Bench shows that LLMs struggle with active reasoning: even GPT-4o solves only 35% of the guessing-number task and 54% of detective cases, while humans reach 80-100%.