HiL-Bench shows frontier AI agents fail to ask for help on incomplete tasks, recovering only a fraction of full-information performance, but RL training on Ask-F1 reward improves judgment and transfers across domains.
ConvAI3 : Generating clarifying questions for open-domain dialogue systems ( ClariQ )
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Post-clarification answering remains the bottleneck in multi-turn QA despite rapid gains in clarification policy via supervised fine-tuning on the PACIFIC benchmark.
citing papers explorer
-
HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?
HiL-Bench shows frontier AI agents fail to ask for help on incomplete tasks, recovering only a fraction of full-information performance, but RL training on Ask-F1 reward improves judgment and transfers across domains.
-
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA
Post-clarification answering remains the bottleneck in multi-turn QA despite rapid gains in clarification policy via supervised fine-tuning on the PACIFIC benchmark.