π-Bench is a new benchmark for evaluating proactive personal assistant agents on 100 multi-turn tasks that include hidden intents, inter-task dependencies, and cross-session continuity.
Learning to clarify: Multi-turn conversations with action-based contrastive self-training
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Post-clarification answering remains the bottleneck in multi-turn QA despite rapid gains in clarification policy via supervised fine-tuning on the PACIFIC benchmark.
citing papers explorer
-
$\pi$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows
π-Bench is a new benchmark for evaluating proactive personal assistant agents on 100 multi-turn tasks that include hidden intents, inter-task dependencies, and cross-session continuity.
-
Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA
Post-clarification answering remains the bottleneck in multi-turn QA despite rapid gains in clarification policy via supervised fine-tuning on the PACIFIC benchmark.