A 209-task Chinese benchmark across six dialogue-failure modes shows that even the top frontier model fully satisfies only 41.1% of long multi-turn requests.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding
A 209-task Chinese benchmark across six dialogue-failure modes shows that even the top frontier model fully satisfies only 41.1% of long multi-turn requests.