A new multi-turn benchmark shows LLMs struggle to keep global constraints satisfied when local constraints are added later, and often sacrifice budget to satisfy soft preferences.
Mint: Evaluating llms in multi-turn interaction with tools and language feedback
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Flex-TravelPlanner: A Benchmark for Flexible Planning with Language Agents
A new multi-turn benchmark shows LLMs struggle to keep global constraints satisfied when local constraints are added later, and often sacrifice budget to satisfy soft preferences.