A new travel-planning benchmark and multi-agent system that still mostly fails, with the best system passing only 2.72% of test cases.
In Proceedings of the 60th Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 1024–1034
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RETAIL: Towards Real-world Travel Planning for Large Language Models
A new travel-planning benchmark and multi-agent system that still mostly fails, with the best system passing only 2.72% of test cases.