REVIEW 8 cited by
Robust Planning with LLM-Modulo Framework: Case Study in Travel Planning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
As the applicability of Large Language Models (LLMs) extends beyond traditional text processing tasks, there is a burgeoning interest in their potential to excel in planning and reasoning assignments, realms traditionally reserved for System 2 cognitive competencies. Despite their perceived versatility, the research community is still unraveling effective strategies to harness these models in such complex domains. The recent discourse introduced by the paper on LLM Modulo marks a significant stride, proposing a conceptual framework that enhances the integration of LLMs into diverse planning and reasoning activities. This workshop paper delves into the practical application of this framework within the domain of travel planning, presenting a specific instance of its implementation. We are using the Travel Planning benchmark by the OSU NLP group, a benchmark for evaluating the performance of LLMs in producing valid itineraries based on user queries presented in natural language. While popular methods of enhancing the reasoning abilities of LLMs such as Chain of Thought, ReAct, and Reflexion achieve a meager 0%, 0.6%, and 0% with GPT3.5-Turbo respectively, our operationalization of the LLM-Modulo framework for TravelPlanning domain provides a remarkable improvement, enhancing baseline performances by 4.6x for GPT4-Turbo and even more for older models like GPT3.5-Turbo from 0% to 5%. Furthermore, we highlight the other useful roles of LLMs in the planning pipeline, as suggested in LLM-Modulo, which can be reliably operationalized such as extraction of useful critics and reformulator for critics.
Forward citations
Cited by 8 Pith papers
-
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
On 800 jointly constrained trip tasks with a deterministic scorer and achievable gold, the best of 15 LLM agents fully solves only 46.2% of feasible plans, with unstated persona needs as the universal bottleneck.
-
We Argue to Agree: Towards Personality-Driven Argumentation-Based Negotiation Dialogue Systems for Tourism
A new LLM-generated dataset (PACT) of 8,687 personality-tagged, argumentation-annotated tourism negotiations, plus a three-part benchmark in which fine-tuned models beat zero-shot and human-human-data baselines.
-
Decompose, Plan in Parallel, and Merge: A Novel Paradigm for Large Language Models based Planning with Multiple Constraints
DPPM, a decompose-plan-in-parallel-and-merge framework with verify-and-refine feedback, improves final pass rates on travel-planning benchmarks over Direct, CoT, and LLM-Modulo.
-
MaskSearch: A Universal Pre-Training Framework to Enhance Agentic Search Capability
A pre-training task called RAMP, where models practice searching to fill masked text spans, improves downstream agentic open-domain QA performance across Qwen and LLaMA models.
-
Is Your LLM-Based Multi-Agent a Reliable Real-World Planner? Exploring Fraud Detection in Travel Planning
Multi-agent LLM travel planners are frequently deceived by injected fake listings, coordinated fake reviews, and multi-round scam conversations, and a simple anti-fraud reviewer helps only some models.
-
Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning
UDOIP combines dual-layer LLM preference extraction with clustering-and-substitution GRASP to produce higher-scoring, constraint-feasible multi-day urban itineraries than LLM-only and adapted solver baselines on two C...
-
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
RLFactory is a plug-and-play RL post-training framework for multi-turn tool use, reporting 0.486 on NQ with Qwen3-4B versus 0.473 for Qwen2.5-7B, and 6.8x faster training throughput.
-
TripTailor: A Real-World Benchmark for Personalized Travel Planning
A travel-planning benchmark is claimed in the abstract, but the full text is an unrelated supernova spectroscopy paper, leaving the central claim completely unsupported.
Discussion (0). Sign in to comment.