Pith. sign in

REVIEW 2 cited by

Leveraging Pre-trained Large Language Models to Construct and Utilize World Models for Model-based Task Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14909 v2 pith:NDIVVNIO submitted 2023-05-24 cs.AI

classification cs.AI
keywords pddlmodelsdomainllmsfeedbacklanguageplanningmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There is a growing interest in applying pre-trained large language models (LLMs) to planning problems. However, methods that use LLMs directly as planners are currently impractical due to several factors, including limited correctness of plans, strong reliance on feedback from interactions with simulators or even the actual environment, and the inefficiency in utilizing human feedback. In this work, we introduce a novel alternative paradigm that constructs an explicit world (domain) model in planning domain definition language (PDDL) and then uses it to plan with sound domain-independent planners. To address the fact that LLMs may not generate a fully functional PDDL model initially, we employ LLMs as an interface between PDDL and sources of corrective feedback, such as PDDL validators and humans. For users who lack a background in PDDL, we show that LLMs can translate PDDL into natural language and effectively encode corrective feedback back to the underlying domain model. Our framework not only enjoys the correctness guarantee offered by the external planners but also reduces human involvement by allowing users to correct domain models at the beginning, rather than inspecting and correcting (through interactive prompting) every generated plan as in previous work. On two IPC domains and a Household domain that is more complicated than commonly used benchmarks such as ALFWorld, we demonstrate that GPT-4 can be leveraged to produce high-quality PDDL models for over 40 actions, and the corrected PDDL models are then used to successfully solve 48 challenging planning tasks. Resources, including the source code, are released at: https://guansuns.github.io/pages/llm-dm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robotouille: An Asynchronous Planning Benchmark for LLM Agents

    cs.RO 2025-02 conditional novelty 7.0 of 10

    Robotouille is a new interactive benchmark showing LLM agents succeed on 47% of synchronous plans but only 11% when tasks require overlapping timed actions.

  2. Implicit Language Models are RNNs: Balancing Parallelization and Expressivity

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Implicit SSMs, which iterate to a fixed point, combine RNN-style state tracking with parallel training and scale to 1.3B parameters.

Pith tools