Pith. sign in

REVIEW 5 cited by

A Survey on Large Language Models for Automated Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12435 v1 pith:I4UGFEJV submitted 2025-02-18 cs.AI cs.CL

classification cs.AIcs.CL
keywords planningllmsmodelsabilityautomatedlanguagelargelimitations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The planning ability of Large Language Models (LLMs) has garnered increasing attention in recent years due to their remarkable capacity for multi-step reasoning and their ability to generalize across a wide range of domains. While some researchers emphasize the potential of LLMs to perform complex planning tasks, others highlight significant limitations in their performance, particularly when these models are tasked with handling the intricacies of long-horizon reasoning. In this survey, we critically investigate existing research on the use of LLMs in automated planning, examining both their successes and shortcomings in detail. We illustrate that although LLMs are not well-suited to serve as standalone planners because of these limitations, they nonetheless present an enormous opportunity to enhance planning applications when combined with other approaches. Thus, we advocate for a balanced methodology that leverages the inherent flexibility and generalized knowledge of LLMs alongside the rigor and cost-effectiveness of traditional planning methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents

    cs.AI 2026-07 conditional novelty 6.0 of 10

    An open-source agent-debugging loop attributes failures to the responsible step and uses the diagnosis to repair failed runs, recovering 13 of 73 GAIA tasks.

  2. SHERPA: A Model-Driven Framework for Large Language Model Execution

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A framework that executes LLM tasks through hierarchical state machines improves output quality in 12 of 15 comparisons, but the evaluation lacks error bars and includes test-set-informed design choices.

  3. Graph-Based Physics-Guided Urban PM2.5 Air Quality Imputation with Constrained Monitoring Data

    cs.LG 2025-06 conditional novelty 5.0 of 10

    GraPhy, a physics-inspired graph neural network with wind-based edge features and learnable diffusion scaling, reports the best PM2.5 imputation accuracy among six baselines on 41 sensors in Fresno, California.

  4. TextAtari: 100K Frames Game Playing with Language Agents

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.

  5. Large Language Models for Planning: A Comprehensive and Systematic Survey

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A structured survey of LLM planning methods, benchmarks, and interpretability work, organized around a three-way taxonomy.

Pith tools