Pith. sign in

REVIEW 4 cited by

On the Planning Abilities of Large Language Models : A Critical Investigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15771 v2 pith:GUX4OV6H submitted 2023-05-25 cs.AI

classification cs.AI
keywords llmsplanningplansllm-moduloautonomouslycapabilitiesdomainsevaluate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intrigued by the claims of emergent reasoning capabilities in LLMs trained on general web corpora, in this paper, we set out to investigate their planning capabilities. We aim to evaluate (1) the effectiveness of LLMs in generating plans autonomously in commonsense planning tasks and (2) the potential of LLMs in LLM-Modulo settings where they act as a source of heuristic guidance for external planners and verifiers. We conduct a systematic study by generating a suite of instances on domains similar to the ones employed in the International Planning Competition and evaluate LLMs in two distinct modes: autonomous and heuristic. Our findings reveal that LLMs' ability to generate executable plans autonomously is rather limited, with the best model (GPT-4) having an average success rate of ~12% across the domains. However, the results in the LLM-Modulo setting show more promise. In the LLM-Modulo setting, we demonstrate that LLM-generated plans can improve the search process for underlying sound planners and additionally show that external verifiers can help provide feedback on the generated plans and back-prompt the LLM for better plan generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 52 citations worldwide. Full citation record

  1. Training Small LLMs as Spatial Multi-Agent Policies

    cs.MA 2026-08 conditional novelty 7.0 of 10

    Small frozen LLMs trained over state-filtered symbolic option menus with per-agent LoRA adapters reach competent play in three cooperative spatial games, while behavioral audits show reward and cooperation decouple.

  2. TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

    cs.CL 2026-07 conditional novelty 6.5 of 10

    On 800 jointly constrained trip tasks with a deterministic scorer and achievable gold, the best of 15 LLM agents fully solves only 46.2% of feasible plans, with unstated persona needs as the universal bottleneck.

  3. Information-seeking failures of large language models in agentic clinical reasoning

    cs.AI 2026-07 conditional novelty 6.5 of 10

    LLMs systematically under-request critical molecular and cytogenetic data in multi-round oncology, capping accuracy at 68% despite high knowledge scores and coherent reasoning traces.

  4. AnnoBench: A Benchmark for Visualization Annotation Generation

    cs.HC 2026-07 conditional novelty 6.0 of 10

    A benchmark for chart annotation generation with a five-dimensional rubric shows current LLMs annotate code-based charts well but distort raster charts and over-rely on explicit instructions.

Pith tools