Pith. sign in

REVIEW 2 cited by

Can only LLMs do Reasoning?: Potential of Small Language Models in Task Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03891 v1 pith:5MD2GFO5 submitted 2024-04-05 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords llmstaskcommandsdatasetsdomainplannersrobotssmall
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In robotics, the use of Large Language Models (LLMs) is becoming prevalent, especially for understanding human commands. In particular, LLMs are utilized as domain-agnostic task planners for high-level human commands. LLMs are capable of Chain-of-Thought (CoT) reasoning, and this allows LLMs to be task planners. However, we need to consider that modern robots still struggle to perform complex actions, and the domains where robots can be deployed are limited in practice. This leads us to pose a question: If small LMs can be trained to reason in chains within a single domain, would even small LMs be good task planners for the robots? To train smaller LMs to reason in chains, we build `COmmand-STeps datasets' (COST) consisting of high-level commands along with corresponding actionable low-level steps, via LLMs. We release not only our datasets but also the prompt templates used to generate them, to allow anyone to build datasets for their domain. We compare GPT3.5 and GPT4 with the finetuned GPT2 for task domains, in tabletop and kitchen environments, and the result shows that GPT2-medium is comparable to GPT3.5 for task planning in a specific domain. Our dataset, code, and more output samples can be found in https://github.com/Gawon-Choi/small-LMs-Task-Planning

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Goal-oriented Communication for Fast and Robust Robotic Fault Detection and Recovery

    cs.RO 2026-01 conditional novelty 5.0 of 10

    A goal-oriented communication framework using 3D scene graphs, edge points, and a fine-tuned small language model substantially reduces simulated fault detection and recovery time while improving task success.

  2. Enhancing the Reasoning Capabilities of Small Language Models via Solution Guidance Fine-Tuning

    cs.CL 2024-12 conditional novelty 4.0 of 10

    SGFT fine-tunes a small model to produce calculation-free solution plans and uses a second model to answer from them, outperforming CoT fine-tuning with roughly 3% of the training data.

Pith tools