Pith. sign in

REVIEW 2 cited by

Few-shot Subgoal Planning with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.14288 v1 pith:3AFTNEK3 submitted 2022-05-28 cs.CL

classification cs.CL
keywords languagesubgoalmodelspre-trainedsequencesinfermethodssupervision
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pre-trained large language models have shown successful progress in many language understanding benchmarks. This work explores the capability of these models to predict actionable plans in real-world environments. Given a text instruction, we show that language priors encoded in pre-trained language models allow us to infer fine-grained subgoal sequences. In contrast to recent methods which make strong assumptions about subgoal supervision, our experiments show that language models can infer detailed subgoal sequences from few training sequences without any fine-tuning. We further propose a simple strategy to re-rank language model predictions based on interaction and feedback from the environment. Combined with pre-trained navigation and visual reasoning components, our approach demonstrates competitive performance on subgoal prediction and task completion in the ALFRED benchmark compared to prior methods that assume more subgoal supervision.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems

    cs.RO 2025-04 conditional novelty 6.0 of 10

    A two-step soft-prompt backdoor attack called Robo-Troj (listed as MuTRAP on arXiv) makes LLM-based robot planners emit malicious plans when hidden trigger words are present, with near-perfect attack success.

  2. LLM-Based Offline Learning for Embodied Agents via Consistency-Guided Reward Ensemble

    cs.AI 2024-11 conditional novelty 6.0 of 10

    CoREN uses an LLM offline to estimate dense action rewards, filters them through three consistency checks, and aligns them to sparse success labels to train a small offline RL agent for household instruction-following tasks.

Pith tools