Pith. sign in

REVIEW 2 cited by

See, Plan, Predict: Language-guided Cognitive Planning with Video Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.03825 v1 pith:Q6VOUKF5 submitted 2022-10-07 cs.AI cs.RO

classification cs.AIcs.RO
keywords planninglanguagenaturalvideocognitivepredictiongroundingability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Cognitive planning is the structural decomposition of complex tasks into a sequence of future behaviors. In the computational setting, performing cognitive planning entails grounding plans and concepts in one or more modalities in order to leverage them for low level control. Since real-world tasks are often described in natural language, we devise a cognitive planning algorithm via language-guided video prediction. Current video prediction models do not support conditioning on natural language instructions. Therefore, we propose a new video prediction architecture which leverages the power of pre-trained transformers.The network is endowed with the ability to ground concepts based on natural language input with generalization to unseen objects. We demonstrate the effectiveness of this approach on a new simulation dataset, where each task is defined by a high-level action described in natural language. Our experiments compare our method again stone video generation baseline without planning or action grounding and showcase significant improvements. Our ablation studies highlight an improved generalization to unseen objects that natural language embeddings offer to concept grounding ability, as well as the importance of planning towards visual "imagination" of a task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Vis2Plan extracts object symbols from unlabeled play videos with vision models, plans symbolically with A* search, and retrieves reachable real images as subgoals for a goal-conditioned robot policy.

  2. Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following

    cs.RO 2025-02 conditional novelty 6.0 of 10

    A temporal alignment auxiliary loss on goal and language representations improves zero-shot compositional generalization in robot instruction following.

Pith tools