Pith. sign in

REVIEW 3 cited by

Chain-of-Thought Predictive Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.00776 v2 pith:CH3C3BE6 submitted 2023-04-03 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords chain-of-thoughtcontroldemoslearningproposegeneralizableguidancemethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study generalizable policy learning from demonstrations for complex low-level control (e.g., contact-rich object manipulations). We propose a novel hierarchical imitation learning method that utilizes sub-optimal demos. Firstly, we propose an observation space-agnostic approach that efficiently discovers the multi-step subskill decomposition of the demos in an unsupervised manner. By grouping temporarily close and functionally similar actions into subskill-level demo segments, the observations at the segment boundaries constitute a chain of planning steps for the task, which we refer to as the chain-of-thought (CoT). Next, we propose a Transformer-based design that effectively learns to predict the CoT as the subskill-level guidance. We couple action and subskill predictions via learnable prompt tokens and a hybrid masking strategy, which enable dynamically updated guidance at test time and improve feature representation of the trajectory for generalizable policy learning. Our method, Chain-of-Thought Predictive Control (CoTPC), consistently surpasses existing strong baselines on challenging manipulation tasks with sub-optimal demos.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M$^3$PC: Test-time Model Predictive Control for Pretrained Masked Trajectory Model

    cs.LG 2024-12 conditional novelty 7.0 of 10

    M3PC runs model predictive control at test time on a pretrained masked trajectory Transformer, improving offline RL returns and enabling goal reaching without extra model training.

  2. Hierarchical Diffusion Policy: manipulation trajectory generation via contact guidance

    cs.RO 2024-11 conditional novelty 5.0 of 10

    A two-layer diffusion policy, where a Guider predicts the next contact point and an Actor generates the trajectory toward it under Q-learning guidance, outperforms end-to-end Diffusion Policy on contact-rich manipulation.

  3. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools