Pith. sign in

REVIEW 1 cited by

TD-M(PC)$^2$: Improving Temporal Difference MPC Through Policy Constraint

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.03550 v1 pith:LGMCAM66 submitted 2025-02-05 cs.LG cs.RO

classification cs.LGcs.RO
keywords policyvaluedatalearningexistingexperimentsimprovinglearned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based reinforcement learning algorithms that combine model-based planning and learned value/policy prior have gained significant recognition for their high data efficiency and superior performance in continuous control. However, we discover that existing methods that rely on standard SAC-style policy iteration for value learning, directly using data generated by the planner, often result in \emph{persistent value overestimation}. Through theoretical analysis and experiments, we argue that this issue is deeply rooted in the structural policy mismatch between the data generation policy that is always bootstrapped by the planner and the learned policy prior. To mitigate such a mismatch in a minimalist way, we propose a policy regularization term reducing out-of-distribution (OOD) queries, thereby improving value learning. Our method involves minimum changes on top of existing frameworks and requires no additional computation. Extensive experiments demonstrate that the proposed approach improves performance over baselines such as TD-MPC2 by large margins, particularly in 61-DoF humanoid tasks. View qualitative results at https://darthutopian.github.io/tdmpc_square/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

    cs.RO 2025-06 conditional novelty 6.0 of 10

    DoublyAware combines conformal trajectory filtering with a group-relative policy constraint to improve sample efficiency of TD-MPC for simulated humanoid locomotion.

Pith tools