TD-GRPC trains a humanoid locomotion policy by combining softmax group-relative Q-value advantages with an MPPI-action matching loss inside TD-MPC, reporting improved sample efficiency on eight of ten HumanoidBench tasks.
Model predictive control,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
TD-GRPC trains a humanoid locomotion policy by combining softmax group-relative Q-value advantages with an MPPI-action matching loss inside TD-MPC, reporting improved sample efficiency on eight of ten HumanoidBench tasks.