TD-GRPC trains a humanoid locomotion policy by combining softmax group-relative Q-value advantages with an MPPI-action matching loss inside TD-MPC, reporting improved sample efficiency on eight of ten HumanoidBench tasks.
Neural network dynamics for model-based deep reinforcement learning with model- free fine-tuning,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.RO 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
TD-GRPC: Temporal Difference Learning with Group Relative Policy Constraint for Humanoid Locomotion
TD-GRPC trains a humanoid locomotion policy by combining softmax group-relative Q-value advantages with an MPPI-action matching loss inside TD-MPC, reporting improved sample efficiency on eight of ten HumanoidBench tasks.