Temporal GRPO splits a robot rollout into detectable task stages and applies separate group-relative policy advantages to each stage's action interval, improving success rates by 7 points on average over matched baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.RO 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning
Temporal GRPO splits a robot rollout into detectable task stages and applies separate group-relative policy advantages to each stage's action interval, improving success rates by 7 points on average over matched baselines.