A trajectory-relative normalization of hindsight distillation signals, giving each turn a multiplier with token-weighted mean one, improves agentic RL performance over GRPO on WebShop and ALFWorld.
Group-in-Group Policy Optimization for
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning
A trajectory-relative normalization of hindsight distillation signals, giving each turn a multiplier with token-weighted mean one, improves agentic RL performance over GRPO on WebShop and ALFWorld.