← back to paper
arxiv: 2608.07371 · 2 revisions
Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning