TurnOPD improves on-policy distillation for long-horizon agents by adaptively budgeting rollout depth and progressively shifting KL loss from token-level to turn-balanced weighting, achieving up to 2.29x faster training with better accuracy on ALFWorld, WebShop, and Multi-Hop Search.
When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
TurnOPD improves on-policy distillation for long-horizon agents by adaptively budgeting rollout depth and progressively shifting KL loss from token-level to turn-balanced weighting, achieving up to 2.29x faster training with better accuracy on ALFWorld, WebShop, and Multi-Hop Search.