Policy improve- ment using language feedback models.Advances in Neural Information Processing Systems, 37:43730–43758, 2024

Victor Zhong, Dipendra Misra, Xingdi Yuan, Marc-Alexandre Côté · 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

browse 1 citing papers

representative citing papers

Learning from Language Feedback via Variational Policy Distillation

cs.LG · 2026-05-14 · unverdicted · novelty 7.0

VPD frames language feedback learning as variational EM so the teacher policy refines itself via trust-region updates on outcomes while the student learns dense token distributions on its own rollouts, outperforming fixed-teacher baselines on reasoning and code tasks.

citing papers explorer

Showing 1 of 1 citing paper.

Learning from Language Feedback via Variational Policy Distillation cs.LG · 2026-05-14 · unverdicted · none · ref 52
VPD frames language feedback learning as variational EM so the teacher policy refines itself via trust-region updates on outcomes while the student learns dense token distributions on its own rollouts, outperforming fixed-teacher baselines on reasoning and code tasks.

Policy improve- ment using language feedback models.Advances in Neural Information Processing Systems, 37:43730–43758, 2024

fields

years

verdicts

representative citing papers

citing papers explorer