RRL replays value-model-selected promising intermediate states during LLM RL training, preserving exploration and improving performance on code, math, and RLHF tasks.
In: 2008 Interna- tional Conference on Computational Intelligence for Modelling Control Automa- tion
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Improving RL Exploration for LLM Reasoning through Retrospective Replay
RRL replays value-model-selected promising intermediate states during LLM RL training, preserving exploration and improving performance on code, math, and RLHF tasks.