A new RL algorithm estimates the optimal value function offline from the reference model and then regresses the policy log-ratio to the optimal advantage, enabling single-rollout per prompt training with competitive accuracy and up to 2x speedup.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
A new RL algorithm estimates the optimal value function offline from the reference model and then regresses the policy log-ratio to the optimal advantage, enabling single-rollout per prompt training with competitive accuracy and up to 2x speedup.