Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Accelerating RL for LLM Reasoning with Optimal Advantage Regression

cs.LG · 2025-05-27 · conditional · novelty 6.0

A new RL algorithm estimates the optimal value function offline from the reference model and then regresses the policy log-ratio to the optimal advantage, enabling single-rollout per prompt training with competitive accuracy and up to 2x speedup.

citing papers explorer

Showing 1 of 1 citing paper.

  • Accelerating RL for LLM Reasoning with Optimal Advantage Regression cs.LG · 2025-05-27 · conditional · none · ref 1

    A new RL algorithm estimates the optimal value function offline from the reference model and then regresses the policy log-ratio to the optimal advantage, enabling single-rollout per prompt training with competitive accuracy and up to 2x speedup.