Pith. sign in

Open-reasoner- zero: An open source approach to scaling reinforcement learning on the base model.https://github.com/Ope n-Reasoner-Zero/Open-Reasoner-Zero, 2025

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Weak-to-Strong Generalization via Direct On-Policy Distillation

cs.LG · 2026-07-06 · conditional · novelty 6.0

Transferring a weak model’s RL-induced log-ratio policy shift on a strong student’s own rollouts raises AIME accuracy more cheaply than imitating the weak teacher or running matched-step RL on the student.

citing papers explorer

Showing 1 of 1 citing paper.

  • Weak-to-Strong Generalization via Direct On-Policy Distillation cs.LG · 2026-07-06 · conditional · none · ref 13

    Transferring a weak model’s RL-induced log-ratio policy shift on a strong student’s own rollouts raises AIME accuracy more cheaply than imitating the weak teacher or running matched-step RL on the student.