TP-GRPO uses thought-level process rewards from a generative judge to reach higher math accuracy than outcome-only GRPO with fewer policy updates.
\n\n" delimiter. However, we observed in practice that the model often generates
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
TP-GRPO uses thought-level process rewards from a generative judge to reach higher math accuracy than outcome-only GRPO with fewer policy updates.