NOVER uses the policy model's perplexity of a reference answer given its own reasoning as a GRPO reward, and reports average gains of about 7.7 percentage points over a 7B R1-distilled model across four task families.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NOVER: Incentive Training for Language Models via Verifier-Free Reinforcement Learning
NOVER uses the policy model's perplexity of a reference answer given its own reasoning as a GRPO reward, and reports average gains of about 7.7 percentage points over a 7B R1-distilled model across four task families.