A self-iterative training loop that combines process reward models with ORPO improves reasoning accuracy of small language models on GSM8K and MBPP.
Any contradictions in the reasoning are incorrect
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning to Reason via Self-Iterative Process Feedback for Small Language Models
A self-iterative training loop that combines process reward models with ORPO improves reasoning accuracy of small language models on GSM8K and MBPP.