A self-iterative training loop that combines process reward models with ORPO improves reasoning accuracy of small language models on GSM8K and MBPP.
In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 9426–9439
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning to Reason via Self-Iterative Process Feedback for Small Language Models
A self-iterative training loop that combines process reward models with ORPO improves reasoning accuracy of small language models on GSM8K and MBPP.