OpenRFT adapts a reasoning foundation model to eight scientific tasks with 100 samples each via data augmentation, self-distilled reasoning SFT, and PPO with a process reward model, achieving 0.447 vs 0.403 average accuracy.
Sciknoweval: Evaluating multi-level scientific knowledge of large language models, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning
OpenRFT adapts a reasoning foundation model to eight scientific tasks with 100 samples each via data augmentation, self-distilled reasoning SFT, and PPO with a process reward model, achieving 0.447 vs 0.403 average accuracy.