VersaPRM, fine-tuned on LLM-generated and auto-labeled multi-domain chain-of-thought data, beats math-only process reward models and majority voting on non-math MMLU-Pro categories.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
VersaPRM, fine-tuned on LLM-generated and auto-labeled multi-domain chain-of-thought data, beats math-only process reward models and majority voting on non-math MMLU-Pro categories.