VersaPRM, fine-tuned on LLM-generated and auto-labeled multi-domain chain-of-thought data, beats math-only process reward models and majority voting on non-math MMLU-Pro categories.
- Verifiable: The step can be verified using common knowledge, simple calculations, or a quick refer- ence (e.g., recalling a basic theorem)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
VersaPRM, fine-tuned on LLM-generated and auto-labeled multi-domain chain-of-thought data, beats math-only process reward models and majority voting on non-math MMLU-Pro categories.