PSPO-WRS adds a Weibull-based reward that depends on reasoning chain length to process-supervised RLHF, and reports accuracy gains over baseline LLMs on six NLI-style reasoning datasets.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
PSPO*: An Effective Process-supervised Policy Optimization for Reasoning Alignment
PSPO-WRS adds a Weibull-based reward that depends on reasoning chain length to process-supervised RLHF, and reports accuracy gains over baseline LLMs on six NLI-style reasoning datasets.