LatentRM treats reasoning traces as discrete latent variables and trains a generator end-to-end so that the scalar reward model's likelihood of the true preference ranking is maximized, outperforming scalar, generative, and hybrid reward models on average.
Saurous, Rif , booktitle =
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning Latent Reasoning Traces for Scalar Reward Models End-to-End
LatentRM treats reasoning traces as discrete latent variables and trains a generator end-to-end so that the scalar reward model's likelihood of the true preference ranking is maximized, outperforming scalar, generative, and hybrid reward models on average.