Evaluators with higher rationality scores showed higher test-retest consistency and lower bias deviation in a small RLHF feedback experiment.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Governance Challenges in Reinforcement Learning from Human Feedback: Evaluator Rationality and Reinforcement Stability
Evaluators with higher rationality scores showed higher test-retest consistency and lower bias deviation in a small RLHF feedback experiment.