In a synthetic preference-learning setup, PA-RLHF is reported to beat a single averaged reward model on group-level alignment, but the experiment appears to use ground-truth group labels to select the reward model, undermining the comparison.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Procedural Fairness Failures in RLHF from Preference Averaging
In a synthetic preference-learning setup, PA-RLHF is reported to beat a single averaged reward model on group-level alignment, but the experiment appears to use ground-truth group labels to select the reward model, undermining the comparison.