The chain of alignment method derives model behavior rules from publicly supported objectives and yields an automated reward that tracks expert ratings of response alignment (r=0.841).
Direct preference-based policy optimization without reward modeling
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.HC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
The chain of alignment method derives model behavior rules from publicly supported objectives and yields an automated reward that tracks expert ratings of response alignment (r=0.841).