AutoRule extracts rules from reasoning chains over preference data, scores responses against them with an LLM verifier, and uses the scores as an auxiliary reward, improving Llama-3-8B alignment metrics and showing reduced reward hacking.
Length-controlled alpacaeval: A simple debiasing of automatic evaluators
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
dataset 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
dataset 1polarities
use dataset 1representative citing papers
citing papers explorer
-
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
AutoRule extracts rules from reasoning chains over preference data, scores responses against them with an LLM verifier, and uses the scores as an auxiliary reward, improving Llama-3-8B alignment metrics and showing reduced reward hacking.