DAR aligns LLMs through advantage-weighted supervised fine-tuning on online AI scalar rewards with dual KL regularization, and reports win-rate gains over online RLHF and online preference methods.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Direct Advantage Regression: Aligning LLMs with Online AI Reward
DAR aligns LLMs through advantage-weighted supervised fine-tuning on online AI scalar rewards with dual KL regularization, and reports win-rate gains over online RLHF and online preference methods.