Adding a frozen POLAR-based similarity penalty to SFT, GRPO, or CHORD improves averaged instruction-following scores by up to 5.77 percent at 0.6B scale in the paper's reported runs.
A survey on post-training of large language models.arXiv e-prints, pages arXiv–2503, 2025
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Adding a frozen POLAR-based similarity penalty to SFT, GRPO, or CHORD improves averaged instruction-following scores by up to 5.77 percent at 0.6B scale in the paper's reported runs.