Fine-tuning a language model on its worst-scoring responses, using a CVaR-style schedule, reduces negative and toxic generations more than standard RLHF on IMDB, Jigsaw, and RealToxicityPrompts.
WHEN I AM UNBLOCKED I SWEAR I WILL GO F**K YOUR M C
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Risk-Averse Finetuning of Large Language Models
Fine-tuning a language model on its worst-scoring responses, using a CVaR-style schedule, reduces negative and toxic generations more than standard RLHF on IMDB, Jigsaw, and RealToxicityPrompts.