Safe Delta preserves LLM safety after fine-tuning by pruning delta parameters with a utility-per-safety ratio and compensating the safety loss with an OBS-style weight adjustment.
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructions
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Safe Delta: Consistently Preserving Safety when Fine-Tuning LLMs on Diverse Datasets
Safe Delta preserves LLM safety after fine-tuning by pruning delta parameters with a utility-per-safety ratio and compensating the safety loss with an OBS-style weight adjustment.