Fine-tuning an LLM on 100 benign samples with the highest normalized self-influence scores breaks its safety alignment, matching harmful fine-tuning.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
Fine-tuning an LLM on 100 benign samples with the highest normalized self-influence scores breaks its safety alignment, matching harmful fine-tuning.