Safety guardrails lasted longer when the alignment data was dissimilar to the downstream fine-tuning data across Llama-2 and Gemma-2 models.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
Safety guardrails lasted longer when the alignment data was dissimilar to the downstream fine-tuning data across Llama-2 and Gemma-2 models.