LARF ranks fine-tuning samples by how close their hidden representations lie to unsafe versus safe reference responses, and removing the top-ranked samples preserves safety alignment.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment
LARF ranks fine-tuning samples by how close their hidden representations lie to unsafe versus safe reference responses, and removing the top-ranked samples preserves safety alignment.