Merging a safety LoRA adapter with a task adapter via weighted fusion reduces the harmfulness rate from 44.2% to 2.0% on the HEx-PHI benchmark, at the cost of increased over-refusal.
In: In- ternational Conference on Learning Representations (ICLR) (2022)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Enhancing AI Safety Through the Fusion of Low Rank Adapters
Merging a safety LoRA adapter with a task adapter via weighted fusion reduces the harmfulness rate from 44.2% to 2.0% on the HEx-PHI benchmark, at the cost of increased over-refusal.