TARS locates feedforward weights most aligned with a model-derived concept vector and replaces them with its reversed form, removing concepts like 'Sherlock Holmes' with a few edits while preserving general model behavior.
Towards safer large language models through machine unlearning
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models
TARS locates feedforward weights most aligned with a model-derived concept vector and replaces them with its reversed form, removing concepts like 'Sherlock Holmes' with a few edits while preserving general model behavior.