SAGE is a source-agnostic post-hoc correction for LLM unlearning updates that suppresses components aligned with high-energy retained activation directions while preserving the forgetting carrier.
Alternate preference optimization for unlearning factual knowledge in large language models.arXiv preprint arXiv:2409.13474
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2representative citing papers
citing papers explorer
-
SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector
SAGE is a source-agnostic post-hoc correction for LLM unlearning updates that suppresses components aligned with high-energy retained activation directions while preserving the forgetting carrier.
- OFMU: Optimization-Driven Framework for Machine Unlearning