Pith. sign in

hub

Who’s harry potter? approximate unlearning in llms.arXiv preprint arXiv:2310.02238

33 Pith papers cite this work, alongside 13 external citations. Polarity classification is still indexing.

33 Pith papers citing it
13 external citations · external index

hub tools

citation-role summary

background 3 method 1

citation-polarity summary

representative citing papers

Revocable Learned State via Process Sidecars

cs.LG · 2026-06-29 · unverdicted · novelty 7.0

Process sidecars use a secant-based two-parameter edit to achieve second-order accurate memory revocation after safety training, outperforming scalar task arithmetic on refusal tasks across three models.

Exact Unlearning in Reinforcement Learning

cs.LG · 2026-06-02 · unverdicted · novelty 7.0

For any ρ>0 there exists a ρ-TV-stable RL algorithm for tabular MDPs supporting exact unlearning at expected cost ρ√(ln T) of retraining from scratch, with regret O(H²√(SAT)+H³S²A+H^{2.5}S²A/ρ) and matching lower bound Ω(H√(SAT)+SAH/ρ).

Improving LLM Unlearning Robustness via Random Perturbations

cs.CL · 2025-01-31 · unverdicted · novelty 7.0

LLM unlearning is reframed as inadvertently installing backdoor triggers on forget-tokens; Random Noise Augmentation is introduced as a defense that improves robustness with theoretical guarantees.

Detecting Pretraining Data from Large Language Models

cs.CL · 2023-10-25 · conditional · novelty 7.0

Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.

Fast Unlearning at Scale via Margin Self-Correction

cs.LG · 2026-06-01 · unverdicted · novelty 6.0

MASC achieves competitive forget-retain trade-offs in language model unlearning at lower computational cost via margin self-correction and an online stopping criterion on TOFU, MUSE News, and MUSE Books.

How Hard Can It Be? Hardness-Aware Multi-Objective Unlearning

cs.LG · 2026-06-01 · unverdicted · novelty 6.0

HAMU is a constrained-optimization unlearning method that uses forget-retain data similarity as a hardness measure to guarantee specified forget-quality gains while minimizing retain degradation.

ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

cs.LG · 2026-05-16 · unverdicted · novelty 6.0 · 3 refs

ZeroUnlearn reformulates machine unlearning as knowledge re-mapping via model editing, using multiplicative updates with closed-form solutions for efficient few-shot removal of sensitive representations while preserving utility.

The Realignment Problem: When Right becomes Wrong in LLMs

cs.CL · 2025-11-04 · unverdicted · novelty 6.0

TRACE is a three-stage optimization framework that realigns LLMs to new policies by categorizing preference conflicts, scoring impact via bi-level optimization, and applying hybrid losses without new human annotations.

citing papers explorer

Showing 33 of 33 citing papers.