Pith. sign in

REVIEW 7 cited by

Does Unlearning Truly Unlearn? A Black Box Evaluation of LLM Unlearning Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12103 v3 pith:GI4JY6EJ submitted 2024-11-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords unlearningmethodsperformancebenchmarksdatasetevaluationgeneralimpact
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language model unlearning aims to remove harmful information that LLMs have learnt to prevent their use for malicious purposes. LLMU and RMU have been proposed as two methods for LLM unlearning, achieving impressive results on unlearning benchmarks. We study in detail the impact of unlearning on LLM performance metrics using the WMDP dataset as well as a new biology dataset we create. We show that unlearning has a notable impact on general model capabilities, with the performance degradation being more significant in general for LLMU. We further test the robustness of the two methods and find that doing 5-shot prompting or rephrasing the question in simple ways can lead to an over ten-fold increase in accuracy on unlearning benchmarks. Finally, we show that training on unrelated data can almost completely recover pre-unlearning performance, demonstrating that these methods fail at truly unlearning. Our methodology serves as an evaluation framework for LLM unlearning methods. The code is available at: https://github.com/JaiDoshi/Knowledge-Erasure.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Matching a retrained oracle on trained probes can certify models that still retain held-out forget knowledge, and oracle-free unlearning certification is only possible for counterfactual, non-inferable facts.

  2. Stochastic Meta-Unlearning: Bridging Language Backbone and Multimodal Unlearning

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A bilevel meta-training method aligns text-only backbone unlearning with full vision-language model forgetting, improving forget-retain trade-offs across two VLMs and two meme datasets.

  3. PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Existing MLLM unlearning methods reduce private-attribute leakage on entangled images but substantially harm co-occurring public figures and landmarks, with private knowledge often re-emerging after public finetuning.

  4. An Example Safety Case for Safeguards Against Misuse

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A proposed framework, built around an 'uplift model' that translates red-team safeguard-evasion data into estimated risk, for justifying that AI misuse safeguards keep large-scale harm risk below a threshold.

  5. DUSK: Do Not Unlearn Shared Knowledge

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DUSK benchmarks machine unlearning under overlapping forget and retain documents, showing existing methods remove surface text but fail to preserve shared knowledge while erasing unique content.

  6. Module-Aware Parameter-Efficient Machine Unlearning on Transformers

    cs.LG 2025-08 conditional novelty 5.0 of 10

    MAPE-Unlearn uses Fisher-information-based scores and greedy search to select important heads and filters, then applies sparse unlearning updates, claiming improved efficacy-fidelity trade-offs on Transformers.

  7. Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

    cs.CR 2025-06 conditional novelty 5.0 of 10

    Prepending a Hindi filler paragraph to WMDP-bio questions restores 57.3% accuracy in ELM-unlearned models, showing the unlearning is superficial output suppression rather than true knowledge removal.

Pith tools