Pith. sign in

Do unlearning methods remove information from language model weights?arXiv preprint arXiv:2410.08827

10 Pith papers cite this work. Polarity classification is still indexing.

10 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 6 cs.CL 4

years

2026 8 2025 2

roles

background 1

polarities

background 1

representative citing papers

Is your algorithm unlearning or untraining?

cs.LG · 2026-04-09 · conditional · novelty 7.0

Machine unlearning conflates reversing the influence of specific training examples (untraining) with removing the full underlying distribution or behavior (unlearning).

Improving LLM Unlearning Robustness via Random Perturbations

cs.CL · 2025-01-31 · unverdicted · novelty 7.0

LLM unlearning is reframed as inadvertently installing backdoor triggers on forget-tokens; Random Noise Augmentation is introduced as a defense that improves robustness with theoretical guarantees.

Safe-RULE: Safe Reinforcement UnLEarning

cs.LG · 2026-06-08 · unverdicted · novelty 5.0

Safe-RULE introduces a reinforcement unlearning defense for offline safe RL that counters data poisoning by removing malicious data influence while preserving task performance and safety.

citing papers explorer

Showing 10 of 10 citing papers.