Pith. sign in

Existing Large Language Model Unlearning Evaluations Are Inconclusive

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Machine unlearning aims to remove sensitive or undesired data from large language models. However, recent studies suggest that unlearning is often shallow, claiming that removed knowledge can easily be recovered. In this work, we critically examine standard unlearning evaluation practices and uncover key limitations that shake our trust in those findings. First, we show that some evaluations introduce substantial new information into the model, potentially masking true unlearning performance by re-teaching the model during testing. Second, we demonstrate that evaluation outcomes vary significantly across tasks, undermining the generalizability of current evaluation routines. Finally, we find that many evaluations rely on spurious correlations, making their results difficult to trust and interpret. Taken together, these issues suggest that current evaluation protocols may both overstate and understate unlearning success. To address this, we propose two principles for future unlearning evaluations: minimal information injection and downstream task awareness. We validate these principles through a series of targeted experiments, showing how violations of each can lead to misleading conclusions.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs

cs.LG · 2025-07-06 · conditional · novelty 6.0

A new method, Partial Model Collapse, iteratively fine-tunes an LLM on its own self-generated responses to conditionally collapse its output distribution on forget queries, removing private answers without the true labels as training targets.

citing papers explorer

Showing 1 of 1 citing paper.

  • Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs cs.LG · 2025-07-06 · conditional · none · ref 11 · internal anchor

    A new method, Partial Model Collapse, iteratively fine-tunes an LLM on its own self-generated responses to conditionally collapse its output distribution on forget queries, removing private answers without the true labels as training targets.