Pith. sign in

REVIEW 10 cited by

Are we making progress in unlearning? Findings from the first NeurIPS unlearning competition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09073 v1 pith:NYK736RV submitted 2024-06-13 cs.LG

classification cs.LG
keywords evaluationunlearningcompetitionalgorithmsanalyzedifferentfindingsframework
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present the findings of the first NeurIPS competition on unlearning, which sought to stimulate the development of novel algorithms and initiate discussions on formal and robust evaluation methodologies. The competition was highly successful: nearly 1,200 teams from across the world participated, and a wealth of novel, imaginative solutions with different characteristics were contributed. In this paper, we analyze top solutions and delve into discussions on benchmarking unlearning, which itself is a research problem. The evaluation methodology we developed for the competition measures forgetting quality according to a formal notion of unlearning, while incorporating model utility for a holistic evaluation. We analyze the effectiveness of different instantiations of this evaluation framework vis-a-vis the associated compute cost, and discuss implications for standardizing evaluation. We find that the ranking of leading methods remains stable under several variations of this framework, pointing to avenues for reducing the cost of evaluation. Overall, our findings indicate progress in unlearning, with top-performing competition entries surpassing existing algorithms under our evaluation framework. We analyze trade-offs made by different algorithms and strengths or weaknesses in terms of generalizability to new datasets, paving the way for advancing both benchmarking and algorithm development in this important area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning

    cs.LG 2025-10 conditional novelty 6.0 of 10

    DiPO is a distribution-level unlearning method that constructs preference distributions from the model's own high-confidence logits and achieves state-of-the-art forget quality on TOFU while preserving utility.

  2. Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design

    cs.CR 2025-08 conditional novelty 6.0 of 10

    Water4MU tunes an invisible watermark on data so that machine unlearning algorithms can remove requested images more effectively, beating prior methods on 'challenging forgets'.

  3. Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective

    cs.CR 2025-06 conditional novelty 6.0 of 10

    RULI is a per-sample, dual-objective inference attack that measures privacy leakage and unlearning efficacy, showing average-case evaluations understate privacy risk.

  4. Certified Unlearning for Neural Networks

    cs.LG 2025-06 reject novelty 6.0 of 10

    Noisy fine-tuning with gradient or model clipping on retained data provably removes the influence of forget data, with guarantees that need no smoothness or convexity assumptions.

  5. Improved Localized Machine Unlearning Through the Lens of Memorization

    cs.LG 2024-12 conditional novelty 6.0 of 10

    DEL, which pairs channel-level weighted-gradient localization with reset-and-finetune, reports state-of-the-art unlearning metrics on CIFAR-10, SVHN, and ImageNet-100.

  6. Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.

  7. When unlearning is free: leveraging low influence points to reduce computational costs

    cs.LG 2025-12 conditional novelty 5.0 of 10

    Low-influence training points can be dropped from forget/retain sets before unlearning, cutting runtime up to ~50% with little measured loss in accuracy or MIA-based privacy.

  8. Verifying Machine Unlearning with Explainable AI

    cs.LG 2024-11 conditional novelty 5.0 of 10

    The authors use SIDU heatmaps and two new metrics to test whether machine unlearning removes reliance on human patterns in a thermal object-counting model.

  9. On the importance of multiple training seeds for evaluating machine unlearning

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Machine-unlearning evaluation with a single training seed can misrepresent method performance, particularly for deterministic unlearning methods, and extra unlearning seeds do not fix it.

  10. Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A position paper urging a shift from data-tracing to knowledge-tracing machine unlearning for foundation models, supported by a CLIP case study that shows current methods struggle to generalize.

Pith tools