Pith. sign in

REVIEW 2 cited by

Potion: Towards Poison Unlearning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09173 v3 pith:HIGFH52K submitted 2024-06-13 cs.LG

classification cs.LG
keywords poisonunlearningmodeldatamethodonlychallengefull
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial attacks by malicious actors on machine learning systems, such as introducing poison triggers into training datasets, pose significant risks. The challenge in resolving such an attack arises in practice when only a subset of the poisoned data can be identified. This necessitates the development of methods to remove, i.e. unlearn, poison triggers from already trained models with only a subset of the poison data available. The requirements for this task significantly deviate from privacy-focused unlearning where all of the data to be forgotten by the model is known. Previous work has shown that the undiscovered poisoned samples lead to a failure of established unlearning methods, with only one method, Selective Synaptic Dampening (SSD), showing limited success. Even full retraining, after the removal of the identified poison, cannot address this challenge as the undiscovered poison samples lead to a reintroduction of the poison trigger in the model. Our work addresses two key challenges to advance the state of the art in poison unlearning. First, we introduce a novel outlier-resistant method, based on SSD, that significantly improves model protection and unlearning performance. Second, we introduce Poison Trigger Neutralisation (PTN) search, a fast, parallelisable, hyperparameter search that utilises the characteristic "unlearning versus model protection" trade-off to find suitable hyperparameters in settings where the forget set size is unknown and the retain set is contaminated. We benchmark our contributions using ResNet-9 on CIFAR10 and WideResNet-28x10 on CIFAR100. Experimental results show that our method heals 93.72% of poison compared to SSD with 83.41% and full retraining with 40.68%. We achieve this while also lowering the average model accuracy drop caused by unlearning from 5.68% (SSD) to 1.41% (ours).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Corrective unlearning for graph neural networks is achieved by alternating contrastive separation of affected neighborhoods with asymmetric gradient ascent and descent, using as little as 5 percent of the manipulated set.

  2. Learning to Forget using Hypernetworks

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A diffusion-based hypernetwork can generate classifier weights with near-zero accuracy on a requested forget class and near-retrained accuracy on retained classes.

Pith tools