REVIEW 1 cited by
Influence-based Attributions can be Manipulated
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Influence Functions are a standard tool for attributing predictions to training data in a principled manner and are widely used in applications such as data valuation and fairness. In this work, we present realistic incentives to manipulate influence-based attributions and investigate whether these attributions can be \textit{systematically} tampered by an adversary. We show that this is indeed possible for logistic regression models trained on ResNet feature embeddings and standard tabular fairness datasets and provide efficient attacks with backward-friendly implementations. Our work raises questions on the reliability of influence-based attributions in adversarial circumstances. Code is available at : \url{https://github.com/infinite-pursuits/influence-based-attributions-can-be-manipulated}
Forward citations
Cited by 1 Pith paper
-
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
ExpProof is the first implemented protocol that lets a holder of a confidential ML model prove to a customer that a LIME explanation was computed from the committed model, with measured proof generation around 1.5 min...
Discussion (0). Continue with ORCID to comment.