Pith. sign in

REVIEW 2 cited by

Fooling Explanations in Text Classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.03178 v1 pith:PPTWWYQH submitted 2022-06-07 cs.LG cs.CLcs.CR

classification cs.LGcs.CLcs.CR
keywords explanationperturbationsmethodstextclassifiersarchitecturesexplanationsapplications
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

State-of-the-art text classification models are becoming increasingly reliant on deep neural networks (DNNs). Due to their black-box nature, faithful and robust explanation methods need to accompany classifiers for deployment in real-life scenarios. However, it has been shown in vision applications that explanation methods are susceptible to local, imperceptible perturbations that can significantly alter the explanations without changing the predicted classes. We show here that the existence of such perturbations extends to text classifiers as well. Specifically, we introduceTextExplanationFooler (TEF), a novel explanation attack algorithm that alters text input samples imperceptibly so that the outcome of widely-used explanation methods changes considerably while leaving classifier predictions unchanged. We evaluate the performance of the attribution robustness estimation performance in TEF on five sequence classification datasets, utilizing three DNN architectures and three transformer architectures for each dataset. TEF can significantly decrease the correlation between unchanged and perturbed input attributions, which shows that all models and explanation methods are susceptible to TEF perturbations. Moreover, we evaluate how the perturbations transfer to other model architectures and attribution methods, and show that TEF perturbations are also effective in scenarios where the target model and explanation method are unknown. Finally, we introduce a semi-universal attack that is able to compute fast, computationally light perturbations with no knowledge of the attacked classifier nor explanation method. Overall, our work shows that explanations in text classifiers are very fragile and users need to carefully address their robustness before relying on them in critical applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Robust and Accurate Stability Estimation of Local Surrogate Models in Text-based Explainable AI

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Many similarity measures used to judge adversarial text-XAI attacks are too sensitive, and synonymity-weighted variants change the measured success rates substantially.

  2. Improving Stability Estimates in Adversarial Explainable AI through Alternate Search Methods

    cs.LG 2025-01 reject novelty 5.0 of 10

    A genetic algorithm finds smaller adversarial perturbations of LIME text explanations than a greedy baseline in some settings, but the stability estimates lack statistical support.

Pith tools