REVIEW 2 cited by
Greedy Attack and Gumbel Attack: Generating Adversarial Examples for Discrete Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a probabilistic framework for studying adversarial attacks on discrete data. Based on this framework, we derive a perturbation-based method, Greedy Attack, and a scalable learning-based method, Gumbel Attack, that illustrate various tradeoffs in the design of attacks. We demonstrate the effectiveness of these methods using both quantitative metrics and human evaluation on various state-of-the-art models for text classification, including a word-based CNN, a character-based CNN and an LSTM. As as example of our results, we show that the accuracy of character-based convolutional networks drops to the level of random selection by modifying only five characters through Greedy Attack.
Forward citations
Cited by 2 Pith papers
-
Natural Adversarial Sentence Generation with Gradient-based Perturbation
Gradient-based perturbation of sentence embeddings, followed by a trained decoder, produces adversarial sentences that fool sentiment classifiers and transfer to Amazon Comprehend.
-
Robustness to Modification with Shared Words in Paraphrase Identification
Paraphrase identification models suffer dramatic accuracy drops on examples modified by replacing or adding shared words, and adversarial training partially restores robustness.
Discussion (0). Continue with ORCID to comment.