Pith. sign in

REVIEW 1 cited by

Adversarial Texts with Gradient Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.07175 v2 pith:XFUQNIUM submitted 2018-01-22 cs.CL cs.CRcs.LG

classification cs.CLcs.CRcs.LG
keywords adversarialtextsmethodsgradientattackingframeworkqualitydifficult
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial samples for images have been extensively studied in the literature. Among many of the attacking methods, gradient-based methods are both effective and easy to compute. In this work, we propose a framework to adapt the gradient attacking methods on images to text domain. The main difficulties for generating adversarial texts with gradient methods are i) the input space is discrete, which makes it difficult to accumulate small noise directly in the inputs, and ii) the measurement of the quality of the adversarial texts is difficult. We tackle the first problem by searching for adversarials in the embedding space and then reconstruct the adversarial texts via nearest neighbor search. For the latter problem, we employ the Word Mover's Distance (WMD) to quantify the quality of adversarial texts. Through extensive experiments on three datasets, IMDB movie reviews, Reuters-2 and Reuters-5 newswires, we show that our framework can leverage gradient attacking methods to generate very high-quality adversarial texts that are only a few words different from the original texts. There are many cases where we can change one word to alter the label of the whole piece of text. We successfully incorporate FGM and DeepFool into our framework. In addition, we empirically show that WMD is closely related to the quality of adversarial texts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WAFBOOSTER: Automatic Boosting of WAF Security Against Mutated Malicious Payloads

    cs.CR 2025-01 reject novelty 6.0 of 10

    WAFBOOSTER combines a shadow model, an RNN payload generator, and automatic signature extraction to harden web application firewalls, but its headline rejection-rate improvement is measured on the same payloads used t...

Pith tools