Pith. sign in

REVIEW 2 cited by

Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1801.04354 v5 pith:ZRRXJTT2 submitted 2018-01-13 cs.CL cs.CRcs.IRcs.LG

classification cs.CLcs.CRcs.IRcs.LG
keywords textblack-boxdeepwordbugadversarialattacksaveragebeenclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although various techniques have been proposed to generate adversarial samples for white-box attacks on text, little attention has been paid to black-box attacks, which are more realistic scenarios. In this paper, we present a novel algorithm, DeepWordBug, to effectively generate small text perturbations in a black-box setting that forces a deep-learning classifier to misclassify a text input. We employ novel scoring strategies to identify the critical tokens that, if modified, cause the classifier to make an incorrect prediction. Simple character-level transformations are applied to the highest-ranked tokens in order to minimize the edit distance of the perturbation, yet change the original classification. We evaluated DeepWordBug on eight real-world text datasets, including text classification, sentiment analysis, and spam detection. We compare the result of DeepWordBug with two baselines: Random (Black-box) and Gradient (White-box). Our experimental results indicate that DeepWordBug reduces the prediction accuracy of current state-of-the-art deep-learning models, including a decrease of 68\% on average for a Word-LSTM model and 48\% on average for a Char-CNN model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WAFBOOSTER: Automatic Boosting of WAF Security Against Mutated Malicious Payloads

    cs.CR 2025-01 reject novelty 6.0 of 10

    WAFBOOSTER combines a shadow model, an RNN payload generator, and automatic signature extraction to harden web application firewalls, but its headline rejection-rate improvement is measured on the same payloads used t...

  2. GenAI Content Detection Task 3: Cross-Domain Machine-Generated Text Detection Challenge

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Top detectors in the shared task achieved above 99% true positive rate at 5% false positive rate on the RAID benchmark when all domains and models were seen during training.

Pith tools