Pith. sign in

REVIEW 3 cited by

Deep Text Classification Can be Fooled

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1704.08006 v2 pith:OECDCGEG submitted 2017-04-26 cs.CR cs.LG

classification cs.CRcs.LG
keywords adversarialsamplestextattackclassificationclassifiersdnn-basedimportant
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we present an effective method to craft text adversarial samples, revealing one important yet underestimated fact that DNN-based text classifiers are also prone to adversarial sample attack. Specifically, confronted with different adversarial scenarios, the text items that are important for classification are identified by computing the cost gradients of the input (white-box attack) or generating a series of occluded test samples (black-box attack). Based on these items, we design three perturbation strategies, namely insertion, modification, and removal, to generate adversarial samples. The experiment results show that the adversarial samples generated by our method can successfully fool both state-of-the-art character-level and word-level DNN-based text classifiers. The adversarial samples can be perturbed to any desirable classes without compromising their utilities. At the same time, the introduced perturbation is difficult to be perceived.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    GAversary, a black-box genetic algorithm with GloVe-based mutations, generates adversarial examples that reduce NLP model accuracy more than BAE or A2T on benchmarks while perturbing more words.

  2. Natural Adversarial Sentence Generation with Gradient-based Perturbation

    cs.IR 2019-09 conditional novelty 6.0 of 10

    Gradient-based perturbation of sentence embeddings, followed by a trained decoder, produces adversarial sentences that fool sentiment classifiers and transfer to Amazon Comprehend.

  3. Cross-Entropy Attacks to Language Models via Rare Event Simulation

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A cross-entropy optimization attack (CEA) with sememe- and MLM-based candidate words improves black-box adversarial attacks on classifiers and machine translation models.

Pith tools