Pith. sign in

REVIEW 3 cited by

Word-level Textual Adversarial Attacking as Combinatorial Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.12196 v4 pith:CERSS52I submitted 2019-10-27 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords adversarialattackattackingmodelmethodsoptimizationtextualword-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial attacks are carried out to reveal the vulnerability of deep neural networks. Textual adversarial attacking is challenging because text is discrete and a small perturbation can bring significant change to the original input. Word-level attacking, which can be regarded as a combinatorial optimization problem, is a well-studied class of textual attack methods. However, existing word-level attack models are far from perfect, largely because unsuitable search space reduction methods and inefficient optimization algorithms are employed. In this paper, we propose a novel attack model, which incorporates the sememe-based word substitution method and particle swarm optimization-based search algorithm to solve the two problems separately. We conduct exhaustive experiments to evaluate our attack model by attacking BiLSTM and BERT on three benchmark datasets. Experimental results demonstrate that our model consistently achieves much higher attack success rates and crafts more high-quality adversarial examples as compared to baseline methods. Also, further experiments show our model has higher transferability and can bring more robustness enhancement to victim models by adversarial training. All the code and data of this paper can be obtained on https://github.com/thunlp/SememePSO-Attack.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SurvAttack: Black-Box Attack On Survival Models through Ontology-Informed EHR Perturbation

    cs.LG 2024-12 conditional novelty 6.0 of 10

    SurvAttack uses ontology-guided additions, removals, and replacements of medical codes to flip the survival-time rankings predicted by black-box survival models.

  2. Towards Action Hijacking of Large Language Model-based Agent

    cs.CR 2024-12 conditional novelty 6.0 of 10

    A RAG-based LLM application can be induced to assemble harmful SQL, code, or medical action plans from knowledge already stored in its database, with the user prompt itself carrying no forbidden words.

  3. Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Pre-trained language models are vulnerable to phonologically and orthographically motivated character substitutions in Indic languages, but less so than to unconstrained random character substitution.

Pith tools