Pith. sign in

REVIEW 2 cited by

Revisiting Character-level Adversarial Attacks for Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.04346 v2 pith:PKLQ3ZR2 submitted 2024-05-07 cs.LG cs.AIcs.CLstat.ML

classification cs.LGcs.AIcs.CLstat.ML
keywords adversarialattackscharmerattackbertcharacter-leveleasilyexamples
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While character-level attacks easily maintain semantics, they have received less attention as they cannot easily adopt popular gradient-based methods, and are thought to be easy to defend. Challenging these beliefs, we introduce Charmer, an efficient query-based adversarial attack capable of achieving high attack success rate (ASR) while generating highly similar adversarial examples. Our method successfully targets both small (BERT) and large (Llama 2) models. Specifically, on BERT with SST-2, Charmer improves the ASR in 4.84% points and the USE similarity in 8% points with respect to the previous art. Our implementation is available in https://github.com/LIONS-EPFL/Charmer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adversarial Robustness through Dynamic Ensemble Learning

    cs.CR 2024-12 reject novelty 4.0 of 10

    A dynamic ensemble of BERT, RoBERTa, and ALBERT with randomized smoothing, masked inference, and TextFooler adversarial training is reported to keep 82-87% accuracy under attack on AG News and IMDB, far above prior defenses.

  2. Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks

    cs.CL 2025-01

Pith tools