Pith. sign in

REVIEW 2 cited by

TransFool: An Adversarial Attack against Neural Machine Translation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.00944 v2 pith:QB6WI4SG submitted 2023-02-02 cs.CL

classification cs.CL
keywords transfooladversarialmodelstranslationattackattacksneuralsemantic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks have been shown to be vulnerable to small perturbations of their inputs, known as adversarial attacks. In this paper, we investigate the vulnerability of Neural Machine Translation (NMT) models to adversarial attacks and propose a new attack algorithm called TransFool. To fool NMT models, TransFool builds on a multi-term optimization problem and a gradient projection step. By integrating the embedding representation of a language model, we generate fluent adversarial examples in the source language that maintain a high level of semantic similarity with the clean samples. Experimental results demonstrate that, for different translation tasks and NMT architectures, our white-box attack can severely degrade the translation quality while the semantic similarity between the original and the adversarial sentences stays high. Moreover, we show that TransFool is transferable to unknown target models. Finally, based on automatic and human evaluations, TransFool leads to improvement in terms of success rate, semantic similarity, and fluency compared to the existing attacks both in white-box and black-box settings. Thus, TransFool permits us to better characterize the vulnerability of NMT models and outlines the necessity to design strong defense mechanisms and more robust NMT systems for real-life applications.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries

    cs.CR 2025-08 conditional novelty 6.0 of 10

    CEMA converts multi-task black-box text attacks into attacks on a binary classifier trained on cluster pseudo-labels, achieving high attack success with 100 queries.

  2. destroR: A Benchmark and Adversarial-Training Defense for Bangla Transfer Models under Meaning-Preserving Attacks

    cs.CL 2025-11 reject novelty 4.0 of 10

    Claims a Bangla adversarial-attack benchmark and defense, but the text delivers a single-model attack study with no defense and no baselines.

Pith tools