Pith. sign in

REVIEW 3 cited by

RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07831 v1 pith:XNTI7L47 submitted 2021-10-15 cs.CL cs.LG

classification cs.CLcs.LG
keywords backdoorrobustness-awareattacksdefensesamplesanalysiscleandefending
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Backdoor attacks, which maliciously control a well-trained model's outputs of the instances with specific triggers, are recently shown to be serious threats to the safety of reusing deep neural networks (DNNs). In this work, we propose an efficient online defense mechanism based on robustness-aware perturbations. Specifically, by analyzing the backdoor training process, we point out that there exists a big gap of robustness between poisoned and clean samples. Motivated by this observation, we construct a word-based robustness-aware perturbation to distinguish poisoned samples from clean samples to defend against the backdoor attacks on natural language processing (NLP) models. Moreover, we give a theoretical analysis about the feasibility of our robustness-aware perturbation-based defense method. Experimental results on sentiment analysis and toxic detection tasks show that our method achieves better defending performance and much lower computational costs than existing online defense methods. Our code is available at https://github.com/lancopku/RAP.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lethe: Purifying Backdoored Large Language Models with Knowledge Dilution

    cs.CL 2025-08 conditional novelty 6.0 of 10

    LETHE uses parameter-level model merging plus prompt-level word definitions to dilute backdoor behavior in LLMs, cutting attack success to below 7% in most tested settings.

  2. Your Agent Can Defend Itself against Backdoor Attacks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    A two-level consistency defense detects backdoored LLM agents by matching thoughts to actions and reconstructed instructions to the user's instruction, reducing attack success rates on tested tasks.

  3. A Systematic Review of Poisoning Attacks Against Large Language Models

    cs.CR 2025-06 conditional novelty 5.0 of 10

    A systematic review that organizes 65 LLM poisoning papers into a threat model with four attack specifications and generalized metrics.

Pith tools