Pith. sign in

REVIEW 2 cited by

Robust Conversational Agents against Imperceptible Toxicity Triggers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.02392 v1 pith:FV7QEDY2 submitted 2022-05-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords languageattacksconversationaldefensetoxicagentsimperceptiblethey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Warning: this paper contains content that maybe offensive or upsetting. Recent research in Natural Language Processing (NLP) has advanced the development of various toxicity detection models with the intention of identifying and mitigating toxic language from existing systems. Despite the abundance of research in this area, less attention has been given to adversarial attacks that force the system to generate toxic language and the defense against them. Existing work to generate such attacks is either based on human-generated attacks which is costly and not scalable or, in case of automatic attacks, the attack vector does not conform to human-like language, which can be detected using a language model loss. In this work, we propose attacks against conversational agents that are imperceptible, i.e., they fit the conversation in terms of coherency, relevancy, and fluency, while they are effective and scalable, i.e., they can automatically trigger the system into generating toxic language. We then propose a defense mechanism against such attacks which not only mitigates the attack but also attempts to maintain the conversational flow. Through automatic and human evaluations, we show that our defense is effective at avoiding toxic language generation even against imperceptible toxicity triggers while the generated language fits the conversation in terms of coherency and relevancy. Lastly, we establish the generalizability of such a defense mechanism on language generation models beyond conversational agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning

    cs.LG 2024-12 conditional novelty 6.0 of 10

    An automated red-teaming method that uses LLM-generated per-goal rewards and multi-step RL with a style-diversity reward to produce diverse and effective attacks on language models.

  2. Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Pre-trained language models are vulnerable to phonologically and orthographically motivated character substitutions in Indic languages, but less so than to unconstrained random character substitution.

Pith tools