DART, a single-step embedding-space perturber trained with reinforcement learning, finds toxic prompts closer to reference prompts than fine-tuned or few-shot baselines on three LLMs.
D.; Ho, J.; Tarlow, D.; and Van Den Berg, R
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
DART, a single-step embedding-space perturber trained with reinforcement learning, finds toxic prompts closer to reference prompts than fine-tuned or few-shot baselines on three LLMs.