Matched harmful and benign wrapper groups plus a refusal-score consistency regularizer allow fine-tuned LLMs to refuse harmful intent under surface-form variation while reducing benign over-refusal.
International Conference on Learning Representations , volume=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Refusing Intent, Not Form: Wrapper-Based Intent-Group Supervision for LLM Safety
Matched harmful and benign wrapper groups plus a refusal-score consistency regularizer allow fine-tuned LLMs to refuse harmful intent under surface-form variation while reducing benign over-refusal.