Reranking an LLM's top-k candidate tokens by cosine similarity to predefined negative concept embeddings reduces unsafe responses and jailbreak success without retraining, but the reported gains are partly tuned to the evaluation data.
, 4. "Containing, describing, enabling, encouraging, or endors- ing or endorse the sexual abuse of children
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
DIESEL -- Dynamic Inference-Guidance via Evasion of Semantic Embeddings in LLMs
Reranking an LLM's top-k candidate tokens by cosine similarity to predefined negative concept embeddings reduces unsafe responses and jailbreak success without retraining, but the reported gains are partly tuned to the evaluation data.