Morality-specific jailbreak attacks expose critical vulnerabilities in both large language models and guardrail systems when handling pluralistic values.
It promotes false information, harmful behaviors, or negative sentiments that could have a serious impact
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
citing papers explorer
-
Jailbreaking Large Language Models with Morality Attacks
Morality-specific jailbreak attacks expose critical vulnerabilities in both large language models and guardrail systems when handling pluralistic values.
- Operator-Valued Hardy Spaces and Kramers--Kronig Relations for Non-Markovian Quantum Memory Kernels