Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.
hub
Title resolution pending
1 Pith paper cite this work, alongside 1,504 external citations. Polarity classification is still indexing.
1
Pith paper citing it
1,504
external citations · OpenAlex
hub tools
citation-role summary
background 1
citation-polarity summary
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.