{"id":"4be06298-9aed-42bb-b346-230c96d0db31","arxiv_id":"2506.06391","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.","lead":"This paper tests eight large language models on 322 prompts asking for actions that violate international humanitarian law, measuring how often the models refuse and whether they explain why. For most models, adding a short safety instruction to the system prompt sharply increased the share of explanatory refusals without retraining.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'helpfulness' metric is partly defined by the behavior the intervention instructs: §3.3 counts legal references as helpful, and §3.4's prompt asks for exactly that, so the headline gains may measure prompt-compliance, not explanation quality; unvalidated LLM judges make this uninterpretable.","rationale":"The reader's conditional verdict is appropriate, but the weakest assumption is slightly broader than LLM-judge validation. The helpfulness rubric itself is confounded with the intervention: the prompt tells the model to cite IHL/IHRL, and the metric counts legal citations as helpful. Even a perfect human judge using the paper's stated criteria would likely reproduce part of the effect because the outcome definition is not independent of the treatment. This is a load-bearing measurement-validity problem, not merely a concern about judge identity. A second, independent issue is the denominator: the reported metric is P(explanatory refusal) over all prompts, not P(explanatory | refusal), so changes in refusal rate can masquerade as changes in explanation quality. The paper never reports post-intervention refusal rates, so the decomposition cannot be checked from the manuscript alone. The qualitative examples and the broad pattern of baseline refusal rates are plausible and worth preserving, so the paper should not be rejected outright. Conditional acceptance with mandatory release of the benchmark, judge definitions, human validation, and a conditional-on-refusal re-analysis is the right disposition. The reader's conditional verdict already captures this, so no change in verdict is recommended.","tokens_in":17344,"tokens_out":9047,"duration_ms":103971,"concrete_test":"Release the judge prompts, judge model identities, and all baseline/intervention outputs; have three human annotators, blind to condition, rate a random sample of 100 baseline and 100 intervention responses with a rubric that scores explanation quality separately from mere legal citation; compute human–LLM agreement and conditional-on-refusal explanatory rates (explanatory refusals / refusals) before and after. If the human-rated conditional improvement is large for most models, the central claim survives; if not, the headline effect is a measurement artifact.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a system prompt improves the quality of explanations within refusals. The evidence in §4.4 is not a clean measure of that. In §3.3, a refusal is deemed 'helpful' if it includes 'reference to relevant legal standards'; in §3.4, the intervention explicitly instructs models to 'mention that the following actions are illegal according to IHL or Human Rights.' The outcome therefore rewards the exact behavior the prompt requests. For a model like claude-3.5-sonnet, the jump from 24.53% to 98.45% may largely reflect the model parroting the instructed legal citation, not a general improvement in explanatory quality. The three LLM judges are unnamed, with no human validation or inter-rater agreement, so the reported rates could be artifacts of judges that equate legalistic phrasing with helpfulness. A second, compounding problem: 'IHL-Compliant Refusal Helpfulness' (§3.4) is the percentage of all prompts resulting in an explanatory refusal, not the percentage of refusals that are explanatory; because post-intervention refusal rates are not reported, part of the gain could be an increase in refusal coverage rather than explanation quality. The paper's own qualitative examples show helpful refusals exist, but the headline quantitative claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces a benchmark of 322 prompts derived from the ICRC customary IHL rules and evaluates eight LLMs on refusal rates and 'helpfulness' of refusals, using majority voting of three undisclosed LLM judges. It then applies a standard system prompt that instructs models to mention that actions are illegal under IHL or human rights and reports that six of eight models show large increases in 'IHL-compliant refusal helpfulness' (Table 1), e.g., claude-3.5-sonnet from 24.53% to 98.45%. The paper concludes that lightweight deployment-time prompting can significantly improve refusal quality without retraining.","tokens_in":17602,"tokens_out":4363,"duration_ms":45278,"significance":"If the central result were established, the paper would make a useful practical contribution: a simple system prompt that improves refusal transparency, plus a benchmark grounded in an external legal standard rather than an ad-hoc safety taxonomy. The traceability of prompts to specific ICRC rules (Section 3.1) is a genuine strength, and the qualitative examples in Section 4.3 are informative. However, the quantitative claims depend entirely on an unvalidated LLM-judge pipeline, and the outcome metric partially encodes the intervention itself, so the headline improvements are not yet interpretable. The paper also contains a significant unsupported claim about benign-prompt response rates in Section 5.","major_comments":[{"comment":"The 'helpfulness' metric is partly defined by the behavior the intervention instructs. In Section 3.3, a refusal is deemed helpful if it includes 'reference to relevant legal standards.' In Section 3.4, the intervention prompt explicitly instructs models to 'mention that the following actions are illegal according to the IHL or Human Rights.' The outcome therefore rewards the exact behavior the prompt requests. For claude-3.5-sonnet, the jump from 24.53% to 98.45% (Table 1) may largely reflect the model parroting the instructed legal citation rather than a general improvement in explanatory quality. The paper should either use a judge rubric that is blind to the intervention content or report separate sub-scores for legal citation versus genuine reasoning, and discuss what the metric actually measures.","section":"Sections 3.3-3.4"},{"comment":"The evaluation uses majority voting among three LLM-based evaluators for both refusal classification and helpfulness classification, but the evaluators are never identified, no inter-rater agreement is reported, no human-validated gold sample is described, and the 'manual spot-checks' mentioned in Section 3.4 are not quantified. Because every number in Table 1 depends on these judges, the absence of validation makes the headline rates uninterpretable. The authors should release the judge identities (or at least model versions), the full evaluation prompts, a human-annotated validation subset, and agreement statistics such as Cohen's kappa.","section":"Section 3.3"},{"comment":"The metric 'IHL-Compliant Refusal Helpfulness' is defined as the percentage of IHL-violating prompts that resulted in explanatory refusals, not the percentage of refusals that are explanatory. Because post-intervention refusal rates are not reported, the increases in Table 1 could partly reflect improved refusal coverage rather than improved explanation quality. For example, mistral-large had a baseline refusal rate of 88.82%; if the intervention also reduces non-refusal compliance, the reported helpfulness of 93.17% would overstate the improvement in explanation quality. The claim in Section 4.4 that the intervention 'corrected prior issues related to the models responding to harmful prompts' requires a separate reporting of refusal rates under the intervention.","section":"Section 3.4 and Table 1"},{"comment":"The Discussion states that the system prompt 'improved the model's response rate to benign prompts from 65.53% to 94.41%' for Claude 3.5 Sonnet. No benign-prompt evaluation appears in the methodology (Section 3) or in the results tables, and the numbers are not otherwise derivable from the reported data. Either the benign-prompt experiment must be fully described and its results reported, or this passage should be deleted.","section":"Section 5"}],"minor_comments":[{"comment":"There is a typo: 'explicitly referenced actions prohibited the IHL and IHRL' should read 'prohibited by IHL and IHRL.'","section":"Section 3.4"},{"comment":"Figure 1 shows a baseline helpfulness of 74.84% for qwen-2.5-72b-instruct, but Table 1 reports 74.12%. These values should be reconciled.","section":"Figure 1"},{"comment":"The paper does not release the 322 prompts or the model outputs. For a proposed benchmark, releasing the prompt set and a sample of outputs, even in an appendix or supplementary material, would substantially aid reproducibility and external validation.","section":"Reproducibility"},{"comment":"The paper gives model names but no exact API versions or access dates (e.g., 'chatgpt-o3-mini' is ambiguous). Reporting the precise model snapshots is important for reproducibility given the rapid pace of model updates.","section":"Section 3.2"},{"comment":"No confidence intervals or significance tests are provided for the headline rates. With 322 prompts and majority voting, the smaller reported differences (e.g., 88.20% vs. 91.93%) may not be statistically meaningful; the authors should either add uncertainty quantification or explicitly label the results as point estimates.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's central claim is plausible and the benchmark has sound external grounding in ICRC rules, but the evaluation pipeline lacks validation and the outcome metric partly encodes the intervention. The unsupported benign-prompt numbers in Section 5 are a red flag that the manuscript was not carefully checked before submission. I recommend major revision with a request for full methodological transparency: identify the judges, report agreement, provide human validation, decompose the helpfulness metric, and either substantiate or remove the benign-prompt claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the paper with the stress-test note in hand. The circularity concern is real but overstated. What's actually new: a 322-prompt benchmark keyed to the ICRC customary IHL rules, and a deployment-time system prompt that, in six of eight models, lifts the rate of explanatory refusals dramatically. The legal grounding is genuinely external and traceable to specific rules; that's the strongest part of the paper. The intervention result is plausible and consistent with prior refusal-benchmark work.\n\nThe soft spots are measurement. Helpfulness is defined in part as 'reference to relevant legal standards,' and the intervention prompt explicitly asks for exactly that, so part of the gain may be prompt-compliance rather than improved explanation. The three LLM judges are unnamed, unvalidated, and reported with no inter-rater agreement or confidence intervals. The metric definition is also internally inconsistent: §3.3 says helpfulness is the proportion of all prompts yielding an explanatory refusal, while §4.4 says it's measured only when a refusal occurs; post-intervention refusal rates are missing, so you can't decompose coverage from quality. For claude-3.5-sonnet, baseline refusal was 100%, so its 24% to 98% gain is clearly explanation quality; but for models like mistral-large, some of the gain could be coverage.\n\nThe single worst issue: the Discussion claims the system prompt improved Claude 3.5 Sonnet's response rate to benign prompts from 65.53% to 94.41%. There is no benign-prompt experiment in the paper. That claim should be retracted or verified.\n\nThis is a paper with a genuinely new application-level benchmark, but the headline numbers are not yet trustworthy. The authors should release prompts, judge definitions, and outputs; validate judges against human labels; report post-intervention refusal rates; and fix the benign-prompt claim. I'd send it to peer review with a request for major revision, not desk-reject it. Anyone working on refusal behavior or legal AI evaluation should keep an eye on this version, but I wouldn't cite the numbers until the artifacts ship.","headline":"IHL-anchored refusal benchmark with plausible intervention results, but unvalidated LLM judges and a stray unsupported claim keep the headline numbers from being fully established.","tokens_in":18139,"tokens_out":4111,"would_cite":false,"duration_ms":43789,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that a standardised system-level safety prompt can lift explanatory refusal rates from roughly a quarter to over 90 percent for six of eight LLMs, without retraining.","keywords":["International Humanitarian Law","LLM refusal behaviour","explanatory refusals","system prompt intervention","AI alignment","content moderation","customary IHL benchmark","harmful content mitigation"],"falsifier":"Re-run the same 322-prompt benchmark with the same intervention but replace the three-judge LLM vote with human annotation of 'refusal' and 'helpful refusal' on a sample of outputs; if human labels diverge from the judge vote, or the post-intervention improvement shrinks, the central claim fails. A cheaper check is to swap in a different set of LLM judges and see whether the before/after gap survives.","tokens_in":17140,"feed_emoji":"⚖️","tokens_out":7502,"duration_ms":73307,"temperature":0.7,"pith_summary":"This paper sets out to establish that explicit, explanatory refusals are a practical lever for aligning large language models with international humanitarian law, and that a standardised system-level safety prompt can activate far better refusal explanations without any retraining. The authors construct 322 prompts that clearly violate specific customary IHL rules, test eight leading instruction-tuned models on them, and measure both raw refusal rates and the rate at which refusals actually explain themselves. Although all eight models refused the great majority of prompts, their refusals were often terse: baseline explanatory refusal rates ranged from under 8% to about 80%. After adding one system-level instruction that flags IHL and human-rights prohibitions, six of the eight models explained their refusals in more than 90% of cases in several instances. The authors argue this shows latent alignment capacity that can be elicited cheaply at deployment time, while noting that technically framed or code-mixed harmful requests still slip through.","feed_headline":"Safety prompt lifts refusal explanations past 90 percent","feed_subtitle":"Six of eight chatbots went from terse denials to explained refusals on war-law prompts, no retraining.","key_machinery":"The load-bearing object is a single standardised system-level safety prompt inserted into each model interaction before the user query; it states that certain acts are illegal under IHL and human rights law and instructs the model to say so when relevant. The effect is measured by the 'IHL-compliant refusal helpfulness rate', the share of IHL-violating prompts that produced a refusal with substantive explanation. Classification is done by majority vote among three LLM-based evaluators on two binary questions: did the model refuse, and was the refusal explanatory. Comparing this metric before and after the prompt isolates the intervention's effect, and the comparison against the 322-prompt benchmark, each prompt tied to a specific customary IHL rule, is what makes the before/after numbers comparable across models.","core_discovery":"The discovery the paper argues for is that refusal quality, not just refusal presence, is the bottleneck in LLM compliance with IHL, and that this bottleneck can be loosened with a lightweight, deployment-time intervention. On the baseline evaluation, refusal rates were high across the board, from 88.82% for mistral-large to 100% for claude-3.5-sonnet, but explanatory refusal rates varied widely, and strength on one dimension did not guarantee strength on the other. The intervention, a standardised high-level system prompt referencing actions prohibited and required under IHL and international human rights law, lifted explanatory refusal rates sharply for most models: claude-3.5-sonnet rose from 24.53% to 98.45%, chatgpt-4o from 36.02% to 91.93%, mistral-large from 70.50% to 93.17%, gemini-2.0-flash from 56.21% to 88.20%, and claude-3.7-sonnet from 80.12% to 91.93%. The two exceptions, llama-3.3-70b-instruct and chatgpt-o3-mini, improved more modestly, showing that the prompt does not fully override a model's entrenched refusal style. In the paper's terms, this demonstrates that well-articulated, legally grounded refusals can be elicited from most current models without additional training.","pith_inferences":["The paper does not test whether the three-judge LLM panel would agree with human raters; if the judges reward length or legalistic phrasing, part of the measured improvement could be an evaluator artefact rather than a genuine gain in refusal quality.","The same one-prompt recipe could plausibly transfer to other codified domains, such as medical ethics or data-protection law, but that transfer is not established by the paper's data.","Because the intervention worked by activating latent behaviour, refusal quality may be more a property of decoding and orchestration than of training, which would make lightweight safety auditing of new models cheaper than the paper explicitly claims.","An extension the paper suggests but does not run is a two-stage system where a second model writes the explanation for a terse refuser; this is directly testable with the existing benchmark."],"forward_implications":["With no retraining, a standardised system prompt can move most of the eight tested models from terse denials to explanatory refusals in over 90% of IHL-violating prompts.","Explanatory refusals that cite legal or safety principles can make a model's boundaries legible to users, which the paper argues reduces ambiguity and makes refusals harder to treat as predictable strings to suppress.","The benchmark of 322 prompts mapped to customary IHL rules provides a reusable protocol for auditing LLM compliance with a codified legal framework rather than general toxicity.","Code-mixed requests that embed harmful intent in technical language or function calls remain a concrete failure mode even for models with near-perfect refusal rates on plain-language violations.","For at least one model, chatgpt-o3-mini, prompt-level intervention alone is not enough to produce explanatory refusals, suggesting a need for complementary mechanisms such as a second model that writes the explanation."],"supporting_citations":[{"why":"Compiles the 161 customary IHL rules that ground the 322-prompt benchmark and the claim that each prompt maps to a codified legal norm.","marker":"[4]"},{"why":"Shows users perceive explanatory refusals as more legitimate and trustworthy, which justifies treating helpfulness as a safety-relevant metric.","marker":"[5]"},{"why":"Describes the model-feedback alignment approach the paper contrasts with its lightweight prompt-only intervention.","marker":"[6]"},{"why":"Demonstrates that short, predictable refusal phrases can be suppressed by adversarial suffixes, motivating the call for explanatory refusals.","marker":"[7]"},{"why":"Argues that explicit reasoning before a refusal improves transparency and safety, supporting the paper's definition of helpful refusals.","marker":"[15]"},{"why":"Provides evidence that technical phrasing lowers refusal rates, used to explain the code-mixed vulnerability cases.","marker":"[18]"},{"why":"Prior work that used a system prompt to improve safety and reported over-refusal; the paper's intervention is positioned against that baseline.","marker":"[31]"}],"fun_headline_variants":["Refusal quality, not just refusals, key to LLM law compliance","System prompt lifts explanatory refusals on war-law prompts","Prompt tweak makes LLM refusals clearer on war-law questions","Lightweight prompt turns terse denials into explained refusals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline result depends on trusting the majority vote of three unnamed LLM judges to determine both whether a response is a refusal and whether the refusal is helpful, with no human validation or inter-rater agreement reported.","fun_headline_variants_meta":{"raw":{"variants":["Refusal quality, not just refusals, key to LLM law compliance","System prompt lifts explanatory refusals on war-law prompts","Prompt tweak makes LLM refusals clearer on war-law questions","Lightweight prompt turns terse denials into explained refusals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2710,"prompt_tokens":1012,"completion_tokens":1698,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":1633}},"tokens_in":628,"tokens_out":1698,"duration_ms":14270,"temperature":1.0,"reasoning_tokens":1633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:22:07.674856+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same 322-prompt benchmark with the same intervention but replace the three-judge LLM vote with human annotation of 'refusal' and 'helpful refusal' on a sample of outputs; if human labels diverge from the judge vote, or the post-intervention improvement shrinks, the central claim fails. A cheaper check is to swap in a different set of LLM judges and see whether the before/after gap survives.","supporting_citations":[{"cited_title":"Henckaerts and L","cited_arxiv_id":null,"evidence_quote":"Compiles the 161 customary IHL rules that ground the 322-prompt benchmark and the claim that each prompt maps to a codified legal norm."},{"cited_title":"Zhang, M","cited_arxiv_id":null,"evidence_quote":"Argues that explicit reasoning before a refusal improves transparency and safety, supporting the paper's definition of helpful refusals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence that technical phrasing lowers refusal rates, used to explain the code-mixed vulnerability cases."}],"review_version":1}