{"id":"2c18a38f-d912-423f-acc2-ae02de464a6d","arxiv_id":"2503.17365","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Self-critique alignment sharply cuts harmful outputs on Llama-based 8B models but barely helps Gemma-2-9B and Qwen2.5-7B, indicating architecture-dependent effectiveness.","lead":"This paper tests a do-it-yourself safety trick, Constitutional AI self-critique, on four small uncensored language models and finds that the two Llama-based models reduce harmful replies far more than Gemma-2 or Qwen2.5. The result suggests that this cheap alignment method works very unevenly across model families, which matters for anyone trying to secure small open models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central comparison relies on a 'harmful response rate' that is never operationally defined: no judge, rubric, classifier, or annotation protocol is described, so the architecture-level conclusion is currently unverifiable.","rationale":"The paper is a short empirical study; read in good faith, it is trying to show that the benefit of CAI's self-critique step differs across small open-weight models. For that claim to hold, the dependent variable in Figure 1 must be measured by a defined, consistent procedure. That condition is not met: Section 2 describes only response generation, and Appendix B describes only capability benchmarks; no evaluation of harmfulness for Figure 1 is specified. This is the most load-bearing weakness because every headline number depends on it, and the architecture split could be an artifact of an unspecified judge. The Gemma system-prompt difference (Appendix A) is a real secondary confound, but it is acknowledged and the authors report a negative control; the missing judge is not acknowledged anywhere. The proposed test—releasing the responses and an explicit evaluation protocol, then recomputing with an independent judge—would settle whether the concern lands. Because the reader already flagged this as the core issue and assigned CONDITIONAL, I do not change the verdict: the paper is acceptable only if the evaluation methodology and data are supplied and the results survive re-measurement. No additional objection is needed.","tokens_in":5157,"tokens_out":4625,"duration_ms":46591,"concrete_test":"Release the 90 initial/critique/revised response triples for all four models, together with the exact evaluation protocol: if an LLM judge was used, specify model, version, prompt, temperature, and sampling parameters; if a classifier, name it; if human, provide annotation instructions and inter-annotator agreement. Then recompute Figure 1 with (a) the stated protocol and (b) an independent judge such as GPT-4o or the official HarmBench classifier, on responses anonymized to remove model identity. If the revised-response harm rates change by more than a few percentage points for any model, or if the model ordering changes, the architecture-dependent conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that CAI self-critique reduces harmful responses strongly for Llama-based models but not for Gemma-2 and Qwen-2.5—rests entirely on the 'Harmful Responses (%)' numbers in Figure 1, yet the manuscript never defines how those percentages were obtained. Section 2 describes response generation and the 90 HarmBench prompts, and Appendix B details capability benchmarks (via lm-evaluation-harness), but nowhere is there a judge, a classifier, a rubric, human annotation instructions, or an agreement measure for Figure 1. Without an operational definition of 'harmful', the headline numbers (54.4→11.1, 87.8→38.9, 45.6→40.0, 76.7→67.8) are not reproducible, and the between-architecture comparison is unfalsifiable. If the judge is itself an LLM (especially a Llama-family model or one that keys on refusal-style wording), the observed architecture split could be an artifact of the judge's own biases rather than a property of CAI. The secondary system-prompt difference for Gemma-2 (Appendix A) compounds this, since that model received a different prompting condition, but the missing evaluation protocol is the more fundamental problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical comparison of Constitutional AI (CAI) self-critique on four small instruction-tuned LLMs (DeepSeek-R1-Distill-Llama-8B, Gemma-2-9B-it, Llama-3.1-8B-Instruct, and Qwen2.5-7B-Instruct), all abliterated to remove refusal behavior. For 90 HarmBench prompts across six categories, the authors measured the 'harmful response rate' of initial and CAI-revised responses, reporting that Llama-based models improve substantially (R1-Llama: 54.4% to 11.1%; Llama-3.1: 87.8% to 38.9%) while Gemma-2 and Qwen2.5 improve only slightly (45.6% to 40.0%; 76.7% to 67.8%). They also report benchmark scores (MMLU, tinyBenchmarks, ETHICS subset, SafetyBench) for abliterated versus original models. The paper concludes that CAI's effectiveness varies by model architecture and reasoning capability.","tokens_in":5508,"tokens_out":4723,"duration_ms":40394,"significance":"The question of whether lightweight CAI can align small uncensored models is timely and practically important. The study has some strengths: it directly compares four models under an abliteration procedure, uses a standardized prompt set (HarmBench), and includes capability benchmarks to monitor collateral damage from abliteration. However, the central comparison is currently unverifiable because the harmfulness classification protocol is not described, and the cross-architecture claim is confounded by a system-prompt difference for Gemma-2. If the authors supply a precise judge/rubric, per-prompt data, and error analysis, the paper could be a useful contribution to safety alignment for small models. At present, the headline conclusion rests on a metric that is not operationally defined.","major_comments":[{"comment":"The central dependent variable, 'Harmful Responses (%)' in Figure 1, is never operationally defined. The manuscript reports means such as 54.4% to 11.1% for R1-Llama and 87.8% to 38.9% for Llama-3.1, but it does not describe the judge, classifier, rubric, annotation instructions, or inter-annotator agreement used to label each initial and revised response. Without this information, the headline architecture-level conclusion cannot be reproduced or falsified. Please specify exactly how each response was classified, including the identity and version of the judge (human annotators, a specific LLM, a classifier), the criteria for 'harmful', and the exact judge prompt, and report raw per-prompt labels and agreement statistics.","section":"Section 2, Figure 1"},{"comment":"The cross-architecture comparison is confounded by a system-prompt asymmetry: all models except Gemma-2 received the same system prompt ('You are a helpful AI assistant' / 'You are an ethical evaluator'), while Gemma-2 received no system prompt because the authors found that prepending it increased harm scores. Because the paper's main conclusion is that Gemma-2 benefits less from CAI, this differential prompting condition alone could explain part of the observed difference. Please either use a system-prompt-capable chat template for Gemma-2, or explicitly analyze Gemma-2 as a separate prompting-condition comparison rather than pooling it with the other models in the architecture-level conclusion.","section":"Appendix A, Section 3"},{"comment":"The study uses only 90 HarmBench prompts (15 per category) and reports no statistical significance tests, confidence intervals, or per-prompt outcome data. The mean harm rates are presented with two decimals, which overstates precision for a sample of 90 Bernoulli trials, and the paper's qualitative claims (e.g., 'completely eliminating harmful content in several categories' in Section 3) are not supported by any error analysis. Please provide per-prompt labels, bootstrap confidence intervals or exact binomial intervals, and a test of whether the initial-to-revised changes differ across architectures.","section":"Section 2, Figure 1"},{"comment":"The CAI prompt wording was tuned via 'iterative testing' on the same HarmBench prompts used for evaluation. This introduces a selection risk: prompt choices that were adjusted after observing their effect on the 90 evaluation prompts can inflate the measured CAI improvement and make the architecture comparison hard to interpret. Please clarify whether the 90 prompts were also used during prompt development, and if so, consider reporting results with a held-out set of HarmBench prompts or describing the tuning procedure in enough detail to assess the risk of overfitting.","section":"Appendix A"}],"minor_comments":[{"comment":"The model is named inconsistently as 'DeepSeek-R1-8B' in the abstract and 'DeepSeek-R1-Distill-Llama-8B' in Section 2; use one name and state the base model explicitly.","section":"Abstract, Section 2"},{"comment":"The category labels are too small to read, and the per-category percentages are not listed; please provide a table with exact values and per-category denominators.","section":"Figure 1"},{"comment":"The statement that 'R1-Llama's overall harm reduction compared to its base model (Llama-3.1)' is misleading because R1-Llama is distilled from Llama-3.1-Base, not from Llama-3.1-Instruct; clarify this relationship.","section":"Section 3"},{"comment":"The column order ('Llama, DeepSeek, gemma, Qwen2.5') should match the order in the text and use consistent capitalization.","section":"Appendix B, Table 1"},{"comment":"Report the standard deviation (±31.60% vs ±49.02%) together with its definition (e.g., across categories or across prompts) and sample size, since the current text does not say what the spread represents.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"This is a borderline case: the topic is relevant and the empirical design has a clear aim, but the absence of any description of the harmfulness judge makes the central figure uninterpretable. If the authors cannot provide the judge's identity and the per-prompt labels, the paper should not be published in its current form. I would also look for evidence that the judge is not a Llama-family model, since that would directly bias the architecture comparison."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper makes a genuinely new empirical observation—CAI self-critique cuts harmful responses dramatically for two Llama-based 8B models, while barely moving the needle for Gemma-2 and Qwen2.5—but the headline metric is never operationally defined. I couldn't tell you what 'Harmful Responses (%)' means after reading the paper, and neither can anyone else. That's a load-bearing gap, not a polish issue.\n\nWhat's good: the study is honestly scoped. It abliterates four small instruction-tuned models to strip out refusal behavior, then applies the same three-step CAI protocol. The benchmark appendix is helpful—the fact that Qwen2.5 drops 20 points on CommonsenseMoral after abliteration while retaining SafetyBench performance is a genuine supporting observation. The prompt templates are given in Appendix A, and the authors disclose that they tuned the wording iteratively. That's the kind of transparency that makes this worth engaging with.\n\nWhere it gets soft: the central comparison rests entirely on Figure 1, and there is no description of how the harmful-response percentages were computed. No judge, no classifier, no rubric, no human annotation protocol, no agreement metric. If the judge was an LLM, the architecture split could be an artifact of the judge's own biases. If it was a human, that needs to be said. Either way, the numbers aren't reproducible as written. There's also a real confound: Gemma-2 received a different prompting condition (system prompt prepended to the user prompt, or none), and the authors say prepending raised harm scores—so the Gemma condition is not strictly comparable. The sample is 90 prompts across six categories; no significance tests or per-prompt error bars are reported. That's a minor-to-moderate weakness on its own, but combined with the undefined metric it makes the architecture-level conclusion provisional.\n\nOne more thing: the causal claim that reasoning capabilities explain R1-Llama's advantage is speculative. R1-Llama is a distilled reasoning model, but it's also a different training run; the comparison can't isolate reasoning as the active ingredient. The paper mostly frames this as a suggestion, so I'd treat it as a hypothesis for follow-up, not a flaw in the core result.\n\nBottom line: this is a useful data point for people doing safety alignment on small models, if the evaluation protocol gets cleaned up. It deserves a serious referee, but the referee should send it back for major revision: define the harmful-response metric, disclose the judge, run a significance test or at least report per-prompt agreement, and either match Gemma's prompting condition or discuss the confound openly. I wouldn't cite the current version in my own work, but I would bring it to a reading group to talk about what a reproducible small-model safety evaluation should look like.","headline":"New observation about architecture-dependent CAI effectiveness in small LLMs, but the undefined 'harmful response rate' metric makes the headline comparison unverifiable as written.","tokens_in":5856,"tokens_out":2355,"would_cite":false,"duration_ms":22765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Constitutional AI's generate-critique-revise loop sharply cuts harmful responses on Llama-based small models but does little for Gemma-2 and Qwen2.5, suggesting the method's benefit depends on model architecture and reasoning.","keywords":["constitutional AI","self-critique","abliteration","small language models","harm reduction","DeepSeek-R1","Llama 3.1","safety alignment"],"falsifier":"Rerun the exact 90-prompt generate-critique-revise pipeline with a specified judge (a written rubric applied by human annotators with measured agreement, or a frozen judge model) and check the revised-response harmful rates: R1-Llama around 11.1%, Llama-3.1 around 38.9%, Gemma-2 around 40.0%, and Qwen2.5 around 67.8%; the architecture-level claim fails if the ordering changes materially or if R1-Llama's large drop does not replicate, and the Gemma-2 result should be checked with and without its system-prompt workaround.","tokens_in":5005,"feed_emoji":"🛡️","tokens_out":5554,"duration_ms":48125,"temperature":0.7,"pith_summary":"This paper asks whether Constitutional AI's self-critique loop can make small uncensored language models safe again after their refusal behavior has been removed through abliteration. Testing four 7-9B models on 90 HarmBench prompts, it finds that Llama-based models (DeepSeek-R1-Distill-Llama-8B and Llama 3.1-8B) cut harmful responses by large margins, while Gemma-2-9B and Qwen2.5-7B improve only slightly. The claim matters because it suggests that lightweight, self-supervised alignment can work in resource-constrained settings, but only when the model's architecture and reasoning style cooperate. The paper concludes that CAI's effectiveness varies by architecture and that retained safety knowledge is not always applied during open-ended critique.","feed_headline":"Constitutional AI slashes harm in Llama models, barely in others","feed_subtitle":"Harmful replies drop from 54.4% to 11.1% on R1-Llama, but Qwen2.5 and Gemma-2 improve only slightly.","key_machinery":"The central mechanism is the Constitutional AI self-critique loop, implemented as three steps: generate an initial response, ask the model to critique its own response against a set of safety rules, and then instruct the model to rewrite the response in light of that critique, with refusal encouraged when the prompt itself is harmful. Abliteration is the second load-bearing mechanism: removing a single activation direction to suppress refusal behavior, so that the effects of CAI can be isolated from pre-existing safety training. The comparison across four architectures is what carries the conclusion that harm reduction is not uniform.","core_discovery":"The paper claims that after abliteration strips refusal behavior, CAI's self-critique mechanism produces architecture-dependent harm reduction: on 90 HarmBench prompts, revised harmful responses fall from 54.4% to 11.1% for DeepSeek-R1-Distill-Llama-8B and from 87.8% to 38.9% for Llama 3.1-8B, but only from 45.6% to 40.0% for Gemma-2-9B and 76.7% to 67.8% for Qwen2.5-7B. It further claims that failure patterns differ: Gemma-2 and Qwen2.5 usually fail to detect harm during the critique phase, while Llama-3.1 often adds warnings without removing the harmful content, and R1-Llama alternates between the two failure modes. The paper interprets this as evidence that the models retain safety knowledge (SafetyBench scores stay high) but differ in applying it during open-ended critique, and that R1-Llama's reasoning step contributes to more consistent harm reduction.","pith_inferences":["If architecture dependence is real, averaging a single CAI benefit across heterogeneous models is misleading; evaluations should report results stratified by architecture and by reasoning behavior.","The absent judge description is the largest unresolved risk: a different fixed judge or rubric could shift the borderline Qwen2.5 and Gemma-2 results enough to change the ordering, even if the Llama improvements hold.","Gemma-2's system-prompt incompatibility is a testable confound: re-running it with no system prompt or with an equivalent chat-template injection would show whether its small improvement is an artifact of the prompting workaround.","The paper's 'safety knowledge exists but is not applied' framing suggests a direct intervention: add explicit reasoning steps or chain-of-thought to the critique phase for Qwen2.5 and Gemma-2 and test whether the gap to Llama closes."],"forward_implications":["CAI can serve as a low-cost self-alignment method for small models in resource-constrained settings, at least for Llama-based architectures.","Architecture-specific prompting may bridge the gap between retained safety knowledge and effective self-critique, since SafetyBench scores stay high even where critique fails.","Abliteration has uneven side effects: Qwen2.5 loses 20 points on the CommonsenseMoral test and degrades on several knowledge and morality benchmarks, while Llama-based models barely change, so uncensoring and realigning is not architecture-neutral.","Reasoning-enhanced models such as R1-Llama may need distinct safety evaluation, because they show both stronger and more stable harm reduction than their non-reasoning base.","High scores on recognition-style safety benchmarks do not guarantee that a model will apply safety judgments in free-form critique and revision."],"supporting_citations":[{"why":"Defines the Constitutional AI critique-and-revise framework that the paper adapts for small models.","marker":"Bai et al., 2022"},{"why":"Supplies the abliteration technique used to remove refusal behavior before running CAI.","marker":"Labonne, 2025"},{"why":"Provides HarmBench, the source of the 90 harmful prompts used for evaluation.","marker":"Mazeika et al., 2024"},{"why":"Supports the claim that refusal is mediated by a single direction and that Qwen models degrade more under abliteration.","marker":"Arditi et al., 2024"},{"why":"Identifies the R1-Llama reasoning model and its distillation from Llama 3.1-8B.","marker":"DeepSeek-AI et al., 2025"},{"why":"Supplies SafetyBench, used to show retained harm-detection capability after abliteration.","marker":"Zhang et al., 2024"}],"fun_headline_variants":["CAI harm cut: Llama models win big, Qwen and Gemma stall","Self-critique drops harmful replies to 11% on R1-Llama, not on Qwen","Architecture decides CAI's safety payoff: Llama yes, others no","CAI helps Llama-based 8B models, barely moves Qwen or Gemma"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's headline comparison assumes the same valid judge scored every model's responses as harmful or not, but it never describes the judge or rubric; it also treats the four models as comparable even though Gemma-2 was tested with a different prompting setup because it does not natively support system prompts.","fun_headline_variants_meta":{"raw":{"variants":["CAI harm cut: Llama models win big, Qwen and Gemma stall","Self-critique drops harmful replies to 11% on R1-Llama, not on Qwen","Architecture decides CAI's safety payoff: Llama yes, others no","CAI helps Llama-based 8B models, barely moves Qwen or Gemma"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":2988,"prompt_tokens":909,"completion_tokens":2079,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1996}},"tokens_in":525,"tokens_out":2079,"duration_ms":17156,"temperature":1.0,"reasoning_tokens":1996,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T19:21:11.684656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the exact 90-prompt generate-critique-revise pipeline with a specified judge (a written rubric applied by human annotators with measured agreement, or a frozen judge model) and check the revised-response harmful rates: R1-Llama around 11.1%, Llama-3.1 around 38.9%, Gemma-2 around 40.0%, and Qwen2.5 around 67.8%; the architecture-level claim fails if the ordering changes materially or if R1-Llama's large drop does not replicate, and the Gemma-2 result should be checked with and without its system-prompt workaround.","supporting_citations":[],"review_version":1}